Thank you for registering for Data Science for Marketers 101. As session materials are processed, we’ll add them here.
Video
Can’t see anything? Watch it on YouTube here.
Audio
Slides
Transcript
Jessica Leitsch 0:00
We have Christopher Penn joining us who is the chief data scientist at Trust Insights dot add. Welcome, Chris.
Christopher Penn 0:56
Thank you very much. It’s a pleasure to be here. Thank you to the entire IBM team and community for having me very briefly about myself I am an IBM champion in data science and AI have been for three years now. And that’s enough of the introduction. If you would like the audio recording the slides the AI based transcript using Watson Watson natural language understanding, you can go to where can I get the slides calm submit your information, and we’ll you’ll get access to the materials as soon as they are available. Let’s start today talking about marketing’s three biggest problems. Generally speaking, marketers are being asked for the same things that we all are right, better, faster and cheaper and marketing is not doing a swell job of this. When you look at the 2019 cmo survey by PricewaterhouseCoopers CMOS CEO sorry, this is a CEO survey were asked what are your priorities for information that you wish you had that you view as critical to your business and then how comprehensive is the information that you are getting across the board, data about your customers and clients preferences needs. 94% of CEOs said this is critical CEO said, financial forecasts and projections 92% 90% for data about your brand and reputation and the list goes on. You can see them here. When CEOs were asked, how is the quality of what I’m getting? it’s appalling. 15% say they are getting comprehensive data about their customers and clients preferences and needs. 24% about data about brand reputation 22% about data to risks that the business is exposed. What this means is that marketing isn’t providing better information, better data, it’s just not something that we’re being particularly good at.
Unknown Speaker 2:46
Are we being
Christopher Penn 2:48
are we using the data Well, when we use data well as marketers, we generate outsized impact, think about how much money your company spends on things like social media, marketing and mobile Marketing. In the August 2019, cmo survey chief marketing officers survey, Duke University asked to what degree does use of marketing analytics contribute to company performance? You know, one on a scale of one to seven one not at all seven very highly. And cmo said 4.1. It’s it’s a reasonably big contributor to company performance. More so than social media, more so than mobile marketing. Where are we putting our money? Where does your company spend money? How much? How much the line item for social media marketing, including ads, how much is the line item for mobile marketing? How much is the line item for analytics for data? I guarantee you it’s much smaller. In the most recent February 2020, cmo survey, I think the spend number was around 5%. For analytics and data. It’s really, really small. Why? Well, a couple of reasons. One, we’re not necessarily comfortable using the data and to even when we are comfortable with Working with data, we don’t make use of it to make decisions. When asked cmo said, what percent of the time do you use marketing analytics for decision making being data driven, you let the data do the driving in August 39.3%, in the February survey, which just came out 37.1%, so just slightly more than one third of the time, the other two thirds of the time, we’re guessing, making things up going with our bad experience or the hippo problem, the highest individually paid person’s opinion in the room. And so this leaves us in a really bad situation. So how do we get out of this? Well, we have to use we should use data science as one of the main components for making our marketing more effective. But what is this stuff? Oh, there are a ton of definitions out there, some of which are coming from very official sources like IBM for example. But what is data science? The simplest way to put it Is that it is meaningful insights from data using the scientific method. This is the essence of data science. And data science is composed of four main parts, business skills, scientific skills, mathematical skills, and technology skills. All four of these when put together and combined to form data science. Because people love especially in marketing, love to put cool labels on things we should probably spend some time talking about what data science is not what isn’t data science, data science is not five things. First, it is not analytics. Analytics is the process of unlocking value from your data from being able to say what happened, right, literally comes from the Greek word on the line, which means to unlock to loosen to shake something loose, you have a thing you shake it and stuff falls out. But that’s to, uh, to loosen. That’s what analytics means. Data Science is not analytics. Analytics is its own discipline. It’s very important. It’s critical for business, but it’s not data science. It’s just the process of analyzing what happened. Data Science is not reporting, reporting and data visualization. Again a central practices are not saying that these are less important, they are not they are equally as important as data science, but it is not reporting recordings, visualization and communication of data, you have to be able to say to somebody, hey, this is what we see this is what’s happening in a format that they can understand it. critical skill, not data science. Data Science is not data engineering. Again, separate profession the ability to work with systems, databases, and other data storage technologies is part of data science, but the management of those systems and Such are the domain of data engineering, data science is not statistics. Statistics is its own profession, discipline school of thought, very different data science uses and leverages a tremendous amount of statistical techniques, everything from simple linear regressions to gradient boosting and all these really cool things, but they have fundamentally different outputs. And of course, even though data science in the IBM portfolio falls under the data and AI collective data science is not artificial intelligence. The fundamental output of AI is a piece of software. It’s a model, right? It’s it’s a software that a machine wrote, The fundamental output of data science is a proven or disproven hypothesis. Now the hypothesis could become
the model. And so that’s why you hear data science and AI grouped together a lot and thought, but they are fundamentally different things as important to keep that in mind because if you don’t keep those two separate, you can be conveniently about what it is you are doing. So those are the things that data science is not we just want to be very clear about that. So that when we’re talking with our colleagues, with vendors, we know what it is we’re talking about. By the way, that knowledge also is super helpful for being able to tell whether a vendor is, is pulling the wool over your eyes. Why do we care as marketers, especially since a lot of people got into marketing, because they wanted to be creative. They wanted to explore their creative capabilities. And they were for those who have a lot of gray haired might remember the old Saturday Night Live sketch with Chevy Chase, where he’s playing a presidential candidate and says my understanding there would be no math on this debate. And a lot of marketers have that same feeling when they come to to marketing analytics and stuff like I was my understanding would be no math and marketing. That was why I didn’t take math in college. Well, it is part and parcel of data science. So why do we care? data science gives us four things helps us accelerate decisions make our decisions go faster. Being able to process a large amount of data and reach a conclusion sooner rather than later, is one of the key benefits. Second, data science helps us lower our costs. I am astonished as a marketer. How many people will say, yeah, we’re going to spend a million dollars on this campaign. But when you say, well, let’s take $25,000 and do an experiment, come up with a hypothesis, test it, run a limited pilot and then prove or disprove that the thing that we’re talking about actually works and stakeholders like nope, nope, not doing that. She’d rather risk a million dollars, then invest $25,000 to find out whether something works. Data Science is one of the pathways to help reduce costs by lowering risk. Data Science helps us beat competitors, especially those competitors don’t have data science capabilities of their own, to be able to find their way weak spots when you’re looking at all the data, SEO data, social media data, email marketing data, if you’ve got it, and you’ve got apples to apples with you and a competitor, you can find weaknesses that they may or may not be aware of, and take advantage of it. This is especially true in SEO, SEO is one of the best things for marketers, because you can see really good apples to apples data a lot and get a very clear understanding of what a company is or is not doing to be successful. And finally, data science gives us new opportunities, the ability to find things that we would not have seen otherwise, unless we were out there testing hypotheses, and building techniques in order to understand what’s in our data. The classic example I use here is the children’s toy line, My Little Pony. Now, for those of you who don’t know, these little like little plastic courses that have names and characters and stuff like that the standard marketing instinct would be we’re going to market this to kids Predominantly female 18, eight to 14 years old. But if you were to put aside your assumption, and use data and explore it, you would have found doing a technical exploratory data analysis, that there’s an entire population of predominantly people who identify as male, aged 26 to 40. We’re spending a gigantic amount of money, because they’re this this subculture called bronies. Guys who love this stuff. And if you were Hasbro and you didn’t know that this was a thing, you would be ignoring an entire market. And so it’s really important to use data science to find those hidden opportunities that can be very, very lucrative. So let’s start digging into these four areas. What do you need to have? What skills do you need to have in each of these four areas? Let’s tackle business skills first, what kind of business skills do you need? Fundamentally, a data scientist has to be able to look at
Unknown Speaker 11:59
data
Christopher Penn 12:00
problem that you’re facing and ask the stakeholders what decision will be made from this? What decision are we making from this data and information? Our rule is data without decisions is distraction. It means nothing. If you got all this data and you’re not making a decision from it, it’s a waste of time to mess around with it, right? It’s fun if you’re a hobbyist, or if you are purely in an r&d role. But if you’re in a marketing role, which I suspect many of us here are, you need to have a clear idea of what decision it is you’re after. Let’s talk about an exercise for this something we call KPI mapping. When you look at all your marketing data, it can be tough to figure out like where should I spend my time? Right, what’s important. The way to get at what’s important is to start with your overall strategic objective, and then work backwards from it. So let’s say your company cares about revenue. I can’t think of many Companies or marketing initiatives that don’t have some kind of revenue component, the exception being the election of political officials. There’s a binary outcome, but there’s no revenue there. Right, you start with revenue. So the next question is what drives revenue? In a lot of technology companies, b2b companies, it’s typically something like closed one deals. And you say to yourself, okay, well, what number or numbers immediately get closed one deals, open deals, you need to have deals in order to close them. Where do we get our deals from? sales qualified leads, where to sales qualified leads come from marketing, qualified leads, where’s that come from prospects? How do we get prospects? Well, we need prospects we prospects to get. We need suspects to get that people who are not sure we’ve identified where we get that from audiences, people who have attracted to our media properties. Great. Where does that come from? The people who could buy our stuff. So what we’ve done here in a very simplified way, is walk through the different numbers that map to that overall KPI or that overall goal, the business goal of revenue. And then the KPIs in here depend on who you are in your organization. If for example, I am the VP of sales, I am measured on closed one deals, right that is the most important number it is not the goal of the company but as the most important number for me to understand as a person. If I am the brand manager, I care about my brand being able to generate audiences. If I am the VP of Marketing, I care about marketing qualified leads, that’s what will help me get my bonus. A KPI is a number for which you will either get a bonus or fired for a KPI is a number for which you get a bonus or fired for. And so when we’re doing data science, a data scientist has to be able to look at the different metrics are available, identify which ones are KPIs, typically by figuring out which one Listen, this is assigned to right? Audience is not a KPI for the sales manager. Close one deals is not a KPI for the CMO. Having deep insight into the business problem you’re trying to solve is how you create success as a data scientist. Without that, you’re just messing around. And again, that’s fine for r&d, but not first, trying to solve real business problems that pay your paycheck. So, if you are aspiring to be a marketing data scientist, you’ve got to have a very clear understanding of your business, how it makes money, how your business works, how who is responsible for things within your organization, where they get their data from. And it all starts from these business skills. The second area that we need to have fluency in as data scientists is scientific skills, scientific skills, all revolves around answering the question is, whatever it is we’re investigating a reproduce usable answer. We are seeing a tremendous amount of discussion right now, because of current events, about research papers and findings and scientific literacy. And it’s astonishing the number of people who do not have any kind of scientific literacy to be able to look at a study or a piece of research and say, Yep, I could see how we could reproduce this study and verify its results or not think about your own marketing if you just do some changes on the website. And you’re testing, you know, buttons and colors and all this stuff. Is are your findings, reproducible? The way you get to reproducible research is through the use of the scientific method. Remember, we talked about this at the beginning data science is meaningful insights from data using the scientific method. If you are not using the scientific method, you are not doing data science. If you’re not using the scientific method, you are not doing data science. starts with asking a question.
What are we trying to solve for right? Which email will perform better? We define the problem and the data we’re going to need, click through rates, open rates, size of the list we sent, click to open rate, etc. And then we create a prediction hypothesis. A hypothesis is a provably true or false single condition. From there we test, collect, analyze, refine our hypothesis, discard it, observe, and then come up with new questions. You have sort of done this already. I guarantee it. You’ve sort of done it. If you’ve done any kind of AV testing as a marketer, if you’ve ever tried to do a website test, social media test, email test, you’ve done this. And the part that people get wrong the most is the hypothesis. Like we’re going to test so try and make our conversion rates better. That’s not a provably true or false statement. Right? That is just kind of a mess. That can be your question, how do we improve our website conversion rates. But our hypothesis has to be something like, we believe that red buttons on the website will increase conversions by 6%. That is a provably true or false single condition that you could then build a reliable test from. Now, the good news is a lot of the the analysis and the collection of data and such is being done by software for you. So you don’t necessarily have to do that part. This is an example of the Free Software Google optimized for doing website testing. And in this case, I had a hypothesis about different things on my website, I wanted to test certain buttons and certain pieces of text and to create a nice multivariate test for me and proved nothing. In the end, it said sorry, there was we couldn’t get enough data that was statistically significant. That a clear lead was found like okay, so that tells me that there’s Some things on my website I can change that will have no meaningful consequence. But I needed to have a provably true or false conditional in this case, a set of conditions in order to be able to do that. Again, I can’t emphasize this enough, if you are not using the scientific method, you are not doing data science, you’re doing something else. A third area of a mathematical skills that you need in data science to be a a functional data scientist. We have to be able to look at our data and decide whether the things we’re doing are mathematically sound and valid. And there’s four core groups of skills that I think are important when it comes to this the ability to look at and do sensuality, distribution, regression and clustering on top of things like math, just basic addition, subtraction, multiplication and division. But more than anything, there’s a certain amount of contextual literacy that comes with your data. Almost every number that you look at in data science requires some kind of comparison and contrast. Here’s an example. Here’s some Google Analytics data from my website. 3343 users 330 100 new users. And here’s the graph. What does this tell you? doesn’t tell me anything, right? There’s no contrast. There’s no comparison. There’s no context to this information is just a bunch of information on screen. Now, if I were to add in last year, the same period of time, year over year, now, I see something I’m down 2% of users down 3.56% on users, the previous year was better. Now I’ve done that contrast. I have a way to start thinking about the state of that makes you go Hmm, is that decline statistically relevant? If it is, why? And if I if there is a reason why is it something I should be fixing something I can fix it So that ability to think about not just the numbers and the data, but the context around the data is so essential as the one of the key mathematical skills you need as a data scientist. There are four sets of techniques, as I mentioned, the first is centrality, to be able to look at data and look at the centrality of it. Let’s take that Google Analytics data. Here it is in a spreadsheet, right? Looks the same just turn outside. What if we rearrange it top to bottom by the number of visitors or number of pageviews per day? We have one day That was really good. And then one day was really bad, and a whole bunch of data in between. If I were to apply a measure of centrality to this data, or several, I could start to learn things about it. So we have the three basic ones from stats one, a one mean, median, and mode, the mean is the arithmetic average, what’s the average number of pageviews the median is the middle number, which in this case is right here in the middle of the table between December 15 and January 3. And then mode is the most frequently occurring number.
When we see these numbers again, with back comparison and contrast, we can start to make some interesting conclusions about this data, the mean is bigger than the median. And the further ahead of the median the mean is, the more it means there are outliers to the upper end of the data set. Here we have one day where there’s a pretty big outlier. But now imagine two or three or four of those days, or one day where there was you know, some 1000 is 10,000, or 100,000. That tells us that there are more anomalies towards the top of the data set, which in marketing is typically indicative of you running very successful campaigns, a campaign, a big splash, a big launch happened and your mean pulled away from your media. On the other side, if you went the opposite way, and your mean was below your meeting. I mean, you have more hours to the bottom things are really broken. And so even just something as simple as having the mean in the media and on a very simple dashboard will tell you, huh I can see what’s happening with my marketing When the mean and the median are close together, it means there aren’t a whole lot of outliers, which if you’re not doing anything with generate outliers, if you don’t have any major campaigns, it’s not a bad thing. But if you’re running a whole bunch of campaigns, you’re pushing really hard and you don’t have outliers to the upper end, means your campaigns are not being successful. So just the simple measures of centrality help give us insight into our data.
second skill. second technique distributions. If I take that same Google Analytics data and put it in buckets, say buckets of roughly 50 visitors, if you look at I put this in a little bar chart by those bins, and draw a line over the top to try and capture the general curve of it. You can see it’s kind of leaning towards the left, this leaning left or right in a distribution, something called kurtosis. And it again, gives us an indicator about how our website is doing. And again, this applies to social media data, SEO data, email, data, you text, message, data, whatever you’ve got You would apply these very simple mathematical and statistical techniques. When you have left leaning kurtosis. That means that you have more data on the lower end than the upper end, it means your marketing is slowing down. And that’s bad. If you’re leaning towards the opposite direction, you’re leaning towards the right side. Ah, things are good. You have more traffic on the operands. So again, this would be a very quick and easy way to visualize how is your marketing doing for any period of time. A third technique, regression. regression, comes in a bunch of different flavors. But fundamentally, it is about taking two or depending on the regression technique more than two data series and comparing the relationship between the two. Here we see website users and then impressions on Twitter for those same days. I put on a regression line, a simple linear regression line and I found there is no relationship What does that tell me was a marketer tells me that what I do on Twitter, and what gets me website traffic are not related. Hmm. Okay, so all that time and effort I’m putting in on say, you know tweeting 18 times a day and using little cute emojis and things. It’s not delivering the website traffic. If that was something I cared about, we could easily put on other variables like conversions or lead fills or emails, whatever. But the regression technique lets us understand the relationship between two different data series. There’s three varieties there’s there’s Pearson, spearmint, and Kendall Tao. You’ll get into more of that and about which one to use in each situation when you’re not the one that’s more of the 201 level. But just knowing that you can look at two data series and see their relationship to each other. If this was a line you set a line on the graph is kind of flat if it was a line going for the lower left hand corner. to the upper right hand corner, that would be a perfect regression. It would be you know, as one series increases, the other series increases with it in step. The opposite is true for this line go from top left down to the lower right, as one goes up, the other goes down. This blind flat, no relationship. Right. So it’s so this is a this testing this information out, configure. Yeah, yeah, no relationship. The fourth technique is clustering, an awful lot of the time, the data we have, can kind of clump together. And it’s informative to dig into that data and understand what is happening are the things that happened together. If we take this data from the same Twitter, the same Twitter and Google Analytics data, we see that they actually are four discrete, different groups. There’s that dark blue group, which is cluster one that all happens to go together is cluster two, cluster three, cluster four When we look at our different groups, we can start to understand Okay, what is this? What are these things have in common that they’re so closely related? And it gives us more nuance into something like a regression. In the big picture, maybe there is no relationship, but there are there certain cases where the Twitter content and website traffic are related. In fact, if I slapped some regression lines, they would see Actually, yeah, there are a couple of cases where there’s a lot of relevance cluster for what’s different about that. So this is a technique that helps us ask more questions, dig deeper into our techniques. If you’re a user of IBM Watson Studio, or SPSS modeler, these techniques are built right into the interface. You’ll find them in the left hand side, regression, the bidding, you’ll find under the histogram widget inside of field, the field widgets, you will find the regression under models and you have use linear linear as and a bunch of other regression models. You’ll find that in Under the modeling side, you will find the mean median mode under basic arithmetic. And clustering is also this k nearest neighbor clustering KNN. Inside the modeling, the modeling section of SPSS modeler and Watson Studio. So if you want to do these techniques with your data you can do right inside of Watson Studio.
Speaking of which, let’s talk about the last area, which is the technology skills that you need in order to be an effective data scientist. And this is not going to be a comprehensive list because the field is changing like crazy. But those technology skills form the underpinning of being able to do a lot of this work. The key question around technology that we want to be able to answer with as data scientists is is what we’re doing scalable and automated, right? repeatable is important and the scientific method helps us get repeatable, scalable is the part that really matters because we need to be able to understand the company competent within eventually master technologies that allow our work to scale. Because while some marketing data is very small, like, you know, number of email opens per week or something, some marketing data is really, really big and exceeds the ability for our tools to process it. Well. what this looks like is kind of an evolution. In the beginning, you’ll spend a lot of time with spreadsheets, and that’s okay. There’s absolutely nothing wrong with spreadsheets, there’s a reason why it’s one of the most popular pieces of software in the world. This is great for early exploratory data analysis messy with subsets or samples to see Is there a thing in here that that is worth pursuing? And it’s so important to be able to do and know the basic functions inside of a spreadsheet? Can you do mean, median and mode? Can you do a simple linear regression? If you can, great. The problem is, this does not scale well. This and this is very hard to repeat. You have to kind of redo everything every week that you’re you’re working on your data. In cases like that, you’re going to want to graduate sort of at the intermediate level to something like Watson Studio, where you’ll be able to take your data, bring it in automatically. And then once you build your workflow, let Watson Studio run it for you and deliver you the same level of analysis, same quality of analysis that you would normally do. But do it in an automated fashion, you can even deploy it as a model and have it just kind of run all the time. And then one level up from that if you’re if if this is, you know, to a one level, then the graduate school level is when you start to write your own code. When you jump in, and you say, you know what, I need to do something very specialized, so specialized, that that it’s just not built in or I need to be able to work with an open source library. That simply is not available in the package traditions or anyone software, you’re going to use languages like Python, or this case, our studio and our statistical language to do that analysis. Again, it’s so important to have an environment that supports ports this and that’s why we strongly recommend for most people most of the time that want to learn data science, sign up for a free account from from IBM, try out Watson Studio and start to learn these different pieces. You’re going to be at roughly three different skill levels in your career right your beginning tools that we think are great jumping off points, but spreadsheet. Google Data Studio is an excellent, basic free tool, just learn the basics of data visualization. Google Analytics is a great free data source for you as a marketer, you probably already have access to it. And these are great ways to start digging into the data science process. At the intermediate level, you want to look at Watson Studio. Tableau software is also excellent. So is IBM Cognos, and SPSS modeler. Those are all great intermediate tools for marketing data science. And then, at the high end, the SpaceX version. You’re going to be creating your own code in our in Python users formats like JSON or command lines, SQL and no SQL databases. Those are the technology skills. I want to provide a bit of caution here. There are a lot, a lot A lot of these, you know, six week crash course in data science curriculums. those courses tend to lean very heavily towards the technology somewhat on the mathematical and omit the scientific skills and omit the business skills. And what they end up creating is a person who’s got a few basic skills which a good important, great jumping off point. But they’re not data scientists. would you go? Would you go to a doctor who’s took a six week crash course in heart surgery? I don’t think I would, I think I’d want the doctor who went through all four years of medical school and did their residency and and they are qualified to be doing surgery. Like that.
There’s a lot of people out there who are bearing data science titles, and resumes that you need to do a bit of digging on. I want to talk about some of the things that are not going to be in a course, but not going to be a certificate that you’re going to need. And those are the soft skills for data science, there’s seven of them that are so critical. The first is you have to be open, you have to be an open person who can communicate well with other people in order to be an effective data scientist because if you’re, if you’re very closed off person, you can’t communicate the value of what it is you’re doing. You can’t communicate your research to somebody else.
Unknown Speaker 33:43
That’s
Christopher Penn 33:45
that’s a real, real distinct problem. I’ve seen that happen with a lot of folks who are brilliant, brilliant, brilliant people, but not open. The second, I would say probably, you know, very, very important personality. You have to to be resilient, you have to be okay with getting things wrong an awful lot of the time you have to be okay with screw ups with experiments blowing up in your face, and not just be okay with it, but then get right back up and keep going. That’s the resilience getting punched in the face, you just get back up and you keep on going. It’s a personality trait that some people don’t have, you know, they have a failure and it just crushes them. And that’s not a personality trait that does well in this particular role. Third, you have to be curious, you have to be willing to want to dig deeper to not settle for the first answer you get and this is a flaw in most business today. We are as marketers as executives we are so caught up in like just can’t get things off my list guy could figure something out. Let’s get things done, get stuff done. You know, there’s courses and this whole people lecturing up there, but how do they get stuff done? seminar but that makes you incurious So you don’t care about the answer. You just need to get the answers you can get something off your to do list. A great data scientist, like Hold on, something doesn’t smell right here. Let’s keep digging on this even though we’ve got an answer may not be the answer. Let’s keep digging. It’s that curiosity is so important. Number four, somebody has you have to be patient. As a data scientist, this stuff takes time, great science takes time. Again, current events are actually very illustrative of this. There are a whole bunch of people who are clamoring for answers to how a virus works and when a vaccine will be available. And people who know science well saying it’s gonna be a while. It’s gonna be a long while because to do it well, to do it, right. When literally lives are on the line, you must be patient. And being able to gauge whether somebody is patient or not, is a critical, soft skill. Number five, you have to be persistent. So resilient, you get punched in the face and get back up. But after a while, you may not be like you know, I’m kind of tired of getting punched in the face. A persistent person is willing to keep doing Over and over again until they get the right answer until they get to where they want to go. And that’s a key critical personality trait. Number six, data scientists have to be humble. This is what I struggle with personally. I’ll say that right now. A lot of the work that we do in data science informs other parts of marketing other parts of business. And it means that guess what somebody else gets to take credit for your work, your work enables their work, they get to look like the stars, they get to be in the spotlight, and you’re like, well, I helped. Right? And you have to be okay with that. You have to be like, you know, I’m okay with somebody else’s naming and lights, as long as you take value and joy in the work itself being humble. And the last trait, and it’s so overused, it’s such a trope, but it’s also completely totally true. You have to be passionate about data science, you have to be passionate about the tools, the math, the business, the scientific method, you have to love this stuff. You have to love it so much. much that you do in your spare time for fun, right? I do a show on Saturday nights on Facebook called Saturday night data party. It is literally me typing in, you know, one or more statistical pieces of software. You can find it on YouTube, if you want. just messing around playing around there like it is there there there this past Saturday, I was looking at SEO data, and whether certain topics had a relationship to organic traffic for something completely unrelated to anything I’m working on right now just because it’s fun. I wanted to see Is there something I love doing this stuff so much that it’s how it’s entertainment for me. That’s a personality trait that you have to look forward to. And again, the reason you need to know these is threefold. One, you need to know it to see if you are a good fit for this. If being a data scientist is important to true, if you are hiring data scientists, these personality traits are things you have to screen And in the interview, because you may find, you know, what a six week Crash Course folks, and they may be technically competent. But they may not be humble, they may not be resilient, they may not love this stuff. And so you may end up making a bad hire if you’re not careful. And three, if you want to assess the likelihood of success of a Data Science Initiative within your company, your company has to have these values to right, if a company
is in curious, if a company is impatient, if a company is not open to sharing and communication of company, and particular the people you work for, they’re not humble. Your data science initiatives are doomed. Absolutely doomed. I have worked with and for a number of companies where some of these personality traits have been deficient and it has never gone well. Ever, never ever, ever gone well. So Understanding the conditions knowing yourself knowing your company, knowing the people you’re working with and trying to hire. These are essential soft skills for data science that cannot be understated. Finally, you need to understand, at least at a broad level, the data science lifecycle and it is gigantic. And the reason for it is because there’s a lot of steps to data science when you do properly. defining your problem ingesting data, analyzing it, repairing it, cleaning it, preparing it for an out for further analysis, augmenting it with additional data sources exploring it, comparing your explorations, creating predictions, hypotheses, prescriptions, creating models, sometimes validating those models refining your hypothesis, deploying a model if you’re doing it in concert with an AI group and then observing it. This life cycle is multi step and each one of these is practically its own profession in and of itself. Being able to see the whole thing and see what things on here. Do Do I know what things I hate? Do I not know which of these things should I learn more about is an essential part of your growth as a data scientist, if there are parts in here where like, I don’t even know what that word means, or I don’t even know how to do that. That gives you a very good roadmap for understanding how to improve your professional development and your professional skills as a data scientist. So how do you get started? Look at those four areas. Which of those areas do you or your company need the most support in? And then how do you obtain that talent? Somebody who is really good at all four of these is kind of a unicorn, and they’re going to be expensive, like hundreds of thousands of dollars a year expensive. What is realistic for most companies is having a team of folks who have discrete skills and then you have a very, very strong management system with it. So you may have a stats and math expert who may you may borrow from the stats team if you have a statistic Statistics team, you have a code or you borrow from it. You have someone who has good scientific mindset, maybe who does a lot of conversion rate optimization. You have your business expert to domain experts, people who know your industry. So well, they can provide nuance and context. And then data engineering, you may borrow from it, you can create data science capability. As long as you understand how the different pieces relate and interact with each other. You don’t have to try and cram it all into one person. It it typically, like I said, they’re very, very hard to find qualified data scientists themselves. give you a sense, I believe it was IBM that did the study in 2017. There were at the time 8500 people who were qualified data scientists with four years or more of experience in the field. And there were 14,000 marine biologists so more people know about whales than data science. This gives you a sense of how rare this talent is. So in a lot of cases, you’re going to be sewing it together. And that’s okay as long as you understand what the things are that you’re supposed to be looking for and how they relate to each other. But most of all, to circle back from where we started, it’s all about understanding the data sciences, the meaningful insights you derive from data, using the scientific method. If you’re not using the scientific method, you’re not doing data science, if you’d like, again, as I said at the beginning, if you’d like to get the slides and the audio and stuff when we post up in just a little while, the slides are actually available right now if you go to where can I get the slides comm there are PDF files, things like 20 megabytes. And if you’d like a book to read about this, you can read AI for marketers books calm.