2019年的经典资料,400多页,内容精彩不容错过!
Over the past decade, business analysis has been disrupted by a new way
of doing things. Spreadsheet models and pivot tables are being replaced by
code scripts in languages like R, Scala, and Python. Tasks that previously
required armies of business analysts are being automated by applied scientists
and software development engineers. The promise of this modern brand of
business analysis is that corporate leaders are able to go deep into every detail
of their operations and customer behavior. They can use tools from machine
learning to not only track what has happened but predict the future for their
businesses.
This revolution has been driven by the rise of big data—specifically, the
massive growth of digitized information tracked in the Internet age and the
development of engineering systems that facilitate the storage and analysis of
this data. There has also been an intellectual convergence across fields—
machine learning and computer science, modern computational and Bayesian
statistics, and data-driven social sciences and economics—that has raised the
breadth and quality of applied analysis everywhere. The machine learners
have taught us how to automate and scale, the economists bring tools for
causal and structural modeling, and the statisticians make sure that everyone
remembers to keep track of uncertainty.
The term data science has been adopted to label this constantly changing,
vaguely defined, cross-disciplinary field. Like many new fields, data science
went through an over-hyped period where crowds of people rebranded
themselves as data scientists. The term has been used to refer to anything
remotely related to data. Indeed, I was hesitant to use the term data science in
this resource because it has been used so inconsistently. However, in the domain
of professional business analysis, we have now seen enough of what works
and what doesn’t for data science to have real meaning as a modern,
scientific, scalable approach to data analysis. Business data science is the
new standard for data analysis at the world’s leading firms and business
schools.
This resource is a primer for those who want to gain the skills to operate as a
data scientist at a sophisticated data-driven firm. They will be able identify
the variables important for business policy, run an experiment to measure
these variables, and mine social media for information about public response
to policy changes. They can connect small changes in a recommender system
to changes in customer experience and use this information to estimate a
demand curve. They will need to do all of these things, scale it to company-
wide data, and explain precisely how uncertain they are about their
conclusions.
These super-analysts will use tools from statistics, economics, and machine
learning to achieve their goals. They will need to adopt the workflow of a
data engineer, organizing end-to-end analyses that pull and aggregate the
needed data and scripting routines that can be automatically repeated as new
data arrives. And they will need to do all of this with an awareness of what
they are measuring and how it is relevant to business decision-making. This
is not a resource about one of machine learning, economics, or statistics, nor is it
a survey of data science as a whole. Rather, this resource pulls from all of these
fields to build a toolset for business data science.
This brand of data science is tightly integrated into the process of business
decision-making. Early “predictive analytics” (a precursor to business data
science) tended to overemphasize showy demonstrations of machine learning
that were removed from the inputs needed to make business decisions.
Detecting patterns in past data can be useful—we will cover a number of
pattern recognition topics—but the necessary analysis for deeper business
problems is about why things happen rather than what has happened. For this
reason, we will spend the time to move beyond correlation to causal analysis.
This resource is closer to economics than to the mainstream of data science,
which should help you have a bigger practical impact through your work.
We can’t cover everything here. This is not an encyclopedia of data
analysis. Indeed, for continuing study, there are a number of excellent resources
covering different areas of contemporary machine learning and data science.1
Instead, this is a highly curated introduction to what I see as the key elements
of business data science. I want you to leave with a set of best practices that
make you confident in what to trust, how to use it, and how to learn more.
I’ve been working in this area for more than a decade, including as a
professor teaching regression (then data mining and then big data) to MBA
students, as a researcher working to bring machine learning to social science,
and as a consultant and employee at some big and exciting tech firms. Over
that time I’ve observed the growth of a class of generalists who can
understand business problems and also dive into the (big) data and run their
own analyses. These people are kicking ass, and every company on Earth
needs more of them. This resource is my attempt to help grow more of these
sorts of people.
The target audience for this resource includes science, business, and
engineering professionals looking to skill up in data science. Since this is a
completely new field, few people come out of college with a data science
degree. Instead, they learn math, programming, and business from other
domains and then need a pathway to enter data science. My initial experience
teaching data science was with MBA students at The University of Chicago
Booth School of Business. We were successful in finding ways to equip
business students with the technical tools necessary to go deep on big data.
However, I have since discovered an even larger pool of future business data
scientists among the legions of tech workers who want to apply their skills to
impactful business problems. Many of these people are scientists: computer
scientists, but also biologists, physicists, meteorologists, and economists. As
machine learning matures into an engineering discipline, many more are
software development engineers.
I’ve tried for a presentation that is accessible to quantitative people from
all of these backgrounds, so long as they have a good foundation in basic
math and a minimal amount of computer programming experience. Teaching
MBAs and career switchers at Chicago has taught me that nonspecialists can
become very capable data scientists. They just need to have the material
presented properly. First, concepts need to be stripped down and unified. The
relevant data science literature is confused and dispersed across academic
papers, conference proceedings, technical manuals, and blogs. To the
newcomer it appears completely disjointed, especially since the people
writing this material are incentivized to make every contribution seem
“completely novel.” But there are simple reasons why good tools work. There
are only a few robust recipes for successful data analysis. For example, make
sure your models predict well on new data, not the data you used to fit the
models. In this resource, we’ll try to identify these best practices, describing
them in clear terms and reinforcing them for every new method or
application.
The other key is to make the material concrete, presenting everything
through application and analogy. As much as possible, the theory and ideas
need to be intuitable in terms of real experience. For example, the crucial idea
of “regularization” is to build algorithms that favor simple models and add
complexity only in response to strong data signals. We’ll introduce this by
analogy to noise canceling on a phone (or the squelch on a VHF radio) and
illustrate its effect when predicting online spending from web browser
history. For some of the more abstract material (e.g., principal components
analysis), we will explain the same ideas from multiple perspectives and
through multiple examples. The main point is that while this is a resource that
uses mathematics (and it is essential that you work through the math as much
as possible), we won’t use math as a crutch to avoid proper explanation.
The final key, and your responsibility as a student of this material, is that
business data science can only be learned by doing. This means writing the
code to run analysis routines on real messy data. In this resource, we will use R
for most of the scripting examples.2 Coded examples are heavily interspersed
throughout the text, and you will not be able to effectively read the resource if
you can’t understand these code snippets. You must write your own code and
run your own analyses as you learn. The easiest way to do this is to focus on
adapting examples from the text, which are available on the resource’s website at
taddylab.com/bds.
I should emphasize that this is not a resource for learning R. There are a ton of
other great resources for that. When teaching this material at Chicago, I found
it best to separate the learning of basic R from the core analytics material, and
this is the model we follow here. As a prerequisite for working through this
resource, you should do whatever tutorials and reading you need to get to a
rudimentary level. Then you can advance by copying, altering, and extending
the in-text examples. You don’t need to be an R expert to read this resource, but
you need be able to read the code.
So, that is what this resource is about. This is a resource about how to do data
science. It is a resource that will gather together all of the exciting things being
done around using data to help run a modern business. We will lay out a set
of core principles and best practices that come from statistics, machine
learning, and economics. You will be working through a ton of real data
analysis examples as you “learn by doing.” It is a resource designed to prepare
scientists, engineers, and business professionals to be exactly what is
promised in the title: business data scientists.
+++++++++++++++++++++++++++++
【课题项目申请申报】马列和D史D建-国家社科基金项目申请书标书范文
https://bbs.pinggu.org/thread-16671251-1-1.html
【课题项目申请申报】社会学、人口学与民族学-国家社科基金项目申请书范文
https://bbs.pinggu.org/thread-16671249-1-1.html
【课题项目申请申报】法学和政治学-国家社科基金项目申请书范文
https://bbs.pinggu.org/thread-16671245-1-1.html
【课题项目申请申报】管理学、教育学和艺术学-国家社科基金项目申请书范文
https://bbs.pinggu.org/thread-16671243-1-1.html
【课题项目申请申报】文学、历史学与哲学-国家社科基金项目申请书范文
https://bbs.pinggu.org/thread-16671240-1-1.html
【课题项目申请申报】经济学-国家社科基金项目申请书范文
https://bbs.pinggu.org/thread-16671238-1-1.html
【基金项目申请】国家社科、国家自科、教育部人文、教改、课程思政、一流本科课程资料
https://bbs.pinggu.org/thread-16315330-1-1.html
【课题申请资料】项目申请书、课题申报书研究框架图合集
https://bbs.pinggu.org/thread-16309537-1-1.html
【课题申报资料】课题项目申请书配套插图、框图、流程图、技术路线图合集
https://bbs.pinggu.org/thread-16308981-1-1.html
【课题基金】1050幅科研项目课题申报申报书插图合集(逻辑结构图、数据示意图、动画)
https://bbs.pinggu.org/thread-16666695-1-1.html
【课题项目申请】2022—2025国家社科基金课题申请书范文
https://bbs.pinggu.org/thread-16620822-1-1.html
【课题项目基金】国家社科基金与教育部人文社科项目申请书申报书标书范文经验合集
https://bbs.pinggu.org/thread-16545433-1-1.html
【重磅课题资料】省中等职业教育教学改革项目课题资料汇总(申请书,开题,中期,结项
https://bbs.pinggu.org/thread-16437471-1-1.html
【教改课题】4份高等学校本科教学改革研究类课题项目申请书申报书
https://bbs.pinggu.org/thread-16437469-1-1.html
【职业学校与专业教育】双高计划申报材料、中期自评报告(院校,专业)超级合集
https://bbs.pinggu.org/thread-16608855-1-1.html
【科研绘图资料】课题项目申请申报书技术路线图与研究框架图合集(适用于各学科)
https://bbs.pinggu.org/thread-16587826-1-1.html
【教师竞赛资料】教学创新大赛参考资料(青教赛、教创赛、教学能力比赛,赠国社科等)
https://bbs.pinggu.org/thread-16586331-1-1.html
【教育教学管理】2025年最新省级市级校级高等教育教学成果奖申报书申报表合集
https://bbs.pinggu.org/thread-16578049-1-1.html
【课题项目】中小学高校教育课题申报立项资料材料模板报告结题材料
https://bbs.pinggu.org/thread-16339763-1-1.html
【课题申请资料】教学教改课题申报书教学能力实施报告模板多色架构图
https://bbs.pinggu.org/thread-16309678-1-1.html
2024年省市级教育教学成果奖合集
https://bbs.pinggu.org/thread-12483283-1-1.html
【科研课题】国家社科基金项目申报综合资料(2022-2005标书,开题,中期,结题等)
https://bbs.pinggu.org/thread-14961969-1-1.html
【课题申请】2020—2005年国家自然科学基金中标标书申请书合集(管理学部)
https://bbs.pinggu.org/thread-15032410-1-1.html
【基金申报】2022—2015年国家社科基金标书合集
https://bbs.pinggu.org/thread-11653799-1-1.html
【基金汇总】(2019—2005)国家社科基金申报资料(申请书、活页、中期检查、开题)等
https://bbs.pinggu.org/thread-11636754-1-1.html
【课题申请】国家社科基金与教育规划项目申请经验与范文合集
https://bbs.pinggu.org/thread-11494635-1-1.html
[学习资料] 【基金课题申报】管理学部、信息学部与数理学部-国家自然科学基金申请书标书合集
https://bbs.pinggu.org/thread-11346860-1-1.html
【基金申请】交叉学科的国家基金与教育部人文社科项目申报书合集
https://bbs.pinggu.org/thread-11564821-1-1.html
【教育教学管理】教学创新大赛超级资源包(青教赛、教创赛、教学能力比赛)
https://bbs.pinggu.org/thread-15126291-1-1.html
【教育教学管理】2025职教宝库(教学能力大赛、教改课题、规划教材、教学成果奖)
https://bbs.pinggu.org/thread-15036935-1-1.html
【课题申请】2020—2005年国家自然科学基金中标标书申请书合集(管理学部)
https://bbs.pinggu.org/thread-15032410-1-1.html
【教育教学管理】本科高校省级立项申报书,课程思政,特等奖,教学成果奖,人才培养
https://bbs.pinggu.org/thread-14978398-1-1.html
【教育教学管理】2018-2024年国家、省教学成果奖超级资料包
https://bbs.pinggu.org/thread-14978393-1-1.html
【教育教学竞赛】教学创新大赛资料(青教赛、教创赛、教学能力比赛)赠送国社科申请
https://bbs.pinggu.org/thread-14978391-1-1.html
【科研课题】国家社科基金项目申报综合资料(2022-2005标书,开题,中期,结题等)
https://bbs.pinggu.org/thread-14961969-1-1.html
【课题申请】历年国家社科基金申请书范例合集
https://bbs.pinggu.org/thread-11562407-1-1.html
https://bbs.pinggu.org/thread-8113337-1-1.html 【重要】国家社科基金申请书与论证活页
https://bbs.pinggu.org/thread-7980590-1-1.html 【基金申请资料】国家社科基金成功范例与经验学习资料
https://bbs.pinggu.org/thread-7855322-1-1.html 【基金申请】国家社科基金项目申请书范例与经验合集
【科研课题】省级教育规划项目申报材料、开题以及中期检查
https://bbs.pinggu.org/thread-11522095-1-1.html
https://bbs.pinggu.org/thread-10868549-1-1.html【重要】国家社科、教育部人文社科基金项目申报范文与经验
https://bbs.pinggu.org/thread-7979953-1-1.html 【重要】国家自然科学基金经管类成功范例
https://bbs.pinggu.org/thread-8019012-1-1.html 【基金申报资料】教育部人文社科基金项目申报书成功范例学习资料
https://bbs.pinggu.org/thread-8006896-1-1.html [学习资料] 教育部人文社科申报案例分析学习资料
https://bbs.pinggu.org/thread-7963726-11-1.html 【重点】教育部人文社科基金项目申请书范文
https://bbs.pinggu.org/thread-8748917-1-1.html [学习资料] 管理类国家自然科学基金申报书撰写学习资料
https://bbs.pinggu.org/thread-7967878-1-1.html [学习资料] 【基金申报资料】经管类国家自然科学基金申报范例学习资料
https://bbs.pinggu.org/thread-7967878-1-1.html 【基金申报资料】经管类国家自然科学基金申报范例学习资料
【课题申请】国家社科基金与教育规划项目申请经验与范文合集
https://bbs.pinggu.org/thread-11494635-1-1.html
【课题申报】国家重点研发计划项目申报书与答辩PPT范文
https://bbs.pinggu.org/thread-12468223-1-1.html