COURSE DETAIL
This course explores the foundations of data science from three perspectives: inferential thinking, computational thinking, and real-world relevance. It focuses on critical concepts and skills in computer programming and statistical inference, in conjunction with hands-on analysis of real-world datasets, including economic data, document collections, geographical data, and social networks. This course also delves into social and legal issues surrounding data analysis, including issues of privacy and data ownership.
The curriculum and format are designed specifically for students who have not previously taken statistics or computer science courses. Students with some prior experience in either statistics or computing are welcome to enroll and often find that this course offers a new perspective that blends computational and inferential thinking. Students who have taken several statistics or computer science courses should instead take a more advanced course.
COURSE DETAIL
This course deals with the application of statistical methods to test hypotheses and draw inferences from data, using maximum likelihood methods. The course starts by developing general-purpose maximum likelihood methods, with interval estimation by means of the information matrix and the bootstrap. It goes on to develop generalized linear models, linear models and analysis of variance models as special cases of maximum likelihood methods. It covers diagnostic methods, including methods for selecting between models, checking assumptions and testing goodness-of-fit. It has an applied focus, with extensive use of R to give students practice in doing inference with real datasets, from problem formulation through to final conclusions.
COURSE DETAIL
Machine learning is concerned with algorithms that process relevant data and then perform some task. Often, performance of machine learning algorithms is measured statistically, and the algorithms themselves are heavily influenced by statistical ideas. For example, after observing several (x,y) pairs an algorithm may be able to predict with high accuracy the corresponding value of y for an unseen x. When the data is complex and/or high-dimensional, a number of statistical and algorithmic issues arise: a sufficiently rich class of statistical models must be used effectively and irrelevant data should be identified and then discarded. Students understand the statistical approach to analyzing data, and how it can be used to effectively perform tasks under appropriate assumptions. This enables students to formulate various real-life problems as statistical learning tasks and use common techniques to develop solutions.
COURSE DETAIL
This course provides a continuation of the study of medical statistics, with emphasis on more advanced topics in epidemiological methods and the design and analysis of clinical trials. Students learn how to model survival data using parametric regression models; to develop and validate a risk prediction model; to analyze clustered data using a regression model; to design and analyze a cross-over trial, cluster randomized trial, equivalence trial, and early phase trial; to understand the issues concerning interim analyses and missing data; and to carry out a meta-analysis.
COURSE DETAIL
This is a basic course in designing experiments and analyzing the resulting data. It is intended for engineers, physical/chemical scientists, and scientists from other fields such as biotechnology and biology. The course deals with the types of experiments that are frequently conducted in industrial settings. Its objective is to learn how to plan, design, and conduct experiments efficiently and effectively, and analyze the resulting data to obtain objective conclusions. Both design and statistical analysis issues are discussed. Opportunities to use the principles taught in the course arise in all phases of engineering and scientific work, including technology development, new product design and development, process development, and manufacturing process improvement. Applications from various fields of engineering (including chemical, mechanical, electrical, materials science, industrial, etc.) will be illustrated throughout the course. Topics include simple design with fixed and random effects. Simultaneous confidence intervals. Requirements for analysis of variance: transformations, model validation, residual analysis. Factorial design with fixed, random, and mixed effects. Additivity and interaction. Complete and incomplete designs. Randomized block designs, Latin squares and confounding. Regression and analysis of covariance. Admission requirements include FMAA20 Linear Algebra with Introduction to Computer Tools or FMAA21 Linear Algebra with Numerical Applications or FMAB20 Linear Algebra or FMAB22 Linear Algebra and FMAB30 Calculus in Several Variables or FMAB35 Calculus in Several Variables or FMSF20 Mathematical Statistics, Basic Course or FMSF25 Mathematical Statistics - Complementary Project or FMSF32 Mathematical Statistics or FMSF45 Mathematical Statistics, Basic Course or FMSF50 Mathematical Statistics, Basic Course or FMSF55 Mathematical Statistics, Basic Course or FMSF70 Mathematical Statistics or FMSF75 Mathematical Statistics, Basic Course or FMSF80 Mathematical Statistics, Basic Course. Assumed prior knowledge: Basic mathematical statistics and programming experience.
COURSE DETAIL
A wide range of phenomena from areas as diverse as physics, economics, and biology can be described by simple probabilistic models. Often, phenomena from different areas share a common mathematical structure. In this course a variety of mathematical structures of wide applicability is described and analyzed. The emphasis is on developing the tools which are useful to anyone modelling applications, rather than the applications themselves Students should have a good knowledge of first year probability and of basic material from first year analysis. As the course builds on Probability 1, it also deepens students' understanding of the basis of probability theory.
COURSE DETAIL
What puts former criminals on the right track? How can we prevent heart disease? Can Twitter predict election outcomes? What does a violent brain look like? How many social classes does 21st century society have? Are hospitals spending too much on health care, or too little? Data analysis is the art and science of tackling questions like these by looking at data. Just as cartographers make maps to see what a country looks like, data analysts explore the hidden structures of data by creating informative pictures and summarizing relationships among variables. And just as doctors diagnose sick patients and advise healthy ones on how to stay healthy, data analysts predict important events and variables so we can act on this knowledge. Methods from statistics, machine learning, and data mining play an important part in this process, as well as visualizations that allow the analyst and other humans to better understand what we can conclude from the available facts. During this course, students actively learn how to apply the main statistical methods in data analysis and how to use machine learning algorithms and visualizing techniques. The course goes beyond linear and logistic regression and thus continue where “Fundamental techniques in data science with R” ended. The course has a strongly practical, hands-on focus: rather than focusing on the mathematics and background of the discussed techniques, student gain hands on experience in using them on real data during the course and interpreting the results. Entry requirements include at least followed an introductory statistics course of 7.5 EC, and familiarity with correlation and regression, comparing means and cross tabulations of categorical variables. It's also expected to have hands on experience in carrying out these analyses, with, for example, SPSS, Stata, R or SAS.
COURSE DETAIL
This course covers: introduction to actuarial modelling; the application of compound interest techniques to financial transactions; generalized cash-models to describe financial transactions such as zero-coupon bonds, fixed interest securities, cash on deposit, equities, interest only loans, repayment loans, annuities certain and others; introduction to R programming for Actuarial Science, and introduction to life insurance.
COURSE DETAIL
This course focuses on data about connections, forming structures known as networks. Networks and network data describe an increasingly vast part of the modern world, through connections on social media, communications, financial transactions, and other ties. This course covers the fundamentals of network structures, network data structures, and the analysis and presentation of network data. Students work directly with network data, and structure and analyze these data using the R statistical programming language. This course develops the theory and methodological tools needed to model and predict social networks and use them in social sciences as diverse as sociology, political science, economics, health, psychology, history, or business. The core of the course comprises the essential tools of network analysis, from centrality, homophily, and community detection, to random graphs, network formation, and information flow.
Pagination
- Page 1
- Next page