| 1. 9/9 |
Wed |
Course overview. Probability review. Linearity of expectation and variance. |
Slides. Compressed slides. Reading: MIT short videos and exercises on probability (go to Unit 4). Khan academy probability lessons (a bit more basic). Chapters 1-3 of Probability and Computing with content and exercises on basic probability, expectation, variance, and concentration bounds. |
| Randomized Methods, Sketching & Streaming |
| 2. 9/14 |
Mon |
Application of linearity of expectation to analyzing random hashing for efficient lookup. Markov's inequality. |
Slides. Compressed slides. Reading: Chapters 1-3 of Probability and Computing with content and exercises on basic probability, expectation, variance, and concentration bounds. |
| 3. 9/16 |
Wed |
Application of Markov's inequality to collision-free hashing and two-level hashing. 2-universal and pairwise independent hashing. Hashing for load balancing and Chebyshev's inequality. The law of large numbers. |
Slides. Compressed slides. Reading: Chapter 2.2 of Foundations of Data Science with content on Markov's inequality and Chebyshev's inequality. Exercises 2.1-2.6. Chapters 1-3 of Probability and Computing with content and exercises on basic probability, expectation, variance, and concentration bounds. Some notes (Arora and Kothari at Princeton) proving that the ax+b mod p hash function described in class in 2-universal.
|
| 4. 9/21 |
Mon |
Union bound. Exponential concentration bounds and the central limit theorem. Application to random hashing. |
Reading: Chapter 4 of Probability and Computing on exponential concentration bounds. Some notes (Goemans at MIT) showing how to prove exponential tail bounds using the moment generating function + Markov's inequality approach. |
| 5. 9/23 |
Wed |
Bloom Filters. |
Reading: Chapter 4 of Mining of Massive Datasets, with content on Bloom filters. See here for full Bloom filter analysis. See Wikipedia for a discussion of the many bloom filter variants, including counting Bloom filters, and Bloom filters with deletions. |
| 6. 9/28 |
Mon |
Streaming algorithms and frequent elements estimation via Count-min sketch. |
Reading: Notes (Amit Chakrabarti at Dartmouth) on streaming algorithms. See Chapters 1 and 5 for frequent elements. Some more notes on the frequent elements problem. A website with lots of resources, implementations, and example applications of count-min sketch. |
| 7. 9/30 |
Wed |
Min-Hashing for Distinct elements. The median trick. Distinct elements in pratice: Flajolet-Martin and HyperLogLog. |
Reading: Chapter 4 of Mining of Massive Datasets, with content on distinct elements counting. The 2007 paper introducing the popular HyperLogLog distinct elements algorithm. |
| 8. 10/5 |
Mon |
Approximate nearest neighbor search, vector databases, and locality sensitive hashing. |
Reading: Chapter 3 of Mining of Massive Datasets, with content on Jaccard similarity, MinHash, and locality sensitive hashing.
|
| 9. 10/7 |
Wed |
Finish up locality sensity hashing -- SimHash for cosine similarity. Graph-based approximate nearest neighbor search. |
Reading: Chapter 3 of Mining of Massive Datasets, with content on Jaccard similarity, MinHash, and locality sensitive hashing.
|
| 10/12 |
Mon |
No Class. Indigenous People’s Day |
|
| 10. 10/14 |
Wed |
Finish up graph-based search. Midterm 1 Review. |
Reading:
|
| 10/19 |
Mon |
Midterm 1. 2:30-3:45pm. In class. |
|
| 11. 10/21 |
Wed |
Compressing high dimensional data: low-distortion embeddings and the Johnson-Lindenstrauss Lemma.
|
Reading: Chapter 2.7 of Foundations of Data Science on the Johnson-Lindenstrauss lemma. Notes on the JL-Lemma (Anupam Gupta (CMU). Sparse random projections which can be multiplied by more quickly. Some good videos for linear algebra review.. See also: Khan academy. |
| 12. 10/26 |
Mon |
Finish up JL Lemma. Vector quantization via random projection. |
Reading:
|
| Spectral Methods |
| 13. 10/28 |
Wed |
Intro to principal component analysis, low-rank approximation, data-dependent dimensionality reduction. Orthogonal bases and projection matrices. |
Reading: Chapter 3 of Foundations of Data Science and Chapter 11 of Mining of Massive Datasets on low-rank approximation and the SVD. Some good videos overviewing the SVD and related topics (like orthogonal projection and low-rank approximation). |
| 14. 11/2 |
Mon |
Finish up low-rank approximation motivation. Dual column/row view of low-rank approximation. Best fit subspaces and optimal low-rank approximation via eigendecomposition. Eigenvalues as a measure of low-rank approximation error.
|
Reading: Chapter 3 of Foundations of Data Science and Chapter 11 of Mining of Massive Datasets on low-rank approximation. Notes on SVD and its connection to eigendecomposition/PCA (Roughgarden and Valiant at Stanford). Proof that optimal low-rank approximation can be found greedily (see Section 1.1). |
| 15. 11/4 |
Wed |
Singular value decomposition and its connection to optimal low-rank approximation. Applications of SVD beyond low-rank approximation. Applications of low-rank approximation beyond compression. Matrix completion and entity embeddings.
|
Reading: Levy Goldberg paper on word embeddings as implicit low-rank approximation. |
| 16. 11/9 |
Mon |
Spectral graph theory and spectral clustering. |
Reading: Chapter 10.4 of Mining of Massive Datasets on spectral graph partitioning. For a lot more interesting material on spectral graph methods see Dan Spielman's lecture notes. Great notes on spectral graph methods (Roughgarden and Valiant at Stanford). |
| 11/11 |
Wed |
No Class. Veterans' Day |
|
| 17. 11/16 |
Mon |
The stochastic block model. |
Reading: Dan Spielman's lecture notes on stochastic block model, including matrix concentration + David-Kahan perturbation analysis.. Further stochastic block model notes (Alessandro Rinaldo at CMU). A survey of the vast literature on the stochastic block model, beyond the spectral methods discussed in class (Emmanuel Abbe at Princeton). |
| 11/18 |
Wed |
Midterm 2. 2:30-3:45pm. In class. |
|
| 18. 11/23 |
Mon |
Computing the SVD: power method. |
Reading: Chapter 3.7 of Foundations of Data Science on the power method for SVD. Some notes on the power method. (Roughgarden and Valiant at Stanford). |
| 11/24 |
Tue |
No Class. Wednesday class schedule followed, but we will have no class. |
|
| 11/25 |
Wed |
No Class. Thanksgiving Recess |
|
| Optimization |
| 19. 11/30 |
Mon |
Intro to gradient descent and assumptions for analysis. |
Reading: Chapters I and III of these notes (Hardt at Berkeley). |
| 20. 12/2 |
Wed |
Analysis of gradient descent for convex Lipschitz functions. |
Reading: Chapters I and III of these notes (Hardt at Berkeley). |
| 21. 12/7 |
Mon |
Constrained optimization and projected gradient descent. Online gradient descent set up and regret definition. High level discussion of stochastic gradient descent. |
Reading: Short notes, proving regret bound for online gradient descent. A good book (by Elad Hazan) on online optimization, including online gradient descent and connection to stochastic gradient descent. |
| 22. 12/9 |
Wed |
TBD |
|
| 23. 12/14 |
Mon |
Course wrap-up and final exam review. |
|
| 12/17 |
Thu |
Final Exam. 3:30-5:30pm. In regular classroom. |
|