Lecture 1
Concentration Inequalities and Randomized Algorithms
Randomness can be super powerful in designing algorithms. It can help an algorithm run faster and avoid worst-case inputs. Concentration inequalities are a key tool to both analyze and design randomized algorithms. We develop basic concentration tools and study two concrete examples in this section. Morris’s algorithm counts a long stream using a small random counter. Explore-Then-Commit uses samples to decide which action to take in a multi-armed bandit problem. The same concentration tools control the accuracy of the counter and the reliability of the decision.
How to read these notes. Part I contains the course content. Part II develops additional reading on martingales, sub-Gaussian and sub-exponential variables, and UCB. Part III introduces four further topics through their background, main guarantees, and references. These are established results to study and explain; their proofs are not included in Part III.
We assume basic probability, expectation, variance, and independence. All logarithms are natural unless the base is displayed. An assertion about a random variable’s range is understood to hold almost surely.
I. Course Content
1.1 Markov and Chebyshev inequalities
We first recall Markov’s inequality and Chebyshev’s inequality we met in the probability theory course.
Proposition 1 (Markov’s inequality) For any non-negative random variable with finite expectation and for any ,
Proof
Since is non-negative, we have This is equivalent to
Proposition 2 (Chebyshev’s inequality) For any random variable with finite variance and for any , it holds that
Proof
Let , then clearly . Therefore
In order to motivate the study of concentration inequalities, let’s first look at the streaming model.
1.2 The streaming model
Suppose we have a router with limited memory, but need to solve some computational tasks with large input data such as monitoring the IDs of devices visiting it. The following questions are usually asked.
- How many numbers are in a given data stream?
- How many distinct numbers?
- What is the most frequent number?
In order to study these problems systematically, we need to formally define the streaming model. In the streaming model, the input is a sequence where each . We should notice that the data arrive one by one as suggested by the word “streaming” in the name. We now focus on the basic problem: How many numbers are in the stream (what is )?
Clearly we can maintain a counter , and whenever a number arrives, increase by one. To represent every possible count from to , we need bits of memory.
Can we design a more clever algorithm with only memory? It turns out that computing the exact answer is impossible with fewer than bits. The reason is as follows: suppose an algorithm uses fewer bits, so it has fewer than memory states. Denote by the memory state of the algorithm after an input of length . Then there exist distinct such that , and the algorithm cannot output the correct length for both inputs.
Even though we cannot get a better algorithm for the exact answer, it is possible to save a lot of memory if approximation is allowed. That is, for every , the algorithm computes a number such that with high probability.
Morris’s algorithm
Morris’s algorithm is presented as follows.
Morris’s Algorithm
Input: An instance where each .
Output: An estimate of the length of the sequence .
- ;
- On each input: with probability ;
- Return .
This is a randomized algorithm with approximately bits of memory. We first look at the expectation of its output.
Theorem 1 (Morris’s estimator is unbiased) The output of Morris’s algorithm satisfies .
Proof. We prove it by induction on . Since when , we have . Assume it is true for smaller , and let denote the value of after processing the th input. We have Therefore, where the last equation holds due to the induction hypothesis.
It is now clear that Morris’s algorithm is an unbiased estimator for . However, for a practical randomized algorithm, we further require its output to concentrate around the expectation. That is, we want to establish a concentration inequality of the form for . It is natural to see that for fixed , the smaller is, the better the algorithm is.
To use Chebyshev’s inequality to analyze the concentration of Morris’s algorithm, we have to compute the variance of .
Lemma 1 (Second moment of the Morris counter)
Proof
We can prove the claim using an induction argument similar to our proof for the expectation. When , . We assume it is true for smaller and use the same notation . Conditioning on gives Hence
With the above lemma, we can compute the variance as follows: Applying Chebyshev’s inequality, we obtain that for every , However, we observe that as becomes smaller, the above bound is not useful. Thus, it is necessary to improve the concentration of the algorithm. We now introduce two common tricks to achieve this.
The averaging trick
Chebyshev’s inequality tells us that we can improve the concentration by reducing the variance. Let’s first review some properties of variances. Let be a random variable. We have for any constant . For any two independent random variables and , we have
We can design a new algorithm by independently running Morris’s algorithm times in parallel. Denote the corresponding outputs by . The final output is By the above two properties, we have . We can apply Chebyshev’s inequality to and obtain For , we have
The new algorithm uses approximately bits of memory. It shows a trade-off between the accuracy of the randomized algorithm and the consumption of memory space. We can further improve the dependence on using the Chernoff bound below.
1.3 Chernoff bounds and success amplification
Like Chebyshev’s inequality, if we choose for and apply Markov’s inequality to , the bound amounts to bounding , which is the moment generating function of . In case can be well bounded, we obtain sharp concentration bounds.
Theorem 2 (Chernoff bound) Let be independent random variables such that for each . Let and denote . For , we have If , then we have
Proof. We only prove the upper tail bound and the proof of the lower tail bound is similar. For every , we have Therefore, we need to estimate the moment generating function . Since is the sum of independent Bernoulli variables, we have Since , we can compute directly: Therefore, Consequently, Note that the above holds for any . Therefore, we can choose so as to minimize the exponent. Its derivative is , which gives . Substituting this value gives
The following form of the Chernoff bound is more convenient to use (but weaker).
Corollary 1 (Corollary) For any ,
Proof
We only prove the upper tail. It suffices to verify that for , Taking logarithms, this is equivalent to Let . Then For , , and for , . Therefore, first decreases and then increases on . Also and , so on . Hence .
The median trick
We can further boost the performance of Morris’s algorithm using the median trick. We choose in the algorithm introduced in the averaging trick and independently run it times in parallel. Denote the outputs by . It holds that for every , At last, we output the median of .
Then we can apply the Chernoff bound to analyze the result obtained by the median trick. For every , let be the indicator of the good event Then satisfies . If the median is bad, then at least half of the are bad. Equivalently, . By the Chernoff bound, Therefore, for and , we have This new algorithm uses bits of memory.
1.4 Hoeffding’s inequality
One annoying restriction of the Chernoff bound is that each needs to be a Bernoulli random variable. Hoeffding’s inequality generalizes the Chernoff bound by allowing to follow any distribution, provided its value is almost surely bounded.
Theorem 3 (Hoeffding’s inequality) Let be independent random variables where each for certain almost surely. Assume for every . Let and . Then for all .
We learnt from the proof of the Chernoff bound that the key to establishing concentration inequalities of this form is to obtain a nice upper bound on the moment generating function. Therefore, the following Hoeffding lemma will be the main technical ingredient to prove the inequality.
Lemma 2 (Hoeffding’s lemma) Let be a random variable with and . Then
Proof. Let . We first compute its derivatives: and When , using the condition , we have . Let be the distribution of . We can interpret as the variance of a tilted random variable: , where Since is supported on , for every , Therefore, by Taylor’s formula, and hence .
Armed with Hoeffding’s lemma, it is routine to prove Hoeffding’s inequality.
Proof of Hoeffding’s inequality
First note that we can assume and therefore ; if not, replace by . Put . By symmetry, we only need to prove the upper tail. Since and applying Hoeffding’s lemma for each factor yields Let . We have
1.5 Multi-armed bandits: Explore-Then-Commit
Suppose there is a -arm bandit, and the reward of each arm follows some distribution supported on with mean ; rewards are independent across arms and rounds. We assume without loss of generality that . Now suppose you can pull the bandit for rounds and the goal is to obtain maximum reward in expectation. If we know , the optimal strategy is to pull arm for times, and the expected reward is . However, in case we do not know the distributions, we have to design some strategy to explore the bandit first.
Denote by the arm pulled in round , and thus the reward in the th round satisfies . The regret of a strategy is defined as the gap between and the expected rewards of the strategy in rounds, namely the regret of not always choosing the first arm:
For every , denote as the gap between the reward of the th arm and the optimal arm. The naive strategy that pulls each arm equally often is bad. When divides , its regret is which is linear in . We consider a strategy/algorithm good if , or equivalently .
Proposition 3 For every , let denote the number of times arm is pulled in the first rounds. Then
Proof
By conditioning on the chosen arm in each round, Since , this equals the stated sum.
We also write for every , and then .
The algorithm and the trade-off
To get small regret, our strategy should identify the best arm as soon as possible. The most straightforward way to find the best arm is to try each arm a few times and pick the one with the best empirical reward. The Explore-Then-Commit algorithm implements this idea: pull every arm for times (so times in total for exploration), and calculate (the average reward gained in those times). After this, always pull an arm with the greatest . We assume and use any fixed rule to break ties. We can write its regret as
Moreover, When ,
We bound the above probability by concentration inequalities. To this end, let be the th reward from , and let be the th reward from . Let ; then . Let ; then . By Hoeffding’s inequality, Therefore,
A bound without knowing the gaps
To further upper bound , define We would like to determine minimizing the upper bound of among all possible , i.e., . First we calculate : We have when , and when . Thus, for all ,
Finally, assuming and setting , we have
The Explore-Then-Commit algorithm enjoys sublinear regret for fixed , which is good, but still suboptimal. The main disadvantage is that it treats all arms equally in the exploration step and pulls each of them for a fixed times regardless of the rewards already obtained.
II. Reading Contents
The classroom arguments leave two natural questions. Can we obtain concentration when observations are dependent, or when variables are not bounded? And can we use concentration to decide how many samples to collect? The following sections develop these questions without changing the basic exponential-moment strategy.
2.1 Concentration with martingales
Our concentration proofs for sums have used mutual independence to factor an MGF. When new observations depend on the past, factorization is no longer available. Conditional expectation gives a replacement: we bound the next factor after fixing everything observed so far, and then work backward.
A fair game and its information
Imagine a fair game in which the size of the next bet may depend on previous wins and losses. The increments need not be independent. What matters is that their conditional expected value is zero.
Definition 1 (Filtrations and martingales) A filtration is an increasing sequence of -algebras representing the information available over time. An integrable process is a martingale if is -measurable and, for every , Equivalently, the increments have conditional mean zero.
For example, if are independent and integrable, then is a martingale with . We have also already met a dependent example: for Morris’s counter, is a martingale, by the conditional expectation calculation in Part I. Its increments are not uniformly small, so it does not directly give a useful bounded-increment estimate.
Revealing a function one coordinate at a time
Let be integrable. Reveal the inputs successively and keep track of the current prediction of : Here and . The tower property gives where . This is the Doob martingale of . The construction itself does not require independent inputs.
This converts a deviation of into a sum of changes in conditional predictions. To prove concentration, we need to control how much one new observation can change that prediction.
Theorem 4 (Azuma–Hoeffding with conditional range bounds) Let be a martingale. Suppose there are -measurable endpoints and deterministic constants such that If , then, for every , Each one-sided bound omits the factor . In particular, the assumption gives .
Proof. Write . Conditional on , the variable has mean zero and lies in an interval of length at most . Hoeffding’s lemma, applied to this conditional distribution, gives Because are already known at time , Iterating yields . The same exponential Markov argument as before, with , proves the upper tail; use for the lower tail.
Notice the distinction between an absolute increment bound and a conditional range length. A variable in has range length , not . Tracking the conditional range is what gives the sharp constant in the next application.
Bounded differences: McDiarmid’s inequality
Suppose are independent. A function has coordinate bounded differences if changing just coordinate changes its value by at most . Writing for with coordinate replaced by , this is The condition is imposed on the product of the input spaces. It describes sensitivity to one input, rather than a Euclidean Lipschitz constant.
Theorem 5 (McDiarmid’s inequality) For independent inputs and an integrable function with coordinate bounded differences , put . Then, for ,
Proof. Use the Doob martingale . Write and . Fix the past and define Independence means that the future inputs in this expectation have the same distribution for every value of . Couple the two expectations using the same future inputs. The bounded-difference assumption then gives Conditional on the past, and is the conditional average of . Subtracting that average does not change the length of its range. Hence the conditional range of has length at most , and Theorem 4 applies.
Example 1 (Empty bins) Throw balls independently and uniformly into bins. Let be the number of empty bins. The empty-bin indicators are dependent, but linearity of expectation still gives Use the positions of the balls as independent inputs. Moving one ball can empty its old bin or fill its new bin. If both occur, the changes cancel; in all cases the total number of empty bins changes by at most . McDiarmid therefore gives The denominator counts independent input coordinates—balls, not bins.
Example 2 (Sampling without replacement) A bag contains balls, of which are red. Draw balls without replacement, let indicate that draw is red, and put . We want concentration for , whose mean is .
The Doob prediction after observations is Let be the conditional probability of red at draw . Direct subtraction gives Given the past, this increment lies in an interval of length . Hence If , the total number of red balls drawn is exactly and no concentration estimate is needed.
Example 3 (Choosing what to reveal: the chromatic number) Let and let be its chromatic number. Revealing all independent edge indicators gives a bounded-difference constant for each edge, but only yields a deviation scale of order .
There is a better representation. For , let encode the edges from vertex to vertices . These inputs are independent because they contain disjoint sets of edge indicators. Changing changes only edges incident to vertex .
Two such graphs have the same graph after deleting vertex . Each chromatic number lies between the chromatic number of that common graph and one more. Their difference is therefore at most . McDiarmid now gives The gain comes from choosing a more informative unit of exposure.
2.2 Sub-Gaussian random variables
Recall the calculation underlying Hoeffding. For a centered variable, a quadratic upper bound on its log-MGF led to a Gaussian-shaped tail. Boundedness was one way to obtain that upper bound; it need not be the only way.
For an integrable , write allowing the value if the MGF diverges.
Definition 2 (Sub-Gaussian variables) We call -sub-Gaussian if, for every , The number is a variance proxy. It is an upper bound on , but need not equal it.
Exponential Markov, optimized at , immediately gives, when , For proxy zero, is constant almost surely.
Example 4 (Gaussian, bounded, and Rademacher variables) If , completing the square in the Gaussian integral gives exactly.
If , Hoeffding’s lemma gives variance proxy .
In particular, a uniform random sign has proxy . One can also see this from .
Addition and the role of independence
If are independent with proxies , then for deterministic real weights , Thus independent variances add at the level of proxies. Hoeffding’s inequality is a special case.
Without independence, this sum-of-squares rule can fail. For example, gives , whose variance is . A useful general comparison follows from Cauchy–Schwarz: if have proxies , with arbitrary dependence, then This gives proxy rather than .
From one variable to a finite family
Suppose are centered and each has proxy at most . A union bound gives There is no independence assumption between the here. For example, distance errors in a random projection can be highly dependent, but a pointwise tail bound still combines with a union bound over a finite set of pairs.
The expected maximum admits a related bound. For , Optimizing gives .
Tail and moment descriptions of sub-Gaussian variables are equivalent to the MGF description up to universal changes in constants. For example, the displayed tail bound implies, for , which gives for a universal . The reverse implications and precise parameter comparisons can be found in Vershynin’s High-Dimensional Probability, Chapter 2. We use the MGF convention throughout these notes.
2.3 Sub-exponential variables and Bernstein bounds
A useful variable may fail to have an MGF for all real parameters. The square of a Gaussian is the simplest example. If and , direct integration gives So is not sub-Gaussian. Nevertheless, its log-MGF is quadratic near zero. We can still run the Chernoff argument, with a restricted choice of parameter.
Definition 3 (Sub-exponential variables: a two-parameter convention) For and , call -sub-exponential if Here is a variance-scale parameter and limits the MGF’s usable range. Some references write in place of our first parameter; the convention matters when adding variables.
For , expanding the logarithm and bounding absolute values gives, when , Thus is -sub-exponential. The absolute values ensure that the calculation also applies to negative .
Proposition 4 (Two tail regimes) If is -sub-exponential, then, for every , The quadratic regime ends at .
Proof
For , exponential Markov gives for the upper tail. If , choose and obtain . If , choose and use Apply the same argument to and add the bounds.
For independent -sub-exponential variables, the weighted sum has parameters Indeed, puts every in its admissible interval, and the log-MGFs add. Zero weights can be omitted; if all weights vanish, the sum is deterministic.
Example 5 (The length of a Gaussian vector) For independent standard Gaussians , the squared norm has mean and sub-exponential parameters . Therefore, For , On the complementary event, . This is a useful precursor to the Gaussian random projections used in the JL lemma.
Using variance information for bounded variables
Hoeffding uses the range of each variable but not its variance. If an event is rare, that can lose useful information, as our Bernoulli example showed. A more careful MGF expansion retains the variance.
Theorem 6 (Bernstein’s inequality) Let be independent, with and for . Set . Then, for , If , every vanishes almost surely.
Proof
Write . For every integer , . Also . For , the power series therefore gives Using and independence yields For the upper tail take , which is strictly less than . Substitution into exponential Markov gives the claimed exponent. Repeat for the negative sum.
For example, each centered bounded above is also -sub-exponential when , by restricting to . The corresponding sum has the two-regime transition at . The smoother Bernstein bound retains more of the MGF estimate.
Example 6 (All degrees in a random graph) In , the degree of a fixed vertex has mean and variance . Put and Since , Bernstein gives . A union bound over vertices gives The degrees are dependent, but the union bound does not require their independence. The linear term is essential in sparse regimes; it cannot be dropped without an additional assumption comparing with .
2.4 Upper Confidence Bounds
ETC explores every arm for the same number of rounds. Ideally, an arm with a large gap should receive fewer observations than an arm whose mean is close to the optimum. But the gaps are unknown. Can observations themselves tell us when an arm has been explored enough?
UCB answers this using optimism. Keep a plausible upper bound on each mean, and choose the arm with the largest upper bound. An arm is attractive either because its observed rewards are high, or because it has been sampled too little to rule out a high mean.
Confidence intervals under adaptive sampling
Use the bounded stochastic bandit model from Part I, with . For each arm , imagine an independent infinite reward stream sampled in advance. The th time arm is pulled, reveal . Define the fixed-prefix means For , let For each fixed with , Hoeffding bounds the probability of by . Taking a union bound gives where This event holds simultaneously for every possible sample count. We may therefore evaluate it at the random count chosen by the policy. Applying a fixed-sample Hoeffding bound directly at an adaptive random sample size would skip this justification.
The algorithm
Pull each arm once. After rounds, choose at round an arm maximizing Break ties by any fixed rule. The policy here is a horizon-aware version of UCB; use , so , for the regret guarantees below.
How many times can a suboptimal arm be selected?
On , an optimal arm always has index at least . If arm is selected after initialization, its index is at least that of an optimal arm. At that moment, Thus, for , This is the number of samples before the next pull. Consequently, No gap was given to the algorithm; the analysis uses the true gap to explain when the confidence interval becomes narrow enough.
Theorem 7 (Regret of the horizon-aware UCB policy) In the bounded stochastic bandit model with , set . The policy above satisfies It also satisfies, for every , and therefore .
Proof. Let be the random sum of gaps incurred. Its expectation is , and . On , the pull-count bound gives the first displayed sum. The complement has probability at most , so its contribution to is at most .
For the second bound, separate arms with from those with . The small-gap arms contribute at most on every path, regardless of their pull counts. On , the other arms contribute at most Add the failure-event contribution and choose . This gives ; since and , it has the claimed order. The trivial bound remains available.
This analysis has two levels of resolution. For a fixed instance, it bounds pulls in terms of each gap. For the worst case, it avoids trying to distinguish arms whose gaps are so small that pulling them has little cost. MOSS, one of the further topics below, changes the confidence schedule to reach the optimal worst-case scale without the extra logarithmic factor.
III. Further Topics
The following four topics continue the two algorithmic examples of this lecture. The first two concern how to allocate samples among alternatives; the other two concern what a small summary of a stream can preserve. Each entry states the problem and a result to understand, followed by sources. Algorithms and proofs are left to the readings.
3.1 MOSS: minimax-optimal stochastic bandits
Background. The UCB analysis above gives worst-case expected regret . Is the logarithmic factor a necessary price for learning unknown means? MOSS—Minimax Optimal Strategy in the Stochastic case—shows that it is not. It belongs to the UCB family but uses a different exploration schedule.
Main result. For , , and stochastic rewards in , a horizon-aware MOSS policy achieves for a universal constant . A matching worst-case lower bound of order holds for every policy. Thus “optimal” here means minimax-optimal expected cumulative regret, up to constants. It does not assert that one policy is best on every individual instance.
References. Jean-Yves Audibert and Sébastien Bubeck, Minimax Policies for Adversarial and Stochastic Bandits, COLT 2009, stochastic-bandit results. For a textbook treatment, see Tor Lattimore and Csaba Szepesvári, Bandit Algorithms, Chapter 9; the minimax lower-bound background is in Part IV.
3.2 Median Elimination: finding a near-best arm
Background. In ETC, exploration is used to choose an arm. We can study this identification task on its own: there is no reward objective during sampling, and we only want to return a good arm with high confidence. Uniformly estimating all means to accuracy and taking the empirical best uses samples. Must we pay that logarithmic dependence on ?
Main result. For independent stochastic arms with rewards in , Median Elimination returns an arm satisfying using samples, for . The guarantee asks for an -optimal arm, not exact identification of the unique best arm. Its performance measure is sample complexity, rather than cumulative regret.
Reference. Eyal Even-Dar, Shie Mannor, and Yishay Mansour, Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems, JMLR 7 (2006), 1079–1105. Read §3, especially the Median Elimination algorithm and Theorem 10. The reinforcement-learning sections are not needed for this topic.
3.3 Count-Min Sketch: estimating individual frequencies
Background. Morris counts the total number of stream items. Suppose instead that items belong to a universe , and after processing the stream we want to query how often a particular item appeared. An exact table of all frequencies may be too large. Let be the frequency of item and .
Main result. For insertion-only streams, Count-Min uses counters and returns, for a fixed item , an estimate satisfying with probability at least . This is an additive error relative to the total stream mass, not a relative-error guarantee for every individual frequency. The probability is over the sketch’s random choices for a fixed stream and fixed query; a simultaneous guarantee for all items requires an adjusted failure probability.
Reference. Graham Cormode and S. Muthukrishnan, An Improved Data Stream Summary: The Count-Min Sketch and its Applications, Journal of Algorithms 55(1) (2005), 58–75. Start with the sketch definition and the point-query guarantee; range queries and other applications are optional.
3.4 AMS sketch: estimating the second frequency moment
Background. Two streams can have the same length but very different repetition patterns. A useful summary of repetition is the second frequency moment It equals the number of ordered pairs of stream positions carrying the same item, including pairs consisting of one position twice. Unlike the stream length, it depends on how the mass is distributed among items.
Main result. For a fixed insertion-only stream, an AMS second-moment sketch produces an estimate with A standard amplified version uses independent sketch counters. Writing , a bit-space bound including bounded-independence seeds is for and . This connects the moment calculations and success amplification used for Morris to a different streaming statistic.
Reference. Noga Alon, Yossi Matias, and Mario Szegedy, The Space Complexity of Approximating the Frequency Moments, STOC 1996; journal version in Journal of Computer and System Sciences 58(1) (1999), 137–147. Focus on the estimator and its moment analysis; the general frequency-moment lower bounds are outside this topic.
Sources and acknowledgments. The main exposition adapts Chihao Zhang’s 2025 Topics in Advanced Algorithms Lecture 1 and Lecture 2, originally scribed by Yuchen He. The martingale reading adapts Lecture 3, originally scribed by Fangke Li and Yuchen He. The 2026 version reorganizes this material and supplies corrections and additional details.
Additional background: Roman Vershynin, High-Dimensional Probability, Chapter 2, for sub-Gaussian and sub-exponential variables; Lattimore and Szepesvári, Bandit Algorithms, Chapters 6–9, for ETC, UCB, and MOSS. Original approximate-counting source: Robert Morris, Counting Large Numbers of Events in Small Registers, Communications of the ACM 21(10) (1978), 840–842.