Topics in Advanced AlgorithmsFall 2026

Lecture 1

Concentration Inequalities and Randomized Algorithms

Randomness can be super powerful in designing algorithms. It can help an algorithm run faster and avoid worst-case inputs. Concentration inequalities are a key tool to both analyze and design randomized algorithms. We develop basic concentration tools and study two concrete examples in this section. Morris’s algorithm counts a long stream using a small random counter. Explore-Then-Commit uses samples to decide which action to take in a multi-armed bandit problem. The same concentration tools control the accuracy of the counter and the reliability of the decision.

How to read these notes. Part I contains the course content. Part II develops additional reading on martingales, sub-Gaussian and sub-exponential variables, and UCB. Part III introduces four further topics through their background, main guarantees, and references. These are established results to study and explain; their proofs are not included in Part III.

We assume basic probability, expectation, variance, and independence. All logarithms are natural unless the base is displayed. An assertion about a random variable’s range is understood to hold almost surely.

I. Course Content

1.1 Markov and Chebyshev inequalities

We first recall Markov’s inequality and Chebyshev’s inequality we met in the probability theory course.

Proposition 1 (Markov’s inequality) For any non-negative random variable XX with finite expectation and for any a>0a>0, P(X≥a)≤EXa. \mathbb P(X\ge a)\le \frac{\mathbb E X}{a}.

Proof

Since XX is non-negative, we have EX≥a⋅P(X≥a)+0⋅P(X<a). \mathbb E X\ge a\cdot\mathbb P(X\ge a)+0\cdot\mathbb P(X<a). This is equivalent to P(X≥a)≤EXa. \mathbb P(X\ge a)\le \frac{\mathbb E X}{a}.

Proposition 2 (Chebyshev’s inequality) For any random variable XX with finite variance and for any a>0a>0, it holds that P(∣X−EX∣≥a)≤Var⁡(X)a2. \mathbb P(|X-\mathbb E X|\ge a) \le \frac{\operatorname{Var}(X)}{a^2}.

Proof

Let Y=∣X−EX∣Y=|X-\mathbb E X|, then clearly Y≥0Y\ge0. Therefore P(∣X−EX∣≥a)=P(Y≥a)=P(Y2≥a2)≤E[Y2]a2=E[(X−EX)2]a2=Var⁡(X)a2. \begin{aligned} \mathbb P(|X-\mathbb E X|\ge a) &=\mathbb P(Y\ge a)=\mathbb P(Y^2\ge a^2)\\ &\le\frac{\mathbb E[Y^2]}{a^2} =\frac{\mathbb E[(X-\mathbb E X)^2]}{a^2} =\frac{\operatorname{Var}(X)}{a^2}. \end{aligned}

In order to motivate the study of concentration inequalities, let’s first look at the streaming model.

1.2 The streaming model

Suppose we have a router with limited memory, but need to solve some computational tasks with large input data such as monitoring the IDs of devices visiting it. The following questions are usually asked.

In order to study these problems systematically, we need to formally define the streaming model. In the streaming model, the input is a sequence σ=⟨a1,a2,…,am⟩\sigma=\langle a_1,a_2,\ldots,a_m\rangle where each ai∈[n]a_i\in[n]. We should notice that the data arrive one by one as suggested by the word “streaming” in the name. We now focus on the basic problem: How many numbers are in the stream (what is mm)?

Clearly we can maintain a counter kk, and whenever a number aia_i arrives, increase kk by one. To represent every possible count from 00 to mm, we need ⌈log⁡2(m+1)⌉\lceil\log_2(m+1)\rceil bits of memory.

Can we design a more clever algorithm with only o(log⁡m)o(\log m) memory? It turns out that computing the exact answer is impossible with fewer than ⌈log⁡2(m+1)⌉\lceil\log_2(m+1)\rceil bits. The reason is as follows: suppose an algorithm uses fewer bits, so it has fewer than m+1m+1 memory states. Denote by M(i)\mathcal M(i) the memory state of the algorithm after an input σ\sigma of length ii. Then there exist distinct i,j∈{0,1,…,m}i,j\in\{0,1,\ldots,m\} such that M(i)=M(j)\mathcal M(i)=\mathcal M(j), and the algorithm cannot output the correct length for both inputs.

Even though we cannot get a better algorithm for the exact answer, it is possible to save a lot of memory if approximation is allowed. That is, for every ε>0\varepsilon>0, the algorithm computes a number m^\widehat m such that 1−ε≤m^m≤1+ε 1-\varepsilon\le\frac{\widehat m}{m}\le1+\varepsilon with high probability.

Morris’s algorithm

Morris’s algorithm is presented as follows.

Morris’s Algorithm

Input: An instance σ=⟨a1,a2,…,am⟩\sigma=\langle a_1,a_2,\ldots,a_m\rangle where each ai∈[n]a_i\in[n].

Output: An estimate of the length mm of the sequence σ\sigma.

  1. X←0X\leftarrow0;
  2. On each input: X←X+1X\leftarrow X+1 with probability 2−X2^{-X};
  3. Return 2X−12^X-1.

This is a randomized algorithm with approximately O(log⁡log⁡m)O(\log\log m) bits of memory. We first look at the expectation of its output.

Theorem 1 (Morris’s estimator is unbiased) The output of Morris’s algorithm m^\widehat m satisfies Em^=m\mathbb E\widehat m=m.

Proof. We prove it by induction on mm. Since X=1X=1 when m=1m=1, we have Em^=1\mathbb E\widehat m=1. Assume it is true for smaller mm, and let XiX_i denote the value of XX after processing the iith input. We have E[2Xm∣Xm−1=i]=(1−2−i)2i+2−i2i+1=2i+1. \mathbb E[2^{X_m}\mid X_{m-1}=i] =(1-2^{-i})2^i+2^{-i}2^{i+1}=2^i+1. Therefore, Em^=E[2Xm]−1=E[2Xm−1+1]−1=E[2Xm−1]=m, \begin{aligned} \mathbb E\widehat m &=\mathbb E[2^{X_m}]-1 =\mathbb E[2^{X_{m-1}}+1]-1\\ &=\mathbb E[2^{X_{m-1}}]=m, \end{aligned} where the last equation holds due to the induction hypothesis.

It is now clear that Morris’s algorithm is an unbiased estimator for mm. However, for a practical randomized algorithm, we further require its output to concentrate around the expectation. That is, we want to establish a concentration inequality of the form P(∣m^−m∣≥εm)≤δ \mathbb P(|\widehat m-m|\ge\varepsilon m) \le\delta for ε,delta>0\varepsilon,delta>0. It is natural to see that for fixed ε\varepsilon, the smaller δ\delta is, the better the algorithm is.

To use Chebyshev’s inequality to analyze the concentration of Morris’s algorithm, we have to compute the variance of m^\widehat m.

Lemma 1 (Second moment of the Morris counter) E[(2Xm)2]=32m2+32m+1. \mathbb E[(2^{X_m})^2] =\frac32m^2+\frac32m+1.

Proof

We can prove the claim using an induction argument similar to our proof for the expectation. When m=1m=1, E[(2Xm)2]=4\mathbb E[(2^{X_m})^2]=4. We assume it is true for smaller mm and use the same notation XiX_i. Conditioning on Xm−1X_{m-1} gives E[22Xm∣Xm−1=i]=(1−2−i)22i+2−i22(i+1)=22i+3⋅2i. \mathbb E[2^{2X_m}\mid X_{m-1}=i] =(1-2^{-i})2^{2i}+2^{-i}2^{2(i+1)} =2^{2i}+3\cdot2^i. Hence E[22Xm]=E[22Xm−1]+3E[2Xm−1]=32(m−1)2+32(m−1)+1+3m=32m2+32m+1. \begin{aligned} \mathbb E[2^{2X_m}] &=\mathbb E[2^{2X_{m-1}}]+3\mathbb E[2^{X_{m-1}}]\\ &=\frac32(m-1)^2+\frac32(m-1)+1+3m\\ &=\frac32m^2+\frac32m+1. \end{aligned}

With the above lemma, we can compute the variance as follows: Var⁡(m^)=E[(2Xm−1)2]−m2=m(m−1)2≤m22. \operatorname{Var}(\widehat m) =\mathbb E[(2^{X_m}-1)^2]-m^2 =\frac{m(m-1)}2\le\frac{m^2}{2}. Applying Chebyshev’s inequality, we obtain that for every ε>0\varepsilon>0, P(∣m^−m∣≥εm)≤12ε2. \mathbb P(|\widehat m-m|\ge\varepsilon m) \le\frac{1}{2\varepsilon^2}. However, we observe that as ε\varepsilon becomes smaller, the above bound is not useful. Thus, it is necessary to improve the concentration of the algorithm. We now introduce two common tricks to achieve this.

The averaging trick

Chebyshev’s inequality tells us that we can improve the concentration by reducing the variance. Let’s first review some properties of variances. Let XX be a random variable. We have Var⁡(aX)=a2Var⁡(X) \operatorname{Var}(aX)=a^2\operatorname{Var}(X) for any constant aa. For any two independent random variables XX and YY, we have Var⁡(X+Y)=Var⁡(X)+Var⁡(Y). \operatorname{Var}(X+Y)=\operatorname{Var}(X)+\operatorname{Var}(Y).

We can design a new algorithm by independently running Morris’s algorithm tt times in parallel. Denote the corresponding outputs by m^1,…,m^t\widehat m_1,\ldots,\widehat m_t. The final output is m^∗:=∑i=1tm^it. \widehat m^*:=\frac{\sum_{i=1}^t\widehat m_i}{t}. By the above two properties, we have Var⁡(m^∗)=Var⁡(m^1)/t\operatorname{Var}(\widehat m^*)=\operatorname{Var}(\widehat m_1)/t. We can apply Chebyshev’s inequality to m^∗\widehat m^* and obtain P(∣m^∗−m∣≥εm)≤12tε2. \mathbb P(|\widehat m^*-m|\ge\varepsilon m) \le\frac{1}{2t\varepsilon^2}. For t≥1/(2ε2δ)t\ge1/(2\varepsilon^2\delta), we have P(∣m^∗−m∣≥εm)≤δ. \mathbb P(|\widehat m^*-m|\ge\varepsilon m)\le\delta.

The new algorithm uses approximately O ⁣(log⁡log⁡mε2δ) O\!\left(\frac{\log\log m}{\varepsilon^2\delta}\right) bits of memory. It shows a trade-off between the accuracy of the randomized algorithm and the consumption of memory space. We can further improve the dependence on δ\delta using the Chernoff bound below.

1.3 Chernoff bounds and success amplification

Like Chebyshev’s inequality, if we choose f(x)=eαxf(x)=e^{\alpha x} for α>0\alpha>0 and apply Markov’s inequality to f(X)f(X), the bound amounts to bounding E[eαX]\mathbb E[e^{\alpha X}], which is the moment generating function of XX. In case E[eαX]\mathbb E[e^{\alpha X}] can be well bounded, we obtain sharp concentration bounds.

Theorem 2 (Chernoff bound) Let X1,…,XnX_1,\ldots,X_n be independent random variables such that Xi∼Ber⁡(pi)X_i\sim\operatorname{Ber}(p_i) for each i=1,2,…,ni=1,2,\ldots,n. Let X=∑i=1nXiX=\sum_{i=1}^nX_i and denote μ:=EX=∑i=1npi\mu:=\mathbb E X=\sum_{i=1}^np_i. For δ>0\delta>0, we have P(X≥(1+δ)μ)≤(eδ(1+δ)1+δ)μ. \mathbb P(X\ge(1+\delta)\mu) \le\left(\frac{e^\delta}{(1+\delta)^{1+\delta}}\right)^\mu. If 0<δ<10<\delta<1, then we have P(X≤(1−δ)μ)≤(e−δ(1−δ)1−δ)μ. \mathbb P(X\le(1-\delta)\mu) \le\left(\frac{e^{-\delta}}{(1-\delta)^{1-\delta}}\right)^\mu.

Proof. We only prove the upper tail bound and the proof of the lower tail bound is similar. For every α>0\alpha>0, we have P(X≥(1+δ)μ)=P(eαX≥eα(1+δ)μ)≤E[eαX]eα(1+δ)μ. \begin{aligned} \mathbb P(X\ge(1+\delta)\mu) &=\mathbb P(e^{\alpha X}\ge e^{\alpha(1+\delta)\mu})\\ &\le\frac{\mathbb E[e^{\alpha X}]}{e^{\alpha(1+\delta)\mu}}. \end{aligned} Therefore, we need to estimate the moment generating function E[eαX]\mathbb E[e^{\alpha X}]. Since X=∑i=1nXiX=\sum_{i=1}^nX_i is the sum of independent Bernoulli variables, we have E[eαX]=E ⁣[∏i=1neαXi]=∏i=1nE[eαXi]. \mathbb E[e^{\alpha X}] =\mathbb E\!\left[\prod_{i=1}^ne^{\alpha X_i}\right] =\prod_{i=1}^n\mathbb E[e^{\alpha X_i}]. Since Xi∼Ber⁡(pi)X_i\sim\operatorname{Ber}(p_i), we can compute E[eαXi]\mathbb E[e^{\alpha X_i}] directly: E[eαXi]=pieα+(1−pi)=1+(eα−1)pi≤exp⁡((eα−1)pi). \mathbb E[e^{\alpha X_i}] =p_ie^\alpha+(1-p_i) =1+(e^\alpha-1)p_i \le\exp((e^\alpha-1)p_i). Therefore, E[eαX]≤∏i=1nexp⁡((eα−1)pi)=exp⁡((eα−1)μ). \mathbb E[e^{\alpha X}] \le\prod_{i=1}^n\exp((e^\alpha-1)p_i) =\exp((e^\alpha-1)\mu). Consequently, P(X≥(1+δ)μ)≤exp⁡ ⁣(μ(eα−1−α(1+δ))). \mathbb P(X\ge(1+\delta)\mu) \le\exp\!\left(\mu(e^\alpha-1-\alpha(1+\delta))\right). Note that the above holds for any α>0\alpha>0. Therefore, we can choose α\alpha so as to minimize the exponent. Its derivative is eα−1−δe^\alpha-1-\delta, which gives α=log⁡(1+δ)\alpha=\log(1+\delta). Substituting this value gives P(X≥(1+δ)μ)≤(eδ(1+δ)1+δ)μ. \mathbb P(X\ge(1+\delta)\mu) \le\left(\frac{e^\delta}{(1+\delta)^{1+\delta}}\right)^\mu.

The following form of the Chernoff bound is more convenient to use (but weaker).

Corollary 1 (Corollary) For any 0<δ<10<\delta<1, P(X≥(1+δ)μ)≤exp⁡ ⁣(−δ23μ),P(X≤(1−δ)μ)≤exp⁡ ⁣(−δ22μ). \begin{aligned} \mathbb P(X\ge(1+\delta)\mu) &\le\exp\!\left(-\frac{\delta^2}{3}\mu\right),\\ \mathbb P(X\le(1-\delta)\mu) &\le\exp\!\left(-\frac{\delta^2}{2}\mu\right). \end{aligned}

Proof

We only prove the upper tail. It suffices to verify that for 0<δ<10<\delta<1, eδ(1+δ)1+δ≤exp⁡ ⁣(−δ23). \frac{e^\delta}{(1+\delta)^{1+\delta}} \le\exp\!\left(-\frac{\delta^2}{3}\right). Taking logarithms, this is equivalent to δ−(1+δ)log⁡(1+δ)≤−δ23. \delta-(1+\delta)\log(1+\delta) \le-\frac{\delta^2}{3}. Let f(δ)=δ−(1+δ)log⁡(1+δ)+δ2/3f(\delta)=\delta-(1+\delta)\log(1+\delta)+\delta^2/3. Then f′(δ)=−log⁡(1+δ)+23δ,f′′(δ)=−11+δ+23. f'(\delta)=-\log(1+\delta)+\frac23\delta, \qquad f''(\delta)=-\frac1{1+\delta}+\frac23. For 0<δ<1/20<\delta<1/2, f′′(δ)<0f''(\delta)<0, and for 1/2<δ<11/2<\delta<1, f′′(δ)>0f''(\delta)>0. Therefore, f′f' first decreases and then increases on [0,1][0,1]. Also f′(0)=0f'(0)=0 and f′(1)<0f'(1)<0, so f′(δ)≤0f'(\delta)\le0 on [0,1][0,1]. Hence f(δ)≤f(0)=0f(\delta)\le f(0)=0.

The median trick

We can further boost the performance of Morris’s algorithm using the median trick. We choose t=⌈3/(2ε2)⌉t=\lceil3/(2\varepsilon^2)\rceil in the algorithm introduced in the averaging trick and independently run it ss times in parallel. Denote the outputs by m^1∗,m^2∗,…,m^s∗\widehat m_1^*,\widehat m_2^*,\ldots,\widehat m_s^*. It holds that for every i=1,…,si=1,\ldots,s, P(∣m^i∗−m∣≥εm)≤13. \mathbb P(|\widehat m_i^*-m|\ge\varepsilon m)\le\frac13. At last, we output the median m^∗∗\widehat m^{**} of m^1∗,…,m^s∗\widehat m_1^*,\ldots,\widehat m_s^*.

Then we can apply the Chernoff bound to analyze the result obtained by the median trick. For every i=1,…,si=1,\ldots,s, let YiY_i be the indicator of the good event ∣m^i∗−m∣<εm. |\widehat m_i^*-m|<\varepsilon m. Then Y:=∑i=1sYiY:=\sum_{i=1}^sY_i satisfies EY≥2s/3\mathbb E Y\ge2s/3. If the median m^∗∗\widehat m^{**} is bad, then at least half of the m^i∗\widehat m_i^* are bad. Equivalently, Y≤s/2Y\le s/2. By the Chernoff bound, P ⁣(∣Y−EY∣≥s6)≤2exp⁡ ⁣(−s72). \mathbb P\!\left(|Y-\mathbb E Y|\ge\frac{s}{6}\right) \le2\exp\!\left(-\frac{s}{72}\right). Therefore, for t=O(1/ε2)t=O(1/\varepsilon^2) and s=O(log⁡(1/δ))s=O(\log(1/\delta)), we have P(∣m^∗∗−m∣≥εm)≤δ. \mathbb P(|\widehat m^{**}-m|\ge\varepsilon m)\le\delta. This new algorithm uses O ⁣(1ε2log⁡1δ⋅log⁡log⁡m) O\!\left(\frac1{\varepsilon^2}\log\frac1\delta\cdot\log\log m\right) bits of memory.

1.4 Hoeffding’s inequality

One annoying restriction of the Chernoff bound is that each XiX_i needs to be a Bernoulli random variable. Hoeffding’s inequality generalizes the Chernoff bound by allowing XiX_i to follow any distribution, provided its value is almost surely bounded.

Theorem 3 (Hoeffding’s inequality) Let X1,…,XnX_1,\ldots,X_n be independent random variables where each Xi∈[ai,bi]X_i\in[a_i,b_i] for certain ai≤bia_i\le b_i almost surely. Assume EXi=pi\mathbb E X_i=p_i for every 1≤i≤n1\le i\le n. Let X=∑i=1nXiX=\sum_{i=1}^nX_i and μ:=EX=∑i=1npi\mu:=\mathbb E X=\sum_{i=1}^np_i. Then P(∣X−μ∣≥t)≤2exp⁡ ⁣(−2t2∑i=1n(bi−ai)2) \mathbb P(|X-\mu|\ge t) \le2\exp\!\left(-\frac{2t^2}{\sum_{i=1}^n(b_i-a_i)^2}\right) for all t>0t>0.

We learnt from the proof of the Chernoff bound that the key to establishing concentration inequalities of this form is to obtain a nice upper bound on the moment generating function. Therefore, the following Hoeffding lemma will be the main technical ingredient to prove the inequality.

Lemma 2 (Hoeffding’s lemma) Let XX be a random variable with EX=0\mathbb E X=0 and X∈[a,b]X\in[a,b]. Then EeλX≤exp⁡ ⁣(λ2(b−a)28)for all λ∈R. \mathbb E e^{\lambda X} \le\exp\!\left(\frac{\lambda^2(b-a)^2}{8}\right) \qquad\text{for all }\lambda\in\mathbb R.

Proof. Let ψ(λ)=log⁡E[eλX]\psi(\lambda)=\log\mathbb E[e^{\lambda X}]. We first compute its derivatives: ψ′(λ)=E[XeλX]EeλX, \psi'(\lambda)=\frac{\mathbb E[Xe^{\lambda X}]}{\mathbb E e^{\lambda X}}, and ψ′′(λ)=E[X2eλX]EeλX−E[XeλX]2E[eλX]2. \begin{aligned} \psi''(\lambda) &=\frac{\mathbb E[X^2e^{\lambda X}]}{\mathbb E e^{\lambda X}}\\ &\quad-\frac{\mathbb E[Xe^{\lambda X}]^2}{\mathbb E[e^{\lambda X}]^2}. \end{aligned} When λ=0\lambda=0, using the condition EX=0\mathbb E X=0, we have ψ(0)=ψ′(0)=0\psi(0)=\psi'(0)=0. Let PP be the distribution of XX. We can interpret ψ′′\psi'' as the variance of a tilted random variable: ψ′′(λ)=Var⁡Qλ(X)\psi''(\lambda)=\operatorname{Var}_{Q_\lambda}(X), where dQλdP(x)=eλxE[eλX]. \frac{\mathrm dQ_\lambda}{\mathrm dP}(x) =\frac{e^{\lambda x}}{\mathbb E[e^{\lambda X}]}. Since XX is supported on [a,b][a,b], for every λ\lambda, Var⁡Qλ(X)≤EQλ ⁣[(X−a+b2)2]≤(b−a)24. \begin{aligned} \operatorname{Var}_{Q_\lambda}(X) &\le\mathbb E_{Q_\lambda}\!\left[\left(X-\frac{a+b}{2}\right)^2\right]\\ &\le\frac{(b-a)^2}{4}. \end{aligned} Therefore, by Taylor’s formula, ψ(λ)=λ2∫01(1−s)ψ′′(sλ) ds \psi(\lambda)=\lambda^2\int_0^1(1-s)\psi''(s\lambda)\,\mathrm ds and hence ψ(λ)≤λ2(b−a)2/8\psi(\lambda)\le\lambda^2(b-a)^2/8.

Armed with Hoeffding’s lemma, it is routine to prove Hoeffding’s inequality.

Proof of Hoeffding’s inequality

First note that we can assume EXi=0\mathbb E X_i=0 and therefore μ=0\mu=0; if not, replace XiX_i by Xi−EXiX_i-\mathbb E X_i. Put V=∑i=1n(bi−ai)2V=\sum_{i=1}^n(b_i-a_i)^2. By symmetry, we only need to prove the upper tail. Since P(X≥t)=P(eλX≥eλt)≤E[eλX]eλt \mathbb P(X\ge t) =\mathbb P(e^{\lambda X}\ge e^{\lambda t}) \le\frac{\mathbb E[e^{\lambda X}]}{e^{\lambda t}} and E[eλX]=E ⁣[eλ∑i=1nXi]=∏i=1nE[eλXi], \mathbb E[e^{\lambda X}] =\mathbb E\!\left[e^{\lambda\sum_{i=1}^nX_i}\right] =\prod_{i=1}^n\mathbb E[e^{\lambda X_i}], applying Hoeffding’s lemma for each factor yields E[eλXi]≤exp⁡ ⁣(λ2(bi−ai)28). \mathbb E[e^{\lambda X_i}] \le\exp\!\left(\frac{\lambda^2(b_i-a_i)^2}{8}\right). Let λ=4t/V\lambda=4t/V. We have P(X≥t)≤exp⁡ ⁣(−λt+λ2V8)=exp⁡ ⁣(−2t2V). \begin{aligned} \mathbb P(X\ge t) &\le\exp\!\left(-\lambda t+\frac{\lambda^2V}{8}\right) =\exp\!\left(-\frac{2t^2}{V}\right). \end{aligned}

1.5 Multi-armed bandits: Explore-Then-Commit

Suppose there is a kk-arm bandit, and the reward of each arm follows some distribution fif_i supported on [0,1][0,1] with mean μi\mu_i; rewards are independent across arms and rounds. We assume without loss of generality that μ1≥μ2≥⋯≥μk\mu_1\ge\mu_2\ge\cdots\ge\mu_k. Now suppose you can pull the bandit for TT rounds and the goal is to obtain maximum reward in expectation. If we know μ1,…,μk\mu_1,\ldots,\mu_k, the optimal strategy is to pull arm 11 for TT times, and the expected reward is Tμ1T\mu_1. However, in case we do not know the distributions, we have to design some strategy to explore the bandit first.

Denote by ata_t the arm pulled in round tt, and thus the reward in the ttth round satisfies Xt∼fatX_t\sim f_{a_t}. The regret of a strategy is defined as the gap between Tμ1T\mu_1 and the expected rewards of the strategy in TT rounds, namely the regret of not always choosing the first arm: R(T):=Tμ1−E ⁣[∑t=1TXt]≥0. R(T):=T\mu_1-\mathbb E\!\left[\sum_{t=1}^TX_t\right]\ge0.

For every i∈[k]i\in[k], denote Δi:=μ1−μi≥0\Delta_i:=\mu_1-\mu_i\ge0 as the gap between the reward of the iith arm and the optimal arm. The naive strategy that pulls each arm equally often is bad. When kk divides TT, its regret is R(T)=∑i=1kΔik T, R(T)=\frac{\sum_{i=1}^k\Delta_i}{k}\,T, which is linear in TT. We consider a strategy/algorithm good if lim⁡T→∞R(T)/T=0\lim_{T\to\infty}R(T)/T=0, or equivalently R(T)=o(T)R(T)=o(T).

Proposition 3 For every t∈[T]t\in[T], let ni(t):=∑s=1t1{as=i} n_i(t):=\sum_{s=1}^t\mathbf 1_{\{a_s=i\}} denote the number of times arm ii is pulled in the first tt rounds. Then R(T)=∑i=2kΔi E[ni(T)]. R(T)=\sum_{i=2}^k\Delta_i\,\mathbb E[n_i(T)].

Proof

By conditioning on the chosen arm in each round, R(T)=Tμ1−E ⁣[∑t=1TXt]=Tμ1−∑t=1TE[μat]=∑t=1T∑i=1kΔi E[1{at=i}]=∑i=1kΔi E ⁣[∑t=1T1{at=i}]=∑i=1kΔi E[ni(T)]. \begin{aligned} R(T) &=T\mu_1-\mathbb E\!\left[\sum_{t=1}^TX_t\right]\\ &=T\mu_1-\sum_{t=1}^T\mathbb E[\mu_{a_t}]\\ &=\sum_{t=1}^T\sum_{i=1}^k\Delta_i\, \mathbb E[\mathbf 1_{\{a_t=i\}}]\\ &=\sum_{i=1}^k\Delta_i\, \mathbb E\!\left[\sum_{t=1}^T\mathbf 1_{\{a_t=i\}}\right]\\ &=\sum_{i=1}^k\Delta_i\,\mathbb E[n_i(T)]. \end{aligned} Since Δ1=0\Delta_1=0, this equals the stated sum.

We also write Ri(T):=ΔiE[ni(T)]R_i(T):=\Delta_i\mathbb E[n_i(T)] for every i∈[k]i\in[k], and then R(T)=∑i=1kRi(T)R(T)=\sum_{i=1}^kR_i(T).

The algorithm and the trade-off

To get small regret, our strategy should identify the best arm as soon as possible. The most straightforward way to find the best arm is to try each arm a few times and pick the one with the best empirical reward. The Explore-Then-Commit algorithm implements this idea: pull every arm ii for LL times (so kLkL times in total for exploration), and calculate μ^i\widehat\mu_i (the average reward gained in those LL times). After this, always pull an arm JJ with the greatest μ^i\widehat\mu_i. We assume kL≤TkL\le T and use any fixed rule to break ties. We can write its regret as

R(T)=L∑i=1kΔi+(T−kL)E[ΔJ]. R(T) =L\sum_{i=1}^k\Delta_i +(T-kL)\mathbb E[\Delta_J]. Moreover, E[ΔJ]=∑i=2kΔiP(J=i). \mathbb E[\Delta_J] =\sum_{i=2}^k\Delta_i\mathbb P(J=i). When i≠1i\ne1, P(J=i)≤P(μ^i≥μ^1). \mathbb P(J=i)\le\mathbb P(\widehat\mu_i\ge\widehat\mu_1).

We bound the above probability by concentration inequalities. To this end, let XjX_j be the jjth reward from fif_i, and let YjY_j be the jjth reward from f1f_1. Let Zj=Xj−Yj∈[−1,1]Z_j=X_j-Y_j\in[-1,1]; then EZj=−Δi≤0\mathbb E Z_j=-\Delta_i\le0. Let Z=∑j=1LZjZ=\sum_{j=1}^LZ_j; then EZ=−LΔi\mathbb E Z=-L\Delta_i. By Hoeffding’s inequality, P(μ^i≥μ^1)=P(Z≥0)=P(Z−EZ≥LΔi)≤exp⁡ ⁣(−2(LΔi)2∑j=1L22)=exp⁡ ⁣(−LΔi22). \begin{aligned} \mathbb P(\widehat\mu_i\ge\widehat\mu_1) &=\mathbb P(Z\ge0)\\ &=\mathbb P(Z-\mathbb E Z\ge L\Delta_i)\\ &\le\exp\!\left(-\frac{2(L\Delta_i)^2}{\sum_{j=1}^L2^2}\right)\\ &=\exp\!\left(-\frac{L\Delta_i^2}{2}\right). \end{aligned} Therefore, R(T)≤L∑i=1kΔi+T∑i=2kΔie−LΔi2/2≤∑i=1k(L+TΔie−LΔi2/2). \begin{aligned} R(T) &\le L\sum_{i=1}^k\Delta_i +T\sum_{i=2}^k\Delta_i e^{-L\Delta_i^2/2}\\ &\le\sum_{i=1}^k \left(L+T\Delta_i e^{-L\Delta_i^2/2}\right). \end{aligned}

A bound without knowing the gaps

To further upper bound R(T)R(T), define g(L,Δi):=L+TΔiexp⁡ ⁣(−LΔi22). g(L,\Delta_i) :=L+T\Delta_i\exp\!\left(-\frac{L\Delta_i^2}{2}\right). We would like to determine LL minimizing the upper bound of R(T)R(T) among all possible Δi\Delta_i, i.e., min⁡Lmax⁡ΔiR(T)\min_L\max_{\Delta_i}R(T). First we calculate max⁡Δig(L,Δi)\max_{\Delta_i}g(L,\Delta_i): ∂g(L,Δi)∂Δi=T(1−LΔi2)exp⁡ ⁣(−LΔi22). \frac{\partial g(L,\Delta_i)}{\partial\Delta_i} =T(1-L\Delta_i^2)\exp\!\left(-\frac{L\Delta_i^2}{2}\right). We have ∂g/∂Δi>0\partial g/\partial\Delta_i>0 when 0≤Δi<1/L0\le\Delta_i<1/\sqrt L, and ∂g/∂Δi<0\partial g/\partial\Delta_i<0 when 1≥Δi>1/L1\ge\Delta_i>1/\sqrt L. Thus, for all L>1L>1, g(L,Δi)≤g ⁣(L,1L)=L+Te−1/2L. g(L,\Delta_i) \le g\!\left(L,\frac1{\sqrt L}\right) =L+\frac{Te^{-1/2}}{\sqrt L}.

Finally, assuming kL≤TkL\le T and setting L=Θ(T2/3)L=\Theta(T^{2/3}), we have R(T)≤∑i=1k(L+Te−1/2L)=O(kT2/3). R(T) \le\sum_{i=1}^k\left(L+\frac{Te^{-1/2}}{\sqrt L}\right) =O(kT^{2/3}).

The Explore-Then-Commit algorithm enjoys sublinear regret for fixed kk, which is good, but still suboptimal. The main disadvantage is that it treats all arms equally in the exploration step and pulls each of them for a fixed LL times regardless of the rewards already obtained.

II. Reading Contents

The classroom arguments leave two natural questions. Can we obtain concentration when observations are dependent, or when variables are not bounded? And can we use concentration to decide how many samples to collect? The following sections develop these questions without changing the basic exponential-moment strategy.

2.1 Concentration with martingales

Our concentration proofs for sums have used mutual independence to factor an MGF. When new observations depend on the past, factorization is no longer available. Conditional expectation gives a replacement: we bound the next factor after fixing everything observed so far, and then work backward.

A fair game and its information

Imagine a fair game in which the size of the next bet may depend on previous wins and losses. The increments need not be independent. What matters is that their conditional expected value is zero.

Definition 1 (Filtrations and martingales) A filtration (Ft)t≥0(\mathcal F_t)_{t\ge0} is an increasing sequence of σ\sigma-algebras representing the information available over time. An integrable process (Zt)(Z_t) is a martingale if ZtZ_t is Ft\mathcal F_t-measurable and, for every t≥1t\ge1, E[Zt∣Ft−1]=Zt−1. \mathbb E[Z_t\mid\mathcal F_{t-1}]=Z_{t-1}. Equivalently, the increments Dt=Zt−Zt−1D_t=Z_t-Z_{t-1} have conditional mean zero.

For example, if X1,X2,…X_1,X_2,\ldots are independent and integrable, then Zt=∑i=1t(Xi−EXi)Z_t=\sum_{i=1}^t(X_i-\mathbb EX_i) is a martingale with Z0=0Z_0=0. We have also already met a dependent example: for Morris’s counter, 2Xt−1−t2^{X_t}-1-t is a martingale, by the conditional expectation calculation in Part I. Its increments are not uniformly small, so it does not directly give a useful bounded-increment estimate.

Revealing a function one coordinate at a time

Let W=f(X1,…,Xn)W=f(X_1,\ldots,X_n) be integrable. Reveal the inputs successively and keep track of the current prediction of WW: Zk=E[W∣X1,…,Xk],0≤k≤n. Z_k=\mathbb E[W\mid X_1,\ldots,X_k], \qquad 0\le k\le n. Here Z0=EWZ_0=\mathbb EW and Zn=WZ_n=W. The tower property gives E[Zk∣Fk−1]=E[W∣Fk−1]=Zk−1, \mathbb E[Z_k\mid\mathcal F_{k-1}] =\mathbb E[W\mid\mathcal F_{k-1}]=Z_{k-1}, where Fk=σ(X1,…,Xk)\mathcal F_k=\sigma(X_1,\ldots,X_k). This is the Doob martingale of WW. The construction itself does not require independent inputs.

This converts a deviation of ff into a sum of changes in conditional predictions. To prove concentration, we need to control how much one new observation can change that prediction.

Theorem 4 (Azuma–Hoeffding with conditional range bounds) Let (Zk)k=0n(Z_k)_{k=0}^n be a martingale. Suppose there are Fk−1\mathcal F_{k-1}-measurable endpoints ak,bka_k,b_k and deterministic constants ck≥0c_k\ge0 such that ak≤Zk−Zk−1≤bk,bk−ak≤ck. a_k\le Z_k-Z_{k-1}\le b_k, \qquad b_k-a_k\le c_k. If C=∑kck2>0C=\sum_kc_k^2>0, then, for every t>0t>0, P(∣Zn−Z0∣≥t)≤2e−2t2/C. \mathbb P(|Z_n-Z_0|\ge t)\le2e^{-2t^2/C}. Each one-sided bound omits the factor 22. In particular, the assumption ∣Zk−Zk−1∣≤dk|Z_k-Z_{k-1}|\le d_k gives 2exp⁡(−t2/(2∑kdk2))2\exp(-t^2/(2\sum_kd_k^2)).

Proof. Write Dk=Zk−Zk−1D_k=Z_k-Z_{k-1}. Conditional on Fk−1\mathcal F_{k-1}, the variable DkD_k has mean zero and lies in an interval of length at most ckc_k. Hoeffding’s lemma, applied to this conditional distribution, gives E[eλDk∣Fk−1]≤eλ2ck2/8. \mathbb E[e^{\lambda D_k}\mid\mathcal F_{k-1}] \le e^{\lambda^2c_k^2/8}. Because D1,…,Dk−1D_1,\ldots,D_{k-1} are already known at time k−1k-1, Eeλ∑j=1kDj=E ⁣[eλ∑j<kDjE[eλDk∣Fk−1]]≤eλ2ck2/8Eeλ∑j<kDj. \begin{aligned} &\mathbb E e^{\lambda\sum_{j=1}^kD_j}\\ &=\mathbb E\!\left[e^{\lambda\sum_{j<k}D_j} \mathbb E[e^{\lambda D_k}\mid\mathcal F_{k-1}]\right]\\ &\le e^{\lambda^2c_k^2/8}\mathbb E e^{\lambda\sum_{j<k}D_j}. \end{aligned} Iterating yields Eeλ(Zn−Z0)≤eλ2C/8\mathbb E e^{\lambda(Z_n-Z_0)}\le e^{\lambda^2C/8}. The same exponential Markov argument as before, with λ=4t/C\lambda=4t/C, proves the upper tail; use −λ-\lambda for the lower tail.

Notice the distinction between an absolute increment bound and a conditional range length. A variable in [−1,1][-1,1] has range length 22, not 11. Tracking the conditional range is what gives the sharp constant in the next application.

Bounded differences: McDiarmid’s inequality

Suppose X1,…,XnX_1,\ldots,X_n are independent. A function ff has coordinate bounded differences c1,…,cnc_1,\ldots,c_n if changing just coordinate ii changes its value by at most cic_i. Writing x(i←y)x^{(i\leftarrow y)} for xx with coordinate ii replaced by yy, this is ∣f(x)−f(x(i←y))∣≤ci. |f(x)-f(x^{(i\leftarrow y)})|\le c_i. The condition is imposed on the product of the input spaces. It describes sensitivity to one input, rather than a Euclidean Lipschitz constant.

Theorem 5 (McDiarmid’s inequality) For independent inputs and an integrable function with coordinate bounded differences c1,…,cnc_1,\ldots,c_n, put C=∑ici2>0C=\sum_i c_i^2>0. Then, for t>0t>0, P(∣f(X)−Ef(X)∣≥t)≤2e−2t2/C. \mathbb P(|f(X)-\mathbb Ef(X)|\ge t)\le2e^{-2t^2/C}.

Proof. Use the Doob martingale Zi=E[f(X)∣X1,…,Xi]Z_i=\mathbb E[f(X)\mid X_1,\ldots,X_i]. Write x<i=(x1,…,xi−1)x_{<i}=(x_1,\ldots,x_{i-1}) and X>i=(Xi+1,…,Xn)X_{>i}=(X_{i+1},\ldots,X_n). Fix the past x<ix_{<i} and define hi(u)=E[f(x<i,u,X>i)]. h_i(u)=\mathbb E[f(x_{<i},u,X_{>i})]. Independence means that the future inputs in this expectation have the same distribution for every value of uu. Couple the two expectations using the same future inputs. The bounded-difference assumption then gives ∣hi(u)−hi(v)∣≤ci. |h_i(u)-h_i(v)|\le c_i. Conditional on the past, Zi=hi(Xi)Z_i=h_i(X_i) and Zi−1Z_{i-1} is the conditional average of hi(Xi)h_i(X_i). Subtracting that average does not change the length of its range. Hence the conditional range of Zi−Zi−1Z_i-Z_{i-1} has length at most cic_i, and Theorem 4 applies.

Example 1 (Empty bins) Throw m≥1m\ge1 balls independently and uniformly into n≥2n\ge2 bins. Let YY be the number of empty bins. The empty-bin indicators are dependent, but linearity of expectation still gives EY=n(1−1n)m. \mathbb EY=n\left(1-\frac1n\right)^m. Use the positions of the mm balls as independent inputs. Moving one ball can empty its old bin or fill its new bin. If both occur, the changes cancel; in all cases the total number of empty bins changes by at most 11. McDiarmid therefore gives P(∣Y−EY∣≥t)≤2e−2t2/m. \mathbb P(|Y-\mathbb EY|\ge t) \le2e^{-2t^2/m}. The denominator counts independent input coordinates—balls, not bins.

Example 2 (Sampling without replacement) A bag contains NN balls, of which rr are red. Draw n<Nn<N balls without replacement, let YiY_i indicate that draw ii is red, and put Si=∑j=1iYjS_i=\sum_{j=1}^iY_j. We want concentration for SnS_n, whose mean is nr/Nnr/N.

The Doob prediction after ii observations is Zi=Si+(n−i)r−SiN−i. Z_i=S_i+(n-i)\frac{r-S_i}{N-i}. Let qi=(r−Si−1)/(N−i+1)q_i=(r-S_{i-1})/(N-i+1) be the conditional probability of red at draw ii. Direct subtraction gives Zi−Zi−1=N−nN−i(Yi−qi). Z_i-Z_{i-1}=\frac{N-n}{N-i}(Y_i-q_i). Given the past, this increment lies in an interval of length (N−n)/(N−i)≤1(N-n)/(N-i)\le1. Hence P(∣Sn−nrN∣≥t)≤2e−2t2/n. \mathbb P\left(\left|S_n-\frac{nr}{N}\right|\ge t\right) \le2e^{-2t^2/n}. If n=Nn=N, the total number of red balls drawn is exactly rr and no concentration estimate is needed.

Example 3 (Choosing what to reveal: the chromatic number) Let G∼G(n,p)G\sim G(n,p) and let χ(G)\chi(G) be its chromatic number. Revealing all (n2)\binom n2 independent edge indicators gives a bounded-difference constant 11 for each edge, but only yields a deviation scale of order nn.

There is a better representation. For i=1,…,ni=1,\ldots,n, let XiX_i encode the edges from vertex ii to vertices 1,…,i−11,\ldots,i-1. These nn inputs are independent because they contain disjoint sets of edge indicators. Changing XiX_i changes only edges incident to vertex ii.

Two such graphs have the same graph after deleting vertex ii. Each chromatic number lies between the chromatic number of that common graph and one more. Their difference is therefore at most 11. McDiarmid now gives P(∣χ(G)−Eχ(G)∣≥t)≤2e−2t2/n. \mathbb P(|\chi(G)-\mathbb E\chi(G)|\ge t)\le2e^{-2t^2/n}. The gain comes from choosing a more informative unit of exposure.

2.2 Sub-Gaussian random variables

Recall the calculation underlying Hoeffding. For a centered variable, a quadratic upper bound on its log-MGF led to a Gaussian-shaped tail. Boundedness was one way to obtain that upper bound; it need not be the only way.

For an integrable XX, write ψX(λ)=log⁡Eeλ(X−EX), \psi_X(\lambda)=\log\mathbb E e^{\lambda(X-\mathbb EX)}, allowing the value +∞+\infty if the MGF diverges.

Definition 2 (Sub-Gaussian variables) We call XX σ2\sigma^2-sub-Gaussian if, for every λ∈R\lambda\in\mathbb R, ψX(λ)≤λ2σ22. \psi_X(\lambda)\le\frac{\lambda^2\sigma^2}{2}. The number σ2\sigma^2 is a variance proxy. It is an upper bound on Var⁡(X)\operatorname{Var}(X), but need not equal it.

Exponential Markov, optimized at λ=t/σ2\lambda=t/\sigma^2, immediately gives, when σ2>0\sigma^2>0, P(∣X−EX∣≥t)≤2e−t2/(2σ2). \mathbb P(|X-\mathbb EX|\ge t) \le2e^{-t^2/(2\sigma^2)}. For proxy zero, XX is constant almost surely.

Example 4 (Gaussian, bounded, and Rademacher variables) If X∼N(μ,σ2)X\sim\mathcal N(\mu,\sigma^2), completing the square in the Gaussian integral gives ψX(λ)=λ2σ2/2\psi_X(\lambda)=\lambda^2\sigma^2/2 exactly.

If X∈[a,b]X\in[a,b], Hoeffding’s lemma gives variance proxy (b−a)2/4(b-a)^2/4.

In particular, a uniform random sign ξ∈{−1,1}\xi\in\{-1,1\} has proxy 11. One can also see this from Eeλξ=cosh⁡λ≤eλ2/2\mathbb Ee^{\lambda\xi}=\cosh\lambda\le e^{\lambda^2/2}.

Addition and the role of independence

If XiX_i are independent with proxies σi2\sigma_i^2, then for deterministic real weights wiw_i, ψ∑iwiXi(λ)=∑iψXi(wiλ)≤λ22∑iwi2σi2. \begin{aligned} \psi_{\sum_iw_iX_i}(\lambda) &=\sum_i\psi_{X_i}(w_i\lambda)\\ &\le\frac{\lambda^2}{2}\sum_iw_i^2\sigma_i^2. \end{aligned} Thus independent variances add at the level of proxies. Hoeffding’s inequality is a special case.

Without independence, this sum-of-squares rule can fail. For example, X1=⋯=Xn=ξX_1=\cdots=X_n=\xi gives ∑iXi=nξ\sum_iX_i=n\xi, whose variance is n2n^2. A useful general comparison follows from Cauchy–Schwarz: if X,YX,Y have proxies σX2,σY2\sigma_X^2,\sigma_Y^2, with arbitrary dependence, then Eeλ[(X−EX)+(Y−EY)]≤(Ee2λ(X−EX))1/2⋅(Ee2λ(Y−EY))1/2≤eλ2(σX2+σY2). \begin{aligned} &\mathbb Ee^{\lambda[(X-\mathbb EX)+(Y-\mathbb EY)]}\\ &\quad\le\big(\mathbb Ee^{2\lambda(X-\mathbb EX)}\big)^{1/2}\\ &\qquad\cdot\big(\mathbb Ee^{2\lambda(Y-\mathbb EY)}\big)^{1/2}\\ &\quad\le e^{\lambda^2(\sigma_X^2+\sigma_Y^2)}. \end{aligned} This gives proxy 2(σX2+σY2)2(\sigma_X^2+\sigma_Y^2) rather than σX2+σY2\sigma_X^2+\sigma_Y^2.

From one variable to a finite family

Suppose Z1,…,ZNZ_1,\ldots,Z_N are centered and each has proxy at most σ2>0\sigma^2>0. A union bound gives P ⁣(max⁡j≤N∣Zj∣≥t)≤2Ne−t2/(2σ2). \mathbb P\!\left(\max_{j\le N}|Z_j|\ge t\right) \le2N e^{-t^2/(2\sigma^2)}. There is no independence assumption between the ZjZ_j here. For example, distance errors in a random projection can be highly dependent, but a pointwise tail bound still combines with a union bound over a finite set of pairs.

The expected maximum admits a related bound. For λ>0\lambda>0, Emax⁡j∣Zj∣≤1λlog⁡Eeλmax⁡j∣Zj∣≤1λlog⁡∑jE(eλZj+e−λZj)≤log⁡(2N)λ+λσ22. \begin{aligned} &\mathbb E\max_j|Z_j|\\ &\quad\le\frac1\lambda\log\mathbb E e^{\lambda\max_j|Z_j|}\\ &\quad\le\frac1\lambda\log\sum_j\mathbb E(e^{\lambda Z_j}+e^{-\lambda Z_j})\\ &\quad\le\frac{\log(2N)}\lambda+\frac{\lambda\sigma^2}{2}. \end{aligned} Optimizing gives Emax⁡j∣Zj∣≤σ2log⁡(2N)\mathbb E\max_j|Z_j|\le\sigma\sqrt{2\log(2N)}.

Tail and moment descriptions of sub-Gaussian variables are equivalent to the MGF description up to universal changes in constants. For example, the displayed tail bound implies, for p≥1p\ge1, E∣X−EX∣p=p∫0∞tp−1P(∣X−EX∣>t) dt≤2p∫0∞tp−1e−t2/(2σ2) dt, \begin{aligned} &\mathbb E|X-\mathbb EX|^p\\ &=p\int_0^\infty t^{p-1}\mathbb P(|X-\mathbb EX|>t)\,\mathrm dt\\ &\le2p\int_0^\infty t^{p-1}e^{-t^2/(2\sigma^2)}\,\mathrm dt, \end{aligned} which gives (E∣X−EX∣p)1/p≤Cσp(\mathbb E|X-\mathbb EX|^p)^{1/p}\le C\sigma\sqrt p for a universal CC. The reverse implications and precise parameter comparisons can be found in Vershynin’s High-Dimensional Probability, Chapter 2. We use the MGF convention throughout these notes.

2.3 Sub-exponential variables and Bernstein bounds

A useful variable may fail to have an MGF for all real parameters. The square of a Gaussian is the simplest example. If g∼N(0,1)g\sim\mathcal N(0,1) and Y=g2Y=g^2, direct integration gives Eeλ(Y−1)={e−λ(1−2λ)−1/2,λ<1/2,+∞,λ≥1/2. \begin{gathered} \mathbb Ee^{\lambda(Y-1)}\\ =\begin{cases} e^{-\lambda}(1-2\lambda)^{-1/2},&\lambda<1/2,\\ +\infty,&\lambda\ge1/2. \end{cases} \end{gathered} So g2g^2 is not sub-Gaussian. Nevertheless, its log-MGF is quadratic near zero. We can still run the Chernoff argument, with a restricted choice of parameter.

Definition 3 (Sub-exponential variables: a two-parameter convention) For ν>0\nu>0 and α>0\alpha>0, call XX (ν,α)(\nu,\alpha)-sub-exponential if ψX(λ)≤νλ22for ∣λ∣≤1α. \psi_X(\lambda)\le\frac{\nu\lambda^2}{2} \qquad\text{for }|\lambda|\le\frac1\alpha. Here ν\nu is a variance-scale parameter and α\alpha limits the MGF’s usable range. Some references write ν2\nu^2 in place of our first parameter; the convention matters when adding variables.

For Y=g2Y=g^2, expanding the logarithm and bounding absolute values gives, when ∣λ∣≤1/4|\lambda|\le1/4, ψY(λ)=12∑k=2∞(2λ)kk≤λ2∑j=0∞(2∣λ∣)j=λ21−2∣λ∣≤2λ2. \begin{aligned} \psi_Y(\lambda) &=\frac12\sum_{k=2}^\infty\frac{(2\lambda)^k}{k}\\ &\le\lambda^2\sum_{j=0}^\infty(2|\lambda|)^j\\ &=\frac{\lambda^2}{1-2|\lambda|}\le2\lambda^2. \end{aligned} Thus g2g^2 is (4,4)(4,4)-sub-exponential. The absolute values ensure that the calculation also applies to negative λ\lambda.

Proposition 4 (Two tail regimes) If XX is (ν,α)(\nu,\alpha)-sub-exponential, then, for every t>0t>0, P(∣X−EX∣≥t)≤2exp⁡ ⁣[−12min⁡ ⁣{t2ν,tα}]. \mathbb P(|X-\mathbb EX|\ge t) \le2\exp\!\left[-\frac12\min\!\left\{\frac{t^2}{\nu},\frac t\alpha\right\}\right]. The quadratic regime ends at t=ν/αt=\nu/\alpha.

Proof

For 0<λ≤1/α0<\lambda\le1/\alpha, exponential Markov gives exp⁡(−λt+νλ2/2)\exp(-\lambda t+\nu\lambda^2/2) for the upper tail. If t≤ν/αt\le\nu/\alpha, choose λ=t/ν\lambda=t/\nu and obtain e−t2/(2ν)e^{-t^2/(2\nu)}. If t>ν/αt>\nu/\alpha, choose λ=1/α\lambda=1/\alpha and use −tα+ν2α2≤−t2α. -\frac t\alpha+\frac{\nu}{2\alpha^2} \le-\frac{t}{2\alpha}. Apply the same argument to −X-X and add the bounds.

For independent (νi,αi)(\nu_i,\alpha_i)-sub-exponential variables, the weighted sum S=∑iwiXiS=\sum_iw_iX_i has parameters νS=∑iwi2νi,αS=max⁡i∣wi∣αi. \nu_S=\sum_iw_i^2\nu_i, \qquad \alpha_S=\max_i|w_i|\alpha_i. Indeed, ∣λ∣≤1/αS|\lambda|\le1/\alpha_S puts every wiλw_i\lambda in its admissible interval, and the log-MGFs add. Zero weights can be omitted; if all weights vanish, the sum is deterministic.

Example 5 (The length of a Gaussian vector) For independent standard Gaussians g1,…,gdg_1,\ldots,g_d, the squared norm Q=∑igi2Q=\sum_i g_i^2 has mean dd and sub-exponential parameters (4d,4)(4d,4). Therefore, P(∣Q−d∣≥t)≤2exp⁡ ⁣[−min⁡ ⁣{t28d,t8}]. \mathbb P(|Q-d|\ge t) \le2\exp\!\left[-\min\!\left\{\frac{t^2}{8d},\frac t8\right\}\right]. For 0<ε<10<\varepsilon<1, P(∣Q/d−1∣≥ε)≤2e−dε2/8. \mathbb P(|Q/d-1|\ge\varepsilon) \le2e^{-d\varepsilon^2/8}. On the complementary event, d(1−ε)≤∥g∥2≤d(1+ε)\sqrt{d(1-\varepsilon)}\le\|g\|_2\le\sqrt{d(1+\varepsilon)}. This is a useful precursor to the Gaussian random projections used in the JL lemma.

Using variance information for bounded variables

Hoeffding uses the range of each variable but not its variance. If an event is rare, that can lose useful information, as our Bernoulli example showed. A more careful MGF expansion retains the variance.

Theorem 6 (Bernstein’s inequality) Let XiX_i be independent, with EXi=0\mathbb EX_i=0 and ∣Xi∣≤b|X_i|\le b for b>0b>0. Set V=∑iEXi2>0V=\sum_i\mathbb EX_i^2>0. Then, for t>0t>0, P ⁣(∣∑iXi∣≥t)≤2exp⁡ ⁣(−t22(V+bt/3)). \mathbb P\!\left(\left|\sum_iX_i\right|\ge t\right) \le2\exp\!\left(-\frac{t^2}{2(V+bt/3)}\right). If V=0V=0, every XiX_i vanishes almost surely.

Proof

Write vi=EXi2v_i=\mathbb EX_i^2. For every integer k≥2k\ge2, E∣Xi∣k≤vibk−2\mathbb E|X_i|^k\le v_i b^{k-2}. Also k!≥2⋅3k−2k!\ge2\cdot3^{k-2}. For ∣λ∣b<3|\lambda|b<3, the power series therefore gives EeλXi≤1+∑k=2∞∣λ∣kvibk−2k!≤1+viλ22(1−b∣λ∣/3). \begin{aligned} \mathbb Ee^{\lambda X_i} &\le1+\sum_{k=2}^\infty\frac{|\lambda|^kv_ib^{k-2}}{k!}\\ &\le1+\frac{v_i\lambda^2}{2(1-b|\lambda|/3)}. \end{aligned} Using log⁡(1+x)≤x\log(1+x)\le x and independence yields log⁡Eeλ∑iXi≤Vλ22(1−b∣λ∣/3). \log\mathbb Ee^{\lambda\sum_iX_i} \le\frac{V\lambda^2}{2(1-b|\lambda|/3)}. For the upper tail take λ=t/(V+bt/3)\lambda=t/(V+bt/3), which is strictly less than 3/b3/b. Substitution into exponential Markov gives the claimed exponent. Repeat for the negative sum.

For example, each centered bounded XiX_i above is also (2vi,2b)(2v_i,2b)-sub-exponential when vi>0v_i>0, by restricting to ∣λ∣≤1/(2b)|\lambda|\le1/(2b). The corresponding sum has the two-regime transition at t=V/bt=V/b. The smoother Bernstein bound retains more of the MGF estimate.

Example 6 (All degrees in a random graph) In G(n,p)G(n,p), the degree DvD_v of a fixed vertex has mean μ=(n−1)p\mu=(n-1)p and variance v=(n−1)p(1−p)v=(n-1)p(1-p). Put x=log⁡(2n/δ)x=\log(2n/\delta) and t=2vx+2x3. t=\sqrt{2vx}+\frac{2x}{3}. Since t2≥2x(v+t/3)t^2\ge2x(v+t/3), Bernstein gives P(∣Dv−μ∣≥t)≤2e−x\mathbb P(|D_v-\mu|\ge t)\le2e^{-x}. A union bound over vertices gives P ⁣(max⁡v∣Dv−μ∣≥t)≤δ. \mathbb P\!\left(\max_v|D_v-\mu|\ge t\right)\le\delta. The degrees are dependent, but the union bound does not require their independence. The linear xx term is essential in sparse regimes; it cannot be dropped without an additional assumption comparing npnp with log⁡n\log n.

2.4 Upper Confidence Bounds

ETC explores every arm for the same number of rounds. Ideally, an arm with a large gap Δi\Delta_i should receive fewer observations than an arm whose mean is close to the optimum. But the gaps are unknown. Can observations themselves tell us when an arm has been explored enough?

UCB answers this using optimism. Keep a plausible upper bound on each mean, and choose the arm with the largest upper bound. An arm is attractive either because its observed rewards are high, or because it has been sampled too little to rule out a high mean.

Confidence intervals under adaptive sampling

Use the bounded stochastic bandit model from Part I, with T≥K≥2T\ge K\ge2. For each arm ii, imagine an independent infinite reward stream Xi,1,Xi,2,…X_{i,1},X_{i,2},\ldots sampled in advance. The ssth time arm ii is pulled, reveal Xi,sX_{i,s}. Define the fixed-prefix means μ^i,s=1s∑j=1sXi,j. \widehat\mu_{i,s}=\frac1s\sum_{j=1}^sX_{i,j}. For 0<δ<10<\delta<1, let ℓ=log⁡2KTδ,rs=ℓ2s. \ell=\log\frac{2KT}{\delta}, \qquad r_s=\sqrt{\frac{\ell}{2s}}. For each fixed (i,s)(i,s) with 1≤s≤T1\le s\le T, Hoeffding bounds the probability of ∣μ^i,s−μi∣>rs|\widehat\mu_{i,s}-\mu_i|>r_s by δ/(KT)\delta/(KT). Taking a union bound gives P(G)≥1−δ, \mathbb P(\mathcal G)\ge1-\delta, where G=⋂i=1K⋂s=1T{∣μ^i,s−μi∣≤rs}. \mathcal G= \bigcap_{i=1}^K\bigcap_{s=1}^T \{|\widehat\mu_{i,s}-\mu_i|\le r_s\}. This event holds simultaneously for every possible sample count. We may therefore evaluate it at the random count Ni(t)N_i(t) chosen by the policy. Applying a fixed-sample Hoeffding bound directly at an adaptive random sample size would skip this justification.

The algorithm

Pull each arm once. After t≥Kt\ge K rounds, choose at round t+1t+1 an arm maximizing Ui(t)=μ^i,Ni(t)+rNi(t). U_i(t)=\widehat\mu_{i,N_i(t)}+r_{N_i(t)}. Break ties by any fixed rule. The policy here is a horizon-aware version of UCB; use δ=1/T\delta=1/T, so ℓ=log⁡(2KT2)\ell=\log(2KT^2), for the regret guarantees below.

How many times can a suboptimal arm be selected?

On G\mathcal G, an optimal arm always has index at least μ∗\mu_*. If arm ii is selected after initialization, its index is at least that of an optimal arm. At that moment, μ∗≤Ui(t)≤μi+2rNi(t). \mu_*\le U_i(t)\le\mu_i+2r_{N_i(t)}. Thus, for Δi>0\Delta_i>0, Δi≤2ℓ2Ni(t)⟹Ni(t)≤2ℓΔi2. \Delta_i\le2\sqrt{\frac{\ell}{2N_i(t)}} \quad\Longrightarrow\quad N_i(t)\le\frac{2\ell}{\Delta_i^2}. This is the number of samples before the next pull. Consequently, Ni(T)≤1+2ℓΔi2on G. N_i(T)\le1+\frac{2\ell}{\Delta_i^2} \qquad\text{on }\mathcal G. No gap was given to the algorithm; the analysis uses the true gap to explain when the confidence interval becomes narrow enough.

Theorem 7 (Regret of the horizon-aware UCB policy) In the bounded stochastic bandit model with T≥K≥2T\ge K\ge2, set ℓ=log⁡(2KT2)\ell=\log(2KT^2). The policy above satisfies RT≤∑i:Δi>0(Δi+2ℓΔi)+1. R_T\le \sum_{i:\Delta_i>0}\left(\Delta_i+\frac{2\ell}{\Delta_i}\right)+1. It also satisfies, for every ε>0\varepsilon>0, RT≤Tε+K+2Kℓε+1, R_T\le T\varepsilon+K+\frac{2K\ell}{\varepsilon}+1, and therefore RT=O(KTlog⁡T)R_T=O(\sqrt{KT\log T}).

Proof. Let QT=∑iΔiNi(T)Q_T=\sum_i\Delta_iN_i(T) be the random sum of gaps incurred. Its expectation is RTR_T, and 0≤QT≤T0\le Q_T\le T. On G\mathcal G, the pull-count bound gives the first displayed sum. The complement has probability at most 1/T1/T, so its contribution to EQT\mathbb EQ_T is at most 11.

For the second bound, separate arms with Δi≤ε\Delta_i\le\varepsilon from those with Δi>ε\Delta_i>\varepsilon. The small-gap arms contribute at most TεT\varepsilon on every path, regardless of their pull counts. On G\mathcal G, the other arms contribute at most ∑i:Δi>ε(Δi+2ℓΔi)≤K+2Kℓε. \begin{aligned} \sum_{i:\Delta_i>\varepsilon} \left(\Delta_i+\frac{2\ell}{\Delta_i}\right) &\le K+\frac{2K\ell}{\varepsilon}. \end{aligned} Add the failure-event contribution and choose ε=2Kℓ/T\varepsilon=\sqrt{2K\ell/T}. This gives 22KTℓ+K+12\sqrt{2KT\ell}+K+1; since K≤TK\le T and ℓ=O(log⁡T)\ell=O(\log T), it has the claimed order. The trivial RT≤TR_T\le T bound remains available.

This analysis has two levels of resolution. For a fixed instance, it bounds pulls in terms of each gap. For the worst case, it avoids trying to distinguish arms whose gaps are so small that pulling them has little cost. MOSS, one of the further topics below, changes the confidence schedule to reach the optimal worst-case scale without the extra logarithmic factor.

III. Further Topics

The following four topics continue the two algorithmic examples of this lecture. The first two concern how to allocate samples among alternatives; the other two concern what a small summary of a stream can preserve. Each entry states the problem and a result to understand, followed by sources. Algorithms and proofs are left to the readings.

3.1 MOSS: minimax-optimal stochastic bandits

Background. The UCB analysis above gives worst-case expected regret O(KTlog⁡T)O(\sqrt{KT\log T}). Is the logarithmic factor a necessary price for learning unknown means? MOSS—Minimax Optimal Strategy in the Stochastic case—shows that it is not. It belongs to the UCB family but uses a different exploration schedule.

Main result. For K≥2K\ge2, T≥KT\ge K, and stochastic rewards in [0,1][0,1], a horizon-aware MOSS policy achieves sup⁡bandit instancesRT≤CKT \sup_{\text{bandit instances}}R_T\le C\sqrt{KT} for a universal constant CC. A matching worst-case lower bound of order KT\sqrt{KT} holds for every policy. Thus “optimal” here means minimax-optimal expected cumulative regret, up to constants. It does not assert that one policy is best on every individual instance.

References. Jean-Yves Audibert and Sébastien Bubeck, Minimax Policies for Adversarial and Stochastic Bandits, COLT 2009, stochastic-bandit results. For a textbook treatment, see Tor Lattimore and Csaba Szepesvári, Bandit Algorithms, Chapter 9; the minimax lower-bound background is in Part IV.

3.2 Median Elimination: finding a near-best arm

Background. In ETC, exploration is used to choose an arm. We can study this identification task on its own: there is no reward objective during sampling, and we only want to return a good arm with high confidence. Uniformly estimating all KK means to accuracy ε/2\varepsilon/2 and taking the empirical best uses O(Kε−2log⁡(K/δ))O(K\varepsilon^{-2}\log(K/\delta)) samples. Must we pay that logarithmic dependence on KK?

Main result. For independent stochastic arms with rewards in [0,1][0,1], Median Elimination returns an arm JJ satisfying P(μJ≥μ∗−ε)≥1−δ \mathbb P(\mu_J\ge\mu_*-\varepsilon)\ge1-\delta using O ⁣(Kε2log⁡1δ) O\!\left(\frac{K}{\varepsilon^2}\log\frac1\delta\right) samples, for 0<ε,δ<10<\varepsilon,\delta<1. The guarantee asks for an ε\varepsilon-optimal arm, not exact identification of the unique best arm. Its performance measure is sample complexity, rather than cumulative regret.

Reference. Eyal Even-Dar, Shie Mannor, and Yishay Mansour, Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems, JMLR 7 (2006), 1079–1105. Read §3, especially the Median Elimination algorithm and Theorem 10. The reinforcement-learning sections are not needed for this topic.

3.3 Count-Min Sketch: estimating individual frequencies

Background. Morris counts the total number of stream items. Suppose instead that items belong to a universe [U][U], and after processing the stream we want to query how often a particular item appeared. An exact table of all frequencies may be too large. Let fj≥0f_j\ge0 be the frequency of item jj and m=∑jfjm=\sum_jf_j.

Main result. For insertion-only streams, Count-Min uses O ⁣(ε−1log⁡1δ) O\!\left(\varepsilon^{-1}\log\frac1\delta\right) counters and returns, for a fixed item jj, an estimate satisfying fj≤f^j≤fj+εm f_j\le\widehat f_j\le f_j+\varepsilon m with probability at least 1−δ1-\delta. This is an additive error relative to the total stream mass, not a relative-error guarantee for every individual frequency. The probability is over the sketch’s random choices for a fixed stream and fixed query; a simultaneous guarantee for all UU items requires an adjusted failure probability.

Reference. Graham Cormode and S. Muthukrishnan, An Improved Data Stream Summary: The Count-Min Sketch and its Applications, Journal of Algorithms 55(1) (2005), 58–75. Start with the sketch definition and the point-query guarantee; range queries and other applications are optional.

3.4 AMS sketch: estimating the second frequency moment

Background. Two streams can have the same length but very different repetition patterns. A useful summary of repetition is the second frequency moment F2=∑j=1Ufj2. F_2=\sum_{j=1}^U f_j^2. It equals the number of ordered pairs of stream positions carrying the same item, including pairs consisting of one position twice. Unlike the stream length, it depends on how the mass is distributed among items.

Main result. For a fixed insertion-only stream, an AMS second-moment sketch produces an estimate with P(∣F^2−F2∣≤εF2)≥1−δ. \mathbb P(|\widehat F_2-F_2|\le\varepsilon F_2)\ge1-\delta. A standard amplified version uses O(ε−2log⁡(1/δ))O(\varepsilon^{-2}\log(1/\delta)) independent sketch counters. Writing B=⌈ε−2log⁡(1/δ)⌉B=\lceil\varepsilon^{-2}\log(1/\delta)\rceil, a bit-space bound including bounded-independence seeds is O ⁣(B[log⁡(U+1)+log⁡(m+1)]) O\!\left(B[\log(U+1)+\log(m+1)]\right) for 0<ε<10<\varepsilon<1 and 0<δ<1/20<\delta<1/2. This connects the moment calculations and success amplification used for Morris to a different streaming statistic.

Reference. Noga Alon, Yossi Matias, and Mario Szegedy, The Space Complexity of Approximating the Frequency Moments, STOC 1996; journal version in Journal of Computer and System Sciences 58(1) (1999), 137–147. Focus on the F2F_2 estimator and its moment analysis; the general frequency-moment lower bounds are outside this topic.

Sources and acknowledgments. The main exposition adapts Chihao Zhang’s 2025 Topics in Advanced Algorithms Lecture 1 and Lecture 2, originally scribed by Yuchen He. The martingale reading adapts Lecture 3, originally scribed by Fangke Li and Yuchen He. The 2026 version reorganizes this material and supplies corrections and additional details.

Additional background: Roman Vershynin, High-Dimensional Probability, Chapter 2, for sub-Gaussian and sub-exponential variables; Lattimore and Szepesvári, Bandit Algorithms, Chapters 6–9, for ETC, UCB, and MOSS. Original approximate-counting source: Robert Morris, Counting Large Numbers of Events in Small Registers, Communications of the ACM 21(10) (1978), 840–842.