Tor Lattimore

Bayesian/minimax duality for adversarial bandits

Posted onMarch 17, 2019March 7, 20201 Comment

The Bayesian approach to learning starts by choosing a prior probability distribution over the unknown parameters of the world. Then, as the learner makes observation, the prior is updated using Bayes rule to form the posterior, which represents the new Continue Reading

The variance of Exp3

Posted onFebruary 16, 2019March 17, 2019Leave a comment

In an earlier post we analyzed an algorithm called Exp3 for $k$-armed adversarial bandits for which the expected regret is bounded by \begin{align*} R_n = \max_{a \in [k]} \E\left[\sum_{t=1}^n y_{tA_t} – y_{ta}\right] \leq \sqrt{2n k \log(k)}\,. \end{align*} The setting of Continue Reading

First order bounds for k-armed adversarial bandits

Posted onFebruary 16, 2019October 28, 20214 Comments

To revive the content on this blog a little we have decided to highlight some of the new topics covered in the book that we are excited about and that were not previously covered in the blog. In this post Continue Reading

Bandit Algorithms Book

Posted onJuly 27, 2018September 10, 20192 Comments

Dear readers After nearly two years since starting to write the blog we have at last completed a first draft of the book, which is to be published by Cambridge University Press. The book is available for free as a Continue Reading

Bandit tutorial slides and update on book

Posted onFebruary 9, 2018March 12, 20192 Comments

This website has been quiet for some time, but we have not given up on bandits just yet. First up, we recently gave a short tutorial at AAAI that covered the basics of finite-armed stochastic bandits and stochastic linear bandits. Continue Reading

Ellipsoidal Confidence Sets for Least-Squares Estimators

Posted onNovember 13, 20168 Comments

Continuing the previous post, here we give a construction for confidence bounds based on ellipsoidal confidence sets. We also put things together and show bound on the regret of the UCB strategy that uses the constructed confidence bounds. Constructing the Continue Reading

Lower Bounds for Stochastic Linear Bandits

Posted onOctober 20, 2016March 17, 20192 Comments

Lower bounds for linear bandits turn out to be more nuanced than the finite-armed case. The big difference is that for linear bandits the shape of the action-set plays a role in the form of the regret, not just the Continue Reading

High probability lower bounds

Posted onOctober 14, 20161 Comment

In the post on adversarial bandits we proved two high probability upper bounds on the regret of Exp-IX. Specifically, we showed: Theorem: There exists a policy $\pi$ such that for all $\delta \in (0,1)$ for any adversarial environment $\nu\in [0,1]^{nK}$, Continue Reading

Instance dependent lower bounds

Posted onSeptember 30, 20166 Comments

In the last post we showed that under mild assumptions ($n = \Omega(K)$ and Gaussian noise), the regret in the worst case is at least $\Omega(\sqrt{Kn})$. More precisely, we showed that for every policy $\pi$ and $n\ge K-1$ there exists Continue Reading

Optimality concepts and information theory

Posted onSeptember 22, 201613 Comments

In this post we introduce the concept of minimax regret and reformulate our previous result on the upper bound on the worst-case regret of UCB in terms of the minimax regret. We briefly discuss the strengths and weaknesses of using Continue Reading