Multi-agent learning

Decentralized Multi-Armed Bandits on Directed Networks

Topology-aware regret analysis for decentralized multi-armed bandits over strongly connected directed graphs.

Year
2025
Status
Ongoing
Topics
Bandits · Directed graphs · Regret analysis
Learning agents sharing uncertain reward estimates over a directed network

A conceptual illustration of the mathematical idea, not a plot of measured experimental results.

01

The question

How does a directed communication topology change what a team of learning agents can discover together?

02

Central insight

Information is not shared symmetrically in a directed network. The regret must account for both exploration uncertainty and the cost of uneven information flow.

03

Approach

  1. 01

    Model local actions and rewards at agents connected by a directed graph.

  2. 02

    Study how information exchange affects sequential decisions under uncertainty.

  3. 03

    Develop the analysis over strongly connected communication graphs.

04

My contribution

  • Studied sequential decision-making under uncertainty across networked agents using decentralized bandit models.
  • Developing a journal extension of decentralized bandit analysis over strongly connected graphs.
05

Result

Journal extension in preparation, building on earlier conference work.

06

What remains

Complete the journal extension and its analysis of networked decision-making.