Skip to content

Reinforcement Learning & Decision Systems

Decision policies that adapt to changing conditions, reward design, training in simulation and safe transfer to the field.

Reinforcement Learning & Decision Systems

Why we work in this area

In some problems the right answer is not fixed: the other side changes its behaviour too. A static rule set, or a classifier trained once, becomes predictable in that setting — and what is predictable gets circumvented.

Reinforcement learning answers this class of problem, but brings difficulties of its own: when the reward is designed wrongly, the system optimises the reward rather than the thing you wanted. That is why the weight of our work sits on reward design and safe exploration.

What we work on

  • Reward design and detecting reward hacking
  • Training in simulation and transferring to the field; closing the sim-to-real gap
  • Decision policies that operate alongside language models
  • Safe exploration: preventing the system from making an irreversible mistake while learning
  • How success is measured across multi-step decisions, and catching policy regressions

From our research projects

Both of our funded projects in this area come from the same idea: a protection strategy becomes predictable once it stays fixed. Our software protection and threat detection systems each update their strategy through reinforcement learning according to the situation they encounter.

A technical assessment for your AI project

Your project's feasibility, risks and timeline are assessed in a technical consultation.