A dialogue policy playground

A dialogue agent often has more than one goal. It may need to move a task forward while keeping the conversation supportive. This small playground makes that trade-off visible: change the task focus and temperature to see how the probabilities of three dialogue acts change.

The scores below are chosen for illustration. This is a teaching example, rather than a trained model or a reproduction of the experiments in the publications.

Try it

  • Task focus: move towards 1 to favour task progress, or towards 0 to favour social support.
  • Temperature: lower values concentrate probability on the highest-scoring act; higher values spread it across the available acts.

The notebook runs Python in your browser through Marimo. Its first load can take a moment. You can also download the notebook to run it locally.

How it works

Each act has a task score and a social score. Task focus sets their relative weight:

\[s(a) = w\,s_{\mathrm{task}}(a) + (1-w)\,s_{\mathrm{social}}(a).\]

A softmax converts those scores into probabilities, with temperature $T$ controlling how concentrated the distribution is:

\[P(a) = \frac{\exp(s(a)/T)}{\sum_b \exp(s(b)/T)}.\]

At high task focus, offering advice receives a larger share of the probability. At low task focus, reflection becomes more likely. Neither setting is universally better: a useful choice depends on the conversation and the people involved.

Real dialogue policies need context, learned representations, and evaluation with users. These fixed scores leave those questions open. For related work on social goals and conversational strategies, see the publications page.

Run or adapt the example locally Install Marimo with `pip install marimo`, download the notebook linked above, then run `marimo edit dialogue-policy.py`. The code uses only Marimo and Python's standard library.