Recent approaches in conversational AI for social assistive robots primarily rely on a human-robot substitution model, where the
robot replaces a human agent. The ĀnandaBot project (PEPR eNSEMBLE, France 2030, 2026–2030) instead investigates a human
(human/robot) approach, in which a robot sidekick accompanies and supports a human agent in a triadic collaborative task — much
as Ananda supported Buddha. The project is coordinated by Fabrice Lefèvre (LIA, Avignon Université) and brings together LIA
(Avignon Université), Inria (RobotLearn team), LISN (CNRS/Université Paris-Saclay), and AP-HP (Broca Hospital), with an application scenario in a gerontology day-hospital setting.
ĀnandaBot develops an audiovisual processing chain enabling triadic verbal and non-verbal interactions between a human user, a
human agent, and their robot sidekick. In this context, the offer is about continual audiovisual robot perception, and aims to develop
methods and algorithms to continuously extract cues about human behaviour from audio and visual data in real-world, socially
situated interactions, with two central requirements: (i) robustness to real-world perturbations, with quantitative estimation of the
reliability of extracted cues, and (ii) continuous adaptation to variations of these perturbations over time. We will cover three tasks:
behaviour understanding from visual inputs, continual audiovisual adaptation, and adaptation to simulated environments.
What do we offer?
- A 24-month postdoctoral contract at Inria Grenoble Rhône-Alpes, within the RobotLearn team.
- Remuneration according to Inria’s postdoctoral salary scale, including standard Inria employee benefits (health insurance, paid leave, restaurant subsidy, etc.).
- Access to Inria’s computing infrastructure and to the ĀnandaBot project’s dedicated hardware (GPU servers, robot platforms).
- Integration in a well-funded, multi-site national consortium (LIA, Inria, LISN, AP-HP) with a clinical deployment site at Broca
Hospital.
- Support for travel to project meetings, conferences, and the yearly ĀnandaBot workshop.
References:
1. M. Marge, C. Espy-Wilson, and N. Ward, "Spoken Language Interaction with Robots: Research Issues and Recommendations," Report from the NSF Future Directions Workshop, 2019.
2. Y. Liu et al., "Continual learning for VLMs: A survey and taxonomy beyond forgetting," 2025.
3. Y. Xu, Y. Ban, G. Delorme, C. Gan, D. Rus, and X. Alameda-Pineda, "Transcenter: Transformers with dense representations for multiple-object tracking," IEEE TPAMI, 2022.
4. L. Vaquero, Y. Xu, X. Alameda-Pineda, V. M. Brea, and M. Mucientes, "Lost and found: Overcoming detector failures in online multi-object tracking," ECCV, 2024.
5. Y. Ban, X. Alameda-Pineda, L. Girin, and R. Horaud, "Variational Bayesian inference for audio-visual tracking of multiple speakers," IEEE TPAMI, 2019.
6. X. Alameda-Pineda et al., "Socially pertinent robots in gerontological healthcare," International Journal of Social Robotics, 2025.
7. A. Golmakani, M. Sadeghi, X. Alameda-Pineda, and R. Serizel, "A weighted-variance variational autoencoder model for speech enhancement," ICASSP, 2024.
8. J.-E. Ayilo, M. Sadeghi, R. Serizel, and X. Alameda-Pineda, "Diffusion-based unsupervised audio-visual speech enhancement," ICASSP, 2025.
9. S. Sadok, S. Leglaive, L. Girin, X. Alameda-Pineda, and R. Séguier, "A multimodal dynamical variational autoencoder for audiovisual speech representation learning," Neural Networks, 2024.
10. A. Ballou, X. Alameda-Pineda, and C. Reinke, "Variational meta reinforcement learning for social robotics," Applied Intelligence, 2023.
11. R. Aljundi, K. Kelchtermans, and T. Tuytelaars, "Task-free continual learning," CVPR, 2019.
12. S. Mo, W. Pian, and Y. Tian, "Class-incremental grouping network for continual audio-visual learning," ICCV, 2023.