Recent progress in self-supervised and generative speech modeling has led to the emergence of textless Speech Language Models (SpeechLMs), capable of learning directly from raw speech without textual supervision [Arora et al., 2025]. These models constitute a promising computational framework for studying language acquisition in a manner closer to how human infants learn language before literacy [Dupoux, 2018], while also offering an opportunity to draw inspiration from human learning mechanisms to design more adaptive and grounded conversational AI systems. Despite recent progress, current textless SpeechLMs still struggle to acquire higher-level linguistic competencies such as robust lexical representations, syntax, or semantics when trained on amounts of speech data comparable to those available to human infants during development, or when learning from ecologically realistic data [Lakhotia et al., 2021; Lavechin et al., 2024]. This discrepancy suggests that scale alone may not be sufficient for language acquisition. One possible explanation is that current models largely lack the multimodal and social grounding mechanisms that characterize early human learning [Dupoux, LeCun & Malik, 2026]. In natural development, language emerges through rich perceptual and communicative interactions combining speech, vision, action, attention, and adaptive social feedback. Understanding the role of such grounding mechanisms therefore constitutes an important challenge for both developmental science and the next generation of conversational AI systems.
Recent computational models of visually grounded speech learning have progressively explored how lexical structure may emerge from the joint statistics of speech and vision. Early work on cross-situational word learning showed that learners can associate spoken forms with visual referents by accumulating evidence across individually ambiguous situations [Smith & Yu, 2008]. Subsequent computational models extended this idea to continuous speech, showing that visual grounding can also support speech segmentation and the discovery of word-like units [Räsänen & Rasilo, 2015; Havard et al., 2019]. More recent visually grounded speech models based on self-supervised and contrastive learning further demonstrated that aligning raw speech with concurrent visual scenes can lead to the emergence of phonological and lexical representations without textual supervision [Harwath et al., 2018; Chrupała, 2022; Khorrami & Räsänen, 2025].
Beyond perceptual grounding, a growing line of computational research emphasizes the role of social interaction in language acquisition. In natural development, children are not merely exposed to multimodal sensory input, but actively engage in communicative exchanges where they produce linguistic behaviors, receive feedback, and progressively adapt their internal representations through interaction. Recent computational models have therefore started to investigate how learning may emerge from the interplay between perception, production, and social feedback. In particular, Nikolaus and Fourtassi [2021] proposed a neural model integrating both perception-based and production-based learning, showing that active language production and interaction-driven feedback improve semantic acquisition beyond passive perceptual learning alone. Their results highlight the importance of modeling language development not only as multimodal statistical learning from sensory input, but also as an interactive and socially guided process in which learners actively participate in the construction of their linguistic knowledge.
Nevertheless, current computational models of multimodal and social grounding still suffer from several important limitations. Most visually grounded models are trained on highly simplified image-captioning datasets that only weakly reflect the richness and ambiguity of real infant experience. Conversely, models attempting to incorporate social interaction often rely on text-based representations as an intermediate backbone, thereby ignoring many crucial communicative signals conveyed directly through speech itself, including the prosodic and interactive cues characteristic of child-directed speech (CDS), such as prosodic emphasis, repetition, corrective feedback, and interactive clarification strategies. More fundamentally, existing approaches rarely model language acquisition as a dynamic co-adaptation process between the child and the caregiver. Modeling how such multimodal and social interactions shape language learning therefore remains a major challenge, and constitutes a central motivation of the present PhD project.
- Arora, S., Chang, K. W., Chien, C. M., Peng, Y., Wu, H., Adi, Y., & Watanabe, S. (2025). On the Landscape of Spoken Language Models: A Comprehensive Survey. arXiv preprint arXiv:2504.08528.
- Chrupała, G. (2022). Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques. Journal of Artificial Intelligence Research, 73, 673–707.
- Dupoux, E. (2018). Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language learner. Cognition, 173, 43–59.
- Dupoux, E., LeCun, Y., & Malik, J. (2026). Why AI systems don’t learn and what to do about it: Lessons on autonomous learning from cognitive science. arXiv preprint arXiv:2603.15381.
- Harwath, D., Recasens, A., Surís, D., Chuang, G., Torralba, A., & Glass, J. (2018). Jointly discovering visual objects and spoken words from raw sensory input. Proceedings of ECCV, 649–665.
- Havard, W. N., Chevrot, J.-P., & Besacier, L. (2019). Word recognition, competition, and activation in a model of visually grounded speech. Proceedings of CoNLL, 339–348.
- Khorrami, K., & Räsänen, O. (2025). A model of early word acquisition based on realistic-scale audiovisual naming events. Speech Communication, 167.
- Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., & Dupoux, E. (2021). On generative spoken language modeling from raw audio. Transactions of the Association for Computational Linguistics, 9, 1336–1354.
- Lavechin, M., de Seyssel, M., Métais, M., Metze, F., Mohamed, A., Bredin, H., Dupoux, E., & Cristia, A. (2024). Modeling early phonetic acquisition from child-centered audio data. Cognition, 245, 105734.
- Nikolaus, M., & Fourtassi, A. (2021). Modeling the interaction between perception-based and production-based learning in children’s early acquisition of semantic knowledge. Proceedings of the 25th Conference on Computational Natural Language Learning (CoNLL), 391–407.
- Räsänen, O., & Rasilo, H. (2015). A joint model of word segmentation and meaning acquisition through cross-situational learning. Psychological Review, 122(4), 792.
- Smith, L., & Yu, C. (2008). Infants rapidly learn word-referent mappings via cross-situational statistics. Cognition, 106(3), 1558–1568.