Resource type
Thesis type
(Thesis) Ph.D.
Date created
2025-10-06
Authors/Contributors
Author: Raychaudhuri, Sonia
Abstract
Instruction following in embodied AI have seen significant progress in recent years, yet the gap between task execution and true language understanding remains unresolved. This thesis investigates how embodied navigation agents can be designed to understand language instructions and generalize across navigation datasets and environments. I pose three key questions: (i) Do navigation agents understand language? (ii) Can they generalize to unknown environments? and (iii) What kind of spatial semantic representations support these abilities? Through a series of completed works, I explore how representation choices, from language-aligned supervision to 2D semantic grid maps, to open-vocabulary topological maps, to multi-layer feature maps, shape both language understanding and generalization. I also discuss how these works tackle one or more challenges - inadequate language grounding, poor generalization, lack of expressive representations, lack of evaluation of language understanding and limited support for spatial reasoning. I find that: (i) what agents can understand depends on how it represents the world; (ii) generalization is shaped by model architecture and decoupling high-level reasoning from low-level control; and (iii) representational design is guided by the linguistic complexity of the instructions and the task requirements.
File
Extent
111 pages.
Identifier
etd24061
Copyright statement
Copyright is held by the author(s).
Academic Supervisor
Thesis advisor: X., Chang, Angel
Language
English
Member of collection
| Download file | Size |
|---|---|
| etd24061.pdf | 9.81 MB |