Curriculum-based Deep Reinforcement Learning for Home Robot Navigation with Natural Language Interaction

Teaching home robots complex tasks progressively—from simple skills to multi-step human instructions
Muhammad Angga Muttaqien
Master’s Program in Computer Science · University of Tsukuba
Advised by Prof. Akihisa Ohya · Submitted in January 2025

Abstract

Human-robot interaction is becoming increasingly important in various aspects of daily life, driving the demand for intelligent robots with navigation and natural language capabilities. This research introduces a novel approach that combines end-to-end deep reinforcement learning with a curriculum-based learning strategy to address this challenge. The proposed method employs Multi-Modal Deep Q-Network (MM-DQN) and Multi-Modal Advantage Actor-Critic (MM-A2C) algorithms to enable robots to autonomously navigate homes, interact with objects, and process users’ instructions, fostering an enhanced human-machine interaction experience. To optimize the learning process, the research designs and evaluates five different curricula, each structured based on its concepts and underlying assumptions. These curricula address specific aspects of robot functionality, such as spatial navigation and natural language understanding. This approach allows the robot to acquire foundational skills individually before integrating them into more complex behaviors. The gradual progression ensures a balance between learning efficiency and adaptability to various tasks and environments. Comprehensive evaluation forms a critical part of this research, focusing on simulations to assess the robot’s performance. Metrics such as accomplishment rate, generalization capability, scalability, and robot behavior inspection are evaluated to validate the approach.

Method

The robot learns an end-to-end policy from what it sees and what a human asks it to do. Visual and language features are combined to guide navigation and object-interaction actions.

RGB Image + Human Instruction Robot's first-person observation
Visual & Language Encoder ResNet-18/32
Word2Vec + LSTM
Multi-modal Deep RL Two model options:
MM-DQN/MM-A2C
Navigation & Manipulation Action Move, pick, place, put,
open, close, break, slice
Architecture of the proposed multi-modal deep reinforcement learning framework

Curriculum Learning

Instead of learning the entire instruction at once, the robot can progressively build the skills needed for the final task. Different curriculum strategies are compared to study how task progression affects learning performance. For example, "find the bread, take it, go to the fridge, and place it inside" can be learned progressively from simpler subtasks.

Find action

"find/discover/identify"

Take action

"take/grab/obtain"

Go action

"go/move/travel"

Place action

"place/put/locate"

This work compares five curriculum designs to study how different task progressions affect learning:

🪜 Incremental CL Incrementally adds task complexity
🔄 Reverse CL Builds the task backward from the final goal
🎲 Random CL Introduces controlled randomness
🏆 Mastery CL Masters one complete instruction at a time
🧩 Constructive CL Constructs complex tasks from simpler skills

Quantitative Results

Curriculum design has a strong effect on whether the agent can successfully learn complex multi-step instructions. Incremental and Random Curriculum Learning emerge as the most effective strategies across the main experiments.

ICL + RCL
The two strongest curriculum strategies based on the learning-curve experiments.
48.89 vs 17.60
Mean cumulative reward for ICL vs RCL on the Stage-4 vase trashcan task.
3 6 9
Instruction-set scaling tested from three tasks to six and nine tasks.

Detailed learning curves, generalization, scalability, and sensitivity analyses are available in the thesis.

Takeaway

What if a robot could handle complex household instructions by first mastering the basic skills they are built from? Incremental Curriculum Learning (ICL) aligns particularly well with this problem structure, because the robot first learns fundamental abilities such as navigation, object detection, and grasping before combining them into complete tasks like finding, taking, and placing an object. By integrating Curriculum Learning with Multi-modal Deep Reinforcement Learning (MM-DQN/MM-A2C), this research provides a structured approach to handling complex human instructions and demonstrates improved task accomplishment through progressive learning.

🚀 As for future works, since the approach shows limited scalability to a larger number of objects and scenes, more advanced Deep RL models, such as world models (e.g., Dreamer/JEPA), could be explored to address generalization across diverse home environments.

Related Publication

Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning
Muhammad A. Muttaqien, Ayanori Yorozu, Akihisa Ohya
IEEE International Conference on Cybernetics and Intelligent Systems (CIS), Hangzhou, China, 2024

BibTeX

@inproceedings{muttaqien2024curriculum,
  author    = {Muhammad A. Muttaqien and Ayanori Yorozu and Akihisa Ohya},
  title     = {Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning},
  booktitle = {2024 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE International Conference on Robotics, Automation and Mechatronics (RAM)},
  year      = {2024},
  doi       = {10.1109/CIS-RAM61939.2024.10673231}
}