Failure-Aware LLM-DWA Replanning for Mobile Robot Navigation in Dynamic Obstacle Environments
Dabin Kim and Youngmin Lee are co-first authors.
Download paper PDFResearch summary
This study evaluates failure-aware LLM-DWA replanning in a simulated maze with three moving obstacles.
The proposed method achieved the highest observed navigation success rate among the tested configurations, reaching the goal in 5/10 trials compared with 1/10 for one-shot LLM-DWA and 0/10 for NavFn-DWA and both periodic replanning baselines.
Ablation experiments further examined how trigger timing, contact-based replanning, and LLM temperature affect navigation outcomes.
My contribution
- Co-defined the follow-up research topic, extending failure-aware LLM-DWA to dynamic obstacle environments.
- Designed and conducted experiments in dynamic obstacle environments, from initial setup through validation of the results.
- Defined and tuned the failure-detection signals, thresholds, and experimental parameters.
Navigation in Dynamic Environments
Moving obstacles can make an initially usable waypoint sequence ineffective during navigation. A local planner may oscillate, stop making progress, or become trapped, while frequent periodic replanning can repeatedly cancel active goals and disrupt execution.
This study investigates event-triggered LLM-DWA replanning as a way to respond to navigation failures without continuously regenerating waypoints. The LLM provides high-level route guidance, while the navigation stack handles path generation and DWA-based local motion.
Failure-Aware Replanning
The monitor tracks Nav2’s remaining path distance, a scalar motion-speed measure incorporating translational and rotational motion, and recent recovery events. The progress reference is updated when the remaining path distance improves by more than 0.2 m.
A monitored replanning request is triggered when either the robot is moving slowly and sufficient time has elapsed since the last progress update, or the number of recent recovery events reaches a threshold. Both conditions are subject to a replanning cooldown.
The main Proposed-40 configuration uses a 40-second stagnation threshold, a speed threshold of 0.016 m/s, five recovery events within a 60-second window, and a 40-second cooldown. The monitor is evaluated on Nav2 feedback updates.
A separate monitor runs every second to detect a 15-second feedback timeout. Additional safeguard paths request replanning following goal rejection, an unsuccessful navigation action other than intentional cancellation, or a feedback timeout.
When replanning is requested, the remaining waypoints are discarded and a new sequence is generated from the robot’s current pose to the final goal. Contact-based triggering is disabled in the main configuration.
Dynamic Maze Setup
The robot navigates from (-8.0, -8.0) to (8.0, 6.0) in a simulated maze containing three box-shaped obstacles moving back and forth at 0.15 m/s.
GPT-4o-mini, with a temperature of 0.5, generates waypoints from the current robot pose, final goal, and map-level obstacle information extracted from the simulation world file. Moving obstacle positions are not supplied to the LLM at every control step.
Instead, the LLM uses the static map structure for high-level guidance, while the local controller and failure monitor respond to dynamic obstacle interactions. The LLM returns a JSON list of two-dimensional waypoints.
These are assigned sequentially as intermediate navigation goals rather than executed directly as a continuous trajectory.
Comparative Evaluation
Five configurations were evaluated over ten trials each: NavFn-DWA, one-shot LLM-DWA, periodic LLM-DWA with 12-second and 25-second intervals, and the proposed failure-aware method.
Success requires reaching the final goal within 600 seconds with a final distance below 0.30 m. Trials fail if the robot exceeds the navigation time limit, exceeds 50 replanning attempts, or the navigation action is aborted by the motion planning system.

Proposed-40 succeeded in 5/10 trials, compared with 1/10 for one-shot LLM-DWA and 0/10 for NavFn-DWA and both periodic configurations.
Observed failure modes included start-stage planning failure, stale waypoint sequences, excessive replanning, dead-end trapping, and timeout loops.

Exploratory two-sided Fisher exact tests yielded p = 0.0325 against each 0/10 baseline and p = 0.1409 against one-shot LLM-DWA. None of the comparisons remained statistically significant after Bonferroni correction for four comparisons.
The results therefore support an initial feasibility finding within the evaluated environment.
Ablation Experiments
The study compared several trigger and LLM configurations. Proposed-45 used a 45-second stagnation threshold, five recovery events, and a 45-second cooldown, succeeding in 2/10 trials.
Proposed-35 used 35 seconds, four recovery events, and a 35-second cooldown, succeeding in 1/10 trials. Proposed-40 achieved the highest observed success rate at 5/10.
The more conservative setting delayed necessary replanning, while the faster setting retained timeout failures. The contact-aware variant succeeded in 0/10 trials and exhibited over-reactive replanning after physical contact.
The temperature-zero variant also succeeded in 0/10 trials, showing that deterministic LLM output alone did not resolve the observed navigation failures.
A representative successful Proposed-40 run covered 71.552 m in 436.315 seconds, using three LLM calls and two replanning events. These measurements describe one selected successful run rather than averages across all trials.

Failure Analysis and Future Work
Remaining failures mainly involved physical interactions, including the robot being pushed by an obstacle or trapped between a moving obstacle and a wall.
These observations motivate combining high-level waypoint replanning with local recovery behaviors that account for physical interactions.
The evaluation covers one controlled dynamic maze with ten trials per configuration.
Future work will test additional layouts and obstacle motion patterns, collect complete trial-level performance distributions, and integrate map- or VLM-grounded perception with local trap-escape behaviors.