publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
2026
- CoRL
Plann3r: Predicting Planning Costs Grounded in 3DAditya Vadali, Krish Pandya, Vansh Garg, and 5 more authorsIn 10th Annual Conference on Robot Learning, 2026Planning paths or trajectories for robot navigation requires understanding scene geometry and traversability. Classic approaches relied on accurate 3D maps to define occupancy-based planning costs. Learning-based alternatives predict planning costs either in terms of distance to goal or temporal distance between images. The former overfits to scene layout and the latter lacks geometric understanding. Most of these methods estimate image-level scalar costs, which are not sufficient to guide the robot. We propose Plann3r, a 3D-grounded method that predicts pixel-level planning costs in terms of geodesic distances for any given set of images with an arbitrary subgoal pixel. We integrate Plann3r within a navigation pipeline, VGGT-Nav, in two ways. The offline mapping cum global-planning phase iteratively uses Plann3r to generate subgoals and reference-image global costmaps. In the execution phase, Plann3r performs simultaneous localization and local planning to generate planning costmaps that directly condition a learnt control policy. Plann3r and VGGT-Nav outperform baselines on the planning and navigation tasks of an existing benchmark, and we show real-world demonstrations of sim-to-real deployment.
- IROS
PixelLoop: Shortcut Topological Navigation with Pixel-Level LoopsSarthak Chittawar, Vansh Garg, Aditya Vadali, and 4 more authorsIn IEEE/RSJ International Conference on Intelligent Robots, 2026Although topological mapping and navigation have been studied extensively, the specific role and downstream effect of loop closures in purely topological representations has received relatively little attention. Importantly, loop closure over topological maps is distinct from loop closure over globally referenced trajectories and metric maps. Building on recent denser topologies grounded in pixel-level, relative 3D geometry, we propose PixelLoop which introduces loop closures directly in pixel space. Unlike sparse image-level edges or pose-graph corrections in SLAM, our pixel-level closures act as dense topological shortcuts that alter planning connectivity and cost propagation rather than merely aligning coordinates. This dense connectivity enables stable any-point-to-any-point navigation and produces costmaps that align accurately with geometric shortest paths. In particular, we showcase the distinct advantage of applying loop closures to fine-grained pixel topologies rather than image-level topologies. Across extensive simulated experiments, PixelLoop achieves over 35% absolute improvement in both Success Rate and SPL compared to image-relative baselines, with the largest gains in scenarios requiring shortcut exploitation. Results are further validated through real-world mobile robot deployments, demonstrating that dense pixel-level loop closures provide a practical and robust foundation for topological visual navigation. Project Page: https://pixelloop-nav.github.io/
- ICRA
MASt3R-Nav: WayPixel Navigation in Relative 3D MapsVansh Garg, Rohit Jayanti, Krish Pandya, and 5 more authorsIn IEEE International Conference on Robotics & Automation, 2026Visual navigation ability is strongly tied to its underlying representation of the world. Unlike classical 3D maps that require globally-consistent geometry, image- or object-relative topological graphs almost entirely do away with geometric understanding. But, this comes at the cost of navigation capability, often limiting it to merely teach-and-repeat. In this work, we propose a novel map representation in the form of pixel-relative connectivity, which is geometrically accurate but does not require global geometric consistency. Inspired by recent progress in 3D grounded image matching, we construct a map from an image sequence through inter-image connectivity based on pixel correspondences in the relative 3D coordinate systems of individual image pairs. We then use this pixel-level graph to perform global path planning by approximating and sparsifying intra-image pixel connectivity. Through this, we derive a ”WayPixel Costmap” representation and train a controller conditioned on it to predict a trajectory rollout. We show that this dense pixel-level costmap based on relative geometry is a more accurate conditioning variable for control prediction than its image- and object-level counterparts. This enables a highly capable navigation system, as validated on four types of navigation tasks in the simulator and through real world demonstrations.
2025
- NeurIPS
SegMASt3R: Geometry Grounded Segment MatchingRohit Jayanti*, Swayam Agrawal*, Vansh Garg*, and 4 more authorsIn The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025* denotes equal contributionSegment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features, segment matching captures structured regions, offering greater robustness to occlusions, lighting variations, and viewpoint changes. In this paper, we leverage the spatial understanding of 3D foundation models to tackle wide-baseline segment matching, a challenging setting involving extreme viewpoint shifts. We propose an architecture that uses the inductive bias of these 3D foundation models to match segments across image pairs with up to 180 degree view-point change rotation. Extensive experiments show that our approach outperforms state-of-the-art methods, including the SAM2 video propagator and local feature matching methods, by up to 30% on the AUPRC metric, on ScanNet++ and Replica datasets. We further demonstrate benefits of the proposed model on relevant downstream tasks, including 3D instance mapping and object-relative navigation.