{"data":[{"id":"10.48550/arxiv.2607.16015","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.16015","identifiers":[{"identifier":"2607.16015","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Leon","familyName":"Jungemeyer","name":"Jungemeyer, Leon","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Alejandro","familyName":"Magaña","name":"Magaña, Alejandro","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Gautham","familyName":"Mohan","name":"Mohan, Gautham","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Matthias","familyName":"Karl","name":"Karl, Matthias","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Daniel","familyName":"Werdehausen","name":"Werdehausen, Daniel","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Computer Vision and Pattern Recognition (cs.CV)","lang":"en","subjectScheme":"arXiv"},{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-17T14:48:43Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-20T00:46:26Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-20T13:35:51Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:37:28Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"6D pose estimation remains a key challenge in robotics and computer vision, particularly in industrial environments. The deployment of currently available data-driven methods is often limited by resource-intensive data pipelines, reliance on textured 3D models, and sensitivity to geometric deviations caused by damages or assembly defects. We present PIXIE, a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model. Synthetic depth and normal maps are rendered from sampled reference viewpoints and matched to the query image via a pretrained cross-modality feature matcher. Matched keypoints are back-projected to obtain 2D--3D correspondences for PnP-based pose estimation. Relying exclusively on geometry makes the method inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object. We evaluate on widely-used public benchmarks, reporting state-of-the-art results on texture-less objects without object-specific training, and introduce a novel dataset with assembly defects, texture variations, and occlusion to demonstrate real-world applicability."},{"descriptionType":"Other","description":"This work has been accepted for publication in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026. The final published version will be available via IEEE Xplore"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.16015","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-20T02:15:13Z","registered":"2026-07-20T02:15:14Z","published":null,"updated":"2026-07-21T03:56:09Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.15659","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.15659","identifiers":[{"identifier":"2607.15659","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","affiliation":["Faculty of Engineering, Monash University, Clayton, VIC 3800, Australia"],"givenName":"Junlong","familyName":"Xiao","name":"Xiao, Junlong","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Faculty of Engineering, Monash University, Clayton, VIC 3800, Australia"],"givenName":"Yaoqiang","familyName":"Pan","name":"Pan, Yaoqiang","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Faculty of Engineering, Monash University, Clayton, VIC 3800, Australia"],"givenName":"Xuan","familyName":"Zhang","name":"Zhang, Xuan","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Michael Yu Wang. School of Engineering, Great Bay University, Songshan Lake, Dongguan, Guangdong, China"],"givenName":"Michael Yu","familyName":"Wang","name":"Wang, Michael Yu","nameIdentifiers":[]},{"nameType":"Personal","affiliation":["Faculty of Engineering, Monash University, Clayton, VIC 3800, Australia"],"givenName":"Chao","familyName":"Chen","name":"Chen, Chao","nameIdentifiers":[]}],"titles":[{"title":"Continuously Stable Structure through Plastic Deformation"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-17T06:10:36Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-20T00:23:05Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-20T04:06:44Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:14:12Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"Soft robots have seen widespread adoption in interactive tasks due to their inherent compliance and adaptability. However, these advantages often come at the cost of stability, posing challenges in a dynamic environment. This limitation is especially critical in soft grippers, where instability under acceleration or external disturbances can result in grasp failure. In this study, we present a continuously stable structure through plastic deformation (CSSPD), integrated into a soft gripper. By leveraging the mechanism of plastic deformation, the gripper maintains continuous configurations without energy input, while the added stiffness ensures both static and dynamic stability. We introduce a bioinspired paw pad that significantly enhances stability and enables sensing-based rapid object grasping. Then we develop the mathematical model and optimize the kirigami structure of the metal layer. Experimental results show that the gripper can sustain a passive holding force of up to 16 N without energy input, achieving performance comparable to pneumatic actuation at 0.3 MPa. When combined with pneumatic actuation, it remains stable under pulsed accelerations of up to 400 m/s^2. It can also passively perch on tree branches for extended periods without power, demonstrating promise for mobile robotic applications."},{"descriptionType":"Other","description":"24 pages, 8 figures"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.15659","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-20T02:07:32Z","registered":"2026-07-20T02:07:33Z","published":null,"updated":"2026-07-21T03:55:55Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.14688","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.14688","identifiers":[{"identifier":"2607.14688","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Mainak","familyName":"Mondal","name":"Mondal, Mainak","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yihang","familyName":"Feng","name":"Feng, Yihang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yangchao","familyName":"Luo","name":"Luo, Yangchao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Song","familyName":"Han","name":"Han, Song","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"MIND-CAVs: Multi-Intelligence Negotiation and Decision System for CAVs based on Intent-Driven Autonomy"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-16T07:46:34Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-17T00:31:51Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T03:23:55Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:18:56Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"Modern autonomous vehicles largely operate as isolated agents: they rely on on-board perception and decision modules and broadcast Basic Safety Messages (BSMs) that expose only low-level kinematic state. While existing cooperative driving frameworks enable limited sensor sharing, they rarely communicate high-level maneuver intentions, and edge computing is primarily used for content delivery rather than decision arbitration. As a result, current connected autonomy lacks a principled mechanism for making globally consistent, intent-aware coordination decisions across vehicles. To address this gap, we propose MIND-CAVs, a Multi-Intelligence Negotiation and Decision framework for connected autonomous vehicles (CAVs) based on intent-driven autonomy. Each vehicle abstracts raw sensor observations into structured intent representations, exchanges them over V2X links, and receives globally consistent coordination plans from roadside edge servers. Edge agents combine learned and rule-based arbitration mechanisms to negotiate conflicting intents among vehicles, while a cloud platform records decisions for auditing and continual retraining. We implement MIND-CAVs in a CARLA-based AI-in-the-loop platform and evaluate it in multi-lane highway scenarios involving conflicting maneuvers and route-constrained exits. Experimental results show improved maneuver completion time and reduced unsafe proximity and unnecessary braking compared with isolated autonomy, first-come-first-served arbitration, and multi-agent reinforcement learning baselines."},{"descriptionType":"Other","description":"8 pages, 5 figures, 1 table"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.14688","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-17T02:17:29Z","registered":"2026-07-17T02:17:30Z","published":null,"updated":"2026-07-21T03:55:37Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.14183","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.14183","identifiers":[{"identifier":"2607.14183","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Zishuo","familyName":"Li","name":"Li, Zishuo","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Bowen","familyName":"Yang","name":"Yang, Bowen","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Changtao","familyName":"Miao","name":"Miao, Changtao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Kai","familyName":"Zhu","name":"Zhu, Kai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hao","familyName":"Chen","name":"Chen, Hao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Qingze","familyName":"Guan","name":"Guan, Qingze","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhengxing","familyName":"Wu","name":"Wu, Zhengxing","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Wanke","familyName":"Zhan","name":"Zhan, Wanke","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yang","familyName":"Sun","name":"Sun, Yang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhiyi","familyName":"Huang","name":"Huang, Zhiyi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zitong","familyName":"Shan","name":"Shan, Zitong","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhenchao","familyName":"Jin","name":"Jin, Zhenchao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiadong","familyName":"Hong","name":"Hong, Jiadong","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Taowen","familyName":"Wang","name":"Wang, Taowen","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yushi","familyName":"Feng","name":"Feng, Yushi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"You","familyName":"Liu","name":"Liu, You","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yibo","familyName":"Wang","name":"Wang, Yibo","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yifan","familyName":"Yang","name":"Yang, Yifan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhaowen","familyName":"Zhou","name":"Zhou, Zhaowen","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Man","familyName":"Luo","name":"Luo, Man","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hao","familyName":"Cheng","name":"Cheng, Hao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Bo","familyName":"Zhang","name":"Zhang, Bo","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jianshu","familyName":"Li","name":"Li, Jianshu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiansheng","familyName":"Cai","name":"Cai, Jiansheng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Guocai","familyName":"Yao","name":"Yao, Guocai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jize","familyName":"Zhang","name":"Zhang, Jize","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Chenhao","familyName":"Lin","name":"Lin, Chenhao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Renjing","familyName":"Xu","name":"Xu, Renjing","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Lequan","familyName":"Yu","name":"Yu, Lequan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Chao","familyName":"Shen","name":"Shen, Chao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Chunhua","familyName":"Shen","name":"Shen, Chunhua","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhe","familyName":"Li","name":"Li, Zhe","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Computer Vision and Pattern Recognition (cs.CV)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-15T14:49:49Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-17T00:01:53Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T10:02:02Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:28:11Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.14183","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-17T02:06:35Z","registered":"2026-07-17T02:06:36Z","published":null,"updated":"2026-07-21T03:55:24Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.13597","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.13597","identifiers":[{"identifier":"2607.13597","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Yuan","familyName":"Xu","name":"Xu, Yuan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Youheng","familyName":"Shi","name":"Shi, Youheng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Chengyang","familyName":"Li","name":"Li, Chengyang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Wentao","familyName":"Zhu","name":"Zhu, Wentao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yizhou","familyName":"Wang","name":"Wang, Yizhou","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Semantic Anchoring for Robotic Action Representations"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Artificial Intelligence (cs.AI)","lang":"en","subjectScheme":"arXiv"},{"subject":"Computer Vision and Pattern Recognition (cs.CV)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-15T08:45:15Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-16T00:32:49Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T04:17:04Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:20:26Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a good action representation? Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structure captured by pretrained encoders. Systematic probing confirms that this structure erodes during finetuning, and that its quality synchronizes with both task success and out-of-distribution generalization. We further introduce a plug-and-play method that anchors action representations to a semantic manifold while decomposing representations into a shared semantic channel and a private channel, all discarded at inference, leaving the deployed model unchanged. Validated on different VLA backbones across simulation and real-world benchmarks, our method yields up to +18.7% on real-world in-distribution tasks and +21.5% on out-of-distribution generalization."},{"descriptionType":"Other","description":"Project Page: https://xy02-05.github.io/SemanticMN"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.13597","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-16T02:19:03Z","registered":"2026-07-16T02:19:04Z","published":null,"updated":"2026-07-21T03:55:06Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.12659","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.12659","identifiers":[{"identifier":"2607.12659","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Zebin","familyName":"Yang","name":"Yang, Zebin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Qi","familyName":"Wang","name":"Wang, Qi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yunhe","familyName":"Wang","name":"Wang, Yunhe","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xiurui","familyName":"Guo","name":"Guo, Xiurui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Bo","familyName":"Yu","name":"Yu, Bo","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Shaoshan","familyName":"Liu","name":"Liu, Shaoshan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiafeng","familyName":"Xu","name":"Xu, Jiafeng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hao","familyName":"Dong","name":"Dong, Hao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Meng","familyName":"Li","name":"Li, Meng","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Artificial Intelligence (cs.AI)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-14T11:38:36Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-15T00:45:30Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-17T02:24:13Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-20T00:15:25Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07-20T13:00:07Z","dateType":"Submitted","dateInformation":"v3"},{"date":"2026-07-21T01:35:38Z","dateType":"Updated","dateInformation":"v3"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"3","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational complexity, which leads to substantial inference latency and low control frequency. Asynchronous inference can partially mask this latency by parallelizing action execution and subsequent inference, but it introduces two critical issues: perception-execution misalignment and long reaction time. In this paper, we propose Jetson-PI, a method for efficient VLA deployment on onboard devices via Foresight-Aligned Asynchronous Correction. To address misalignment, we train a lightweight future correction module that predicts future environment representation conditioned on committed actions, enabling the action expert to directly predict actions from the future time step. To reduce reaction time, we introduce confidence-based scheduling optimization that adaptively balances VLM and action expert invocations, complemented by system-level accelerations including CUDA graph reuse, GPU-resident intermediate buffering, and flow unrolling. Extensive experiments demonstrate that Jetson-PI achieves 8.66x and 5.41x improvements in control frequency compared with naive PyTorch and vla.cpp on NVIDIA Jetson Orin, while outperforming VLASH by 14.8\\% in average success rate on the LIBERO benchmark. The code of our asynchronous algorithm is available on https://github.com/PKU-SEC-Lab/Jetson-PI, and our efficient llama.cpp-based inference engine is available on https://github.com/PKU-SEC-Lab/Jetson-PI-Edge."},{"descriptionType":"Other","description":"16 pages, 10 figures"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.12659","contentUrl":null,"metadataVersion":2,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-15T02:06:12Z","registered":"2026-07-15T02:06:12Z","published":null,"updated":"2026-07-21T03:54:52Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.12115","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.12115","identifiers":[{"identifier":"2607.12115","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Christopher A.","familyName":"Tucker","name":"Tucker, Christopher A.","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"A Behavioral State Vocabulary in Sony ERS-111 R-CODE"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-13T19:42:54Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-15T00:08:38Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-17T19:02:53Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:07:05Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"This paper presents a corpus-level analysis of generated behavior diagrams derived from Sony's R-CODE sample distribution for the ERS-111 AIBO. Rather than reading each script in isolation, the study compares named states across the corpus to identify the recurring control vocabulary that structures the sample set. The resulting aggregate shows that many superficially different routines are built from a compact embodied grammar centered on initialization, sensing, iterative action, synchronization, and recovery. It further shows that this vocabulary supports a graded scale of rising behavioral complexity, from capability activation and startup regularization to monitored locomotion, environmental decision loops, and fuller mode-based control. In addition to historical analysis, the paper argues that this form of state-based abstraction is useful as an intermediate representation for constructing new encapsulated behavior routines, especially on constrained native robotic systems where deterministic control, direct hardware access, and modular behavioral composition remain important."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.12115","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-15T01:56:28Z","registered":"2026-07-15T01:56:29Z","published":null,"updated":"2026-07-21T03:54:44Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.10879","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.10879","identifiers":[{"identifier":"2607.10879","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Siyi","familyName":"Hu","name":"Hu, Siyi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jared","familyName":"Strader","name":"Strader, Jared","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hyungtae","familyName":"Lim","name":"Lim, Hyungtae","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Luca","familyName":"Carlone","name":"Carlone, Luca","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"3D Scene Graph Prediction: Generating Hierarchical Models from Partially Observed Environments"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Computer Vision and Pattern Recognition (cs.CV)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-12T18:52:05Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-14T00:57:19Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-17T19:41:25Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:08:13Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"Generating realistic 3D indoor scenes is an area of growing interest in computer vision and robotics. Existing methods, often motivated by applications such as interior design, generally focus on object layout generation within a single room. The generation of high-level scene structure, such as room-level layout and traversability, remains underexplored despite its importance for robotics applications. In this paper, we consider the case where a robot has explored part of an environment and needs to predict the unexplored parts to support downstream tasks such as exploration or object search. We propose a top-down framework for synthesizing hierarchical 3D scene graphs, including a room layer -- describing the floor plan and traversability -- and an object layer modeling object layouts within each room. For the room layer, we propose a novel mixed-domain graph diffusion model jointly predicting room categories, floor boundaries, and traversability between rooms. Via corruption and masking, this model supports partial constraints such as incomplete floor plans, avoiding the need for partially observed training data. For the object layer, we integrate an existing mixed discrete-continuous diffusion model for joint prediction of object categories, locations, sizes, and orientations within each room given the floor plan. We compare our method with state-of-the-art occupancy-based and LLM-based floor plan generation methods on a standard benchmark. Compared with an occupancy-based learning baseline, our method generalizes substantially better to out-of-distribution partial floor plans. We also demonstrate our integrated prediction pipeline on real-world scenes from robot-collected data, enabling prediction beyond explored areas."},{"descriptionType":"Other","description":"Accepted at IROS 2026. Main paper: 8 pages, 3 figures, 3 tables. Includes a supplementary appendix"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.10879","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-14T03:33:55Z","registered":"2026-07-14T03:33:55Z","published":null,"updated":"2026-07-21T03:54:31Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.10436","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.10436","identifiers":[{"identifier":"2607.10436","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Alex","familyName":"Borisevich","name":"Borisevich, Alex","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Tracking Through Decoupling Singularities: A Singularity-Robust Homotopy-Continuation Extension of Feedback Linearization"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Systems and Control (eess.SY)","lang":"en","subjectScheme":"arXiv"},{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Optimization and Control (math.OC)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Electrical engineering, electronic engineering, information engineering","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"FOS: Mathematics","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"G.1.5; G.1.7; I.2.9; J.2","lang":"en","subjectScheme":"ACM"},{"subject":"Primary 93C10, Secondary 93C15, 93B52, 93D15, 65H20, 65H10","lang":"en","subjectScheme":"MSC"}],"contributors":[],"dates":[{"date":"2026-07-11T18:41:04Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-14T00:35:38Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T21:46:41Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:42:44Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Input--output feedback linearization fails at decoupling singularities, where the decoupling matrix loses rank, the relative degree is lost, and the linearizing control becomes unbounded. This paper develops a singularity-robust trajectory-tracking controller for square nonlinear control-affine systems that tracks through isolated decoupling singularities with bounded control. The method recasts tracking as real-time arc-length homotopy continuation, equivalently a continuous-time Newton/Davidenko flow, and replaces the inverse decoupling matrix by the least-norm Moore--Penrose solution of an augmented matrix $A=[Λ\\mid b]$, where $b$ is the homotopy direction. A transversality condition $w^T b \\ne 0$, with $w$ in the left null space of the decoupling matrix, keeps the augmented matrix full row rank through a generic rank-one loss. The resulting flow agrees with feedback linearization away from the singular set, tracks with $O(1/k)$ error, and re-locks after each crossing. The theory also characterizes the reflection-versus-branch-crossing dichotomy at Whitney folds and relates the reflection case to a Filippov sliding mode. Extensions cover dynamic relative-degree-one minimum-phase systems and arbitrary relative degree via filtered-error reduction. Simulations include a redundant 2-DOF manipulator, relative-degree-one and relative-degree-two plants, and a dual-active-bridge series-resonant DC/DC converter, where the method performs bounded inversion across buck/boost and resonance singularities while preserving zero-voltage soft switching."},{"descriptionType":"Other","description":"Python code to reproduce all numerical results is included as ancillary files"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.10436","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-14T03:25:47Z","registered":"2026-07-14T03:25:47Z","published":null,"updated":"2026-07-21T03:54:26Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.03920","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.03920","identifiers":[{"identifier":"2607.03920","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Rufeng","familyName":"Chen","name":"Chen, Rufeng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yue","familyName":"Chang","name":"Chang, Yue","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zili","familyName":"Shao","name":"Shao, Zili","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhaofan","familyName":"Zhang","name":"Zhang, Zhaofan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Li","familyName":"Chen","name":"Chen, Li","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hechang","familyName":"Chen","name":"Chen, Hechang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hui","familyName":"Xiong","name":"Xiong, Hui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Sihong","familyName":"Xie","name":"Xie, Sihong","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-04T15:25:25Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-07T01:18:29Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-20T09:22:48Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:25:12Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, a benchmark for Long-Horizon Audio-Visual-Language Navigation that combines multi-goal mission execution, heterogeneous goal specifications, and persistent spatialized acoustic cues. In LH-AVLN, an agent receives a global mission of two to four goals specified by category, language description, or reference image, and navigates with RGB-D observations, pose, and binaural audio in indoor 3D environments. The benchmark supports both ordered and unordered missions, where alternating goal-associated sounds can guide non-line-of-sight search but may also become distractors as mission progress changes. We further develop PAG-Nav, a training-free reference agent that maintains a temporal uniform semantic map and performs progressive goal-state planning, using sound for search while reserving completion for visual-semantic verification. Experiments show that existing vision-language, memory-based, and audio-visual agents struggle to complete full LH-AVLN missions, and that PAG-Nav provides a stronger diagnostic baseline while leaving substantial room for future progress."},{"descriptionType":"Other","description":"12 pages, 3 figures"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.03920","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-07T04:22:59Z","registered":"2026-07-07T04:23:00Z","published":null,"updated":"2026-07-21T03:53:43Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2607.03387","type":"dois","attributes":{"doi":"10.48550/arxiv.2607.03387","identifiers":[{"identifier":"2607.03387","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Pengwei","familyName":"Zhang","name":"Zhang, Pengwei","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Bin","familyName":"Xie","name":"Xie, Bin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xinpan","familyName":"Meng","name":"Meng, Xinpan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xinyu","familyName":"Guo","name":"Guo, Xinyu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Ce","familyName":"Hao","name":"Hao, Ce","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Fang","familyName":"Deng","name":"Deng, Fang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Long","familyName":"Cheng","name":"Cheng, Long","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Tiancai","familyName":"Wang","name":"Wang, Tiancai","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Feeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-07-03T14:42:30Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-07T00:49:25Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-07T03:40:42Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-08T00:19:52Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07-19T08:29:52Z","dateType":"Submitted","dateInformation":"v3"},{"date":"2026-07-21T00:52:44Z","dateType":"Updated","dateInformation":"v3"},{"date":"2026-07","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"3","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"Tactile perception is indispensable for contact-rich manipulation, yet integrating it into Vision-Language-Action (VLA) models often induces modality collapse, where high-bandwidth visual features overshadow sparse tactile cues. Inspired by Predictive Coding, a neural mechanism where the brain attenuates predictable inputs to prioritize surprising stimuli, we propose ResTacVLA. Rather than treating tactile data as raw input, we reformulate it as a Residual Tactile Representation capturing the discrepancy between visual priors and physical sensations. By filtering out visually predictable dynamics, this formulation transforms sparse tactile signals into dense, high-value information gain, thereby inherently resolving the bandwidth mismatch. These residuals are discretized through a Vector Quantized (VQ) bottleneck into Latent Contact Primitives that capture critical events missed by vision. Analogous to the neural surprise signal, we leverage the uncertainty of the visual prior to adaptively gate tactile integration, prioritizing residuals specifically during visually unreliable phases to explicitly prevent visual dominance. Experimental results show that ResTacVLA consistently outperforms all baselines on a diverse set of contact-rich manipulation tasks, while remaining robust to unexpected dynamic disturbances. Project page: https://awilekong.github.io/ResTacVLA/"},{"descriptionType":"Other","description":"8 pages, 6 figures, 3 tables. Accepted by IROS 2026, Project page: https://awilekong.github.io/ResTacVLA/"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2607.03387","contentUrl":null,"metadataVersion":2,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-07-07T04:13:10Z","registered":"2026-07-07T04:13:11Z","published":null,"updated":"2026-07-21T03:53:35Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.30362","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.30362","identifiers":[{"identifier":"2606.30362","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Xiao","familyName":"Chen","name":"Chen, Xiao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Weishuai","familyName":"Zeng","name":"Zeng, Weishuai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xiaojie","familyName":"Niu","name":"Niu, Xiaojie","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zirui","familyName":"Wang","name":"Wang, Zirui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jianan","familyName":"Li","name":"Li, Jianan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Huayi","familyName":"Wang","name":"Wang, Huayi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Furui","familyName":"Xu","name":"Xu, Furui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiahe","familyName":"Chen","name":"Chen, Jiahe","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Weixiang","familyName":"Zhong","name":"Zhong, Weixiang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Lihe","familyName":"Ding","name":"Ding, Lihe","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Kailin","familyName":"Li","name":"Li, Kailin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiangmiao","familyName":"Pang","name":"Pang, Jiangmiao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Tai","familyName":"Wang","name":"Wang, Tai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Tianfan","familyName":"Xue","name":"Xue, Tianfan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jingbo","familyName":"Wang","name":"Wang, Jingbo","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Artificial Intelligence (cs.AI)","lang":"en","subjectScheme":"arXiv"},{"subject":"Computer Vision and Pattern Recognition (cs.CV)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-29T14:27:27Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-30T02:01:39Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-19T19:31:30Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:05:49Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motions. As a result, they are vulnerable to environmental shifts and incapable of reactive whole-body coordination. Naively cascading them with generative motion planners fails to achieve true reactivity, as inevitable tracking discrepancies induce fatal cumulative exposure bias. To bridge this gap, we propose ReactiveBFM, a real-time closed-loop planning-control framework. At its core, we effectively mitigate exposure bias via a scheduled prefix sampling curriculum, forcing the generative planner to actively learn error-recovery behaviors from imperfect physical states rather than ground-truth trajectories. Systematically, to reconcile the severe latency mismatch between auto-regressive planning and high-frequency tracking, we introduce an asynchronous replanning mechanism. Combined with trajectory chunking to temporally ensemble spatial references, our system guarantees spatio-temporally fluid execution without physical jitter. Deployed on the Unitree G1 humanoid, ReactiveBFM demonstrates unprecedented physical agility across a vast repertoire of text-conditioned closed-loop motions. Notably, ReactiveBFM achieves zero-shot moving target reaching, showcasing intricate whole-body coordination and on-the-fly replanning. In sim-to-sim benchmarking under severe perturbations, ReactiveBFM achieves a 93.1% success rate, significantly outperforming cascaded open-loop baselines by 28.6%."},{"descriptionType":"Other","description":"Project page: https://xiao-chen.tech/reactivebfm/"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.30362","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-30T03:55:28Z","registered":"2026-06-30T03:55:28Z","published":null,"updated":"2026-07-21T03:53:12Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.28196","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.28196","identifiers":[{"identifier":"2606.28196","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Ha Thang Long","familyName":"Doan","name":"Doan, Ha Thang Long","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Hikaru","familyName":"Arita","name":"Arita, Hikaru","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Kazuto","familyName":"Nakashima","name":"Nakashima, Kazuto","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Kenji","familyName":"Tahara","name":"Tahara, Kenji","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Learning Stable In-Grasp Manipulation in a Non-Dropping Action Space"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-26T15:43:55Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-29T00:54:46Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T14:08:53Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:33:12Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Traditionally, dexterous manipulation controllers are designed using analytic models constrained by strong assumptions about the hand and the objects being manipulated. Reinforcement learning (RL) has become another common approach in which skills are explored openly in an end-to-end manner but is inefficient because of unnoticeable instability and conflicts in learning objectives. This paper attempts to efficiently explore stable and accurate manipulation skills by decomposing dexterous skills into multiple simpler/analyzable components. Each skill component is subsequently learned with constraints and guidance from classical physics and control theory. Our work shows that for stable grasp, in-grasp reposition/reorientation with different objects, sensor/motor noise, latency, and frictional conditions, skill learning becomes efficient and stable with prior knowledge from theory."},{"descriptionType":"Other","description":"This work has been submitted to the Taylor \u0026amp; Francis for possible publication"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.28196","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-29T01:53:04Z","registered":"2026-06-29T01:53:05Z","published":null,"updated":"2026-07-21T03:53:06Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.27163","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.27163","identifiers":[{"identifier":"2606.27163","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Ilia","familyName":"Larchenko","name":"Larchenko, Ilia","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Artificial Intelligence (cs.AI)","lang":"en","subjectScheme":"arXiv"},{"subject":"Machine Learning (cs.LG)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-25T15:31:23Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-26T00:57:00Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T14:45:34Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:34:24Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a sim-to-real recipe with camera-alignment tooling, heavy augmentation and DAgger-like HIL data collection."},{"descriptionType":"Other","description":"Solution of the LeHome Challenge at ICRA 2026"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.27163","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-26T02:07:59Z","registered":"2026-06-26T02:08:00Z","published":null,"updated":"2026-07-21T03:53:03Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.22881","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.22881","identifiers":[{"identifier":"2606.22881","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Param","familyName":"Patel","name":"Patel, Param","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jay","familyName":"Dave","name":"Dave, Jay","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Pratyush","familyName":"Chakraborty","name":"Chakraborty, Pratyush","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"A Vendor-Agnostic LiDAR Data Conversion System with Multi-Signal Detection and Multi-Format Output"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Signal Processing (eess.SP)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"FOS: Electrical engineering, electronic engineering, information engineering","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-22T05:46:41Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-23T02:15:27Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-20T08:22:57Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:22:28Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"LiDAR (Light Detection and Ranging) sensors capture the surrounding environment as dense 3D point clouds by measuring the time-of-flight of emitted laser pulses, making them foundational across autonomous vehicles, robotics, and large-scale mapping. PCAP (Packet Capture) files from these sensors are the starting point of most 3D perception pipelines, yet internal packet structures, UDP (User Datagram Protocol) port conventions and encoding schemes differ enough across manufacturers that no single tool reads them all. Ouster, Velodyne, Hesai, and Livox each require their own SDK (Software Development Kit), their own environment setup, and their own conversion workflow. Supporting all four means maintaining four disconnected pipelines with no shared infrastructure. The pipeline described here takes a raw PCAP as input and handles vendor identification automatically, scoring six independent file characteristics through a weighted multi-signal approach to determine the source sensor. C++ SDKs handle Ouster and Velodyne, while Hesai and Livox rely on Python-based dpkt parsing where no open source SDK exists. From there, a single command writes output to any of five industry-standard formats. We tested on real outdoor captures. Ouster peaks at 2.08M points per second, Velodyne at 1.47M, both running through native C++ packet decoding. Hesai and Livox land at 110K and 150K respectively, where Python-layer parsing introduces overhead that compounds under sustained load. The 8-10x gap held consistently across runs. Tested on a consumer-grade i3 with 8GB RAM, no vendor configuration required"},{"descriptionType":"Other","description":"Manuscript submitted at Software: Practice and Experience"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.22881","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-23T05:05:24Z","registered":"2026-06-23T05:05:25Z","published":null,"updated":"2026-07-21T03:52:45Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.16467","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.16467","identifiers":[{"identifier":"2606.16467","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Alberto","familyName":"Giaretta","name":"Giaretta, Alberto","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"A Formal Resilience Framework for Cyber-Physical Embodied Systems under Device-Level Cyberattacks"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Cryptography and Security (cs.CR)","lang":"en","subjectScheme":"arXiv"},{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-15T09:35:04Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-16T01:34:07Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T00:20:04Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:15:01Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"In cyber-physical systems (CPSs), fault tolerance is traditionally achieved by analysing sensor and actuator outputs, detecting progressive drift or sudden failures, and initiating suitable tolerance mechanisms. Reasonable under general failure models, this approach fails to capture nuanced disruptions caused by cyberattacks, which may employ subtle strategies. This is particularly critical in embodied CPSs, where computational and physical devices not only have an active role in task completion, but also in embodiment preservation (that is, maintaining the system's physical integrity). To prevent structural physical damage, embodied CPSs require a framework that enables proactive response to cyberattacks. This paper proposes a formal dependability framework that incorporates IDS information into resilience evaluation predicates, enabling assessment of tolerance to disruption and degradation. The framework supports structured reasoning about how cyberattacks affect task execution and embodiment preservation, and whether mitigation strategies must be deployed. Analytical examples demonstrate its analytical capability and soundness, establishing a theoretical foundation for dependable and secure embodied CPSs."},{"descriptionType":"Other","description":"8 pages, 2 tables"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.16467","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-16T03:45:58Z","registered":"2026-06-16T03:45:59Z","published":null,"updated":"2026-07-21T03:52:25Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.14409","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.14409","identifiers":[{"identifier":"2606.14409","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"He","familyName":"Zhang","name":"Zhang, He","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Lingzhu","familyName":"Xiang","name":"Xiang, Lingzhu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Haitao","familyName":"Lin","name":"Lin, Haitao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zeyu","familyName":"Huang","name":"Huang, Zeyu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Minghui","familyName":"Wang","name":"Wang, Minghui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Dingyan","familyName":"Zhong","name":"Zhong, Dingyan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yubo","familyName":"Dong","name":"Dong, Yubo","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yihao","familyName":"Wu","name":"Wu, Yihao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yongming","familyName":"Rao","name":"Rao, Yongming","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Dongsheng","familyName":"Zhang","name":"Zhang, Dongsheng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Wanjia","familyName":"He","name":"He, Wanjia","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Ling","familyName":"Chen","name":"Chen, Ling","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Kai","familyName":"Huang","name":"Huang, Kai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiahao","familyName":"Chen","name":"Chen, Jiahao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Sichang","familyName":"Su","name":"Su, Sichang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xumin","familyName":"Yu","name":"Yu, Xumin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Ziyi","familyName":"Wang","name":"Wang, Ziyi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Chengwei","familyName":"Zhu","name":"Zhu, Chengwei","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xiao","familyName":"Teng","name":"Teng, Xiao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yuchun","familyName":"Guo","name":"Guo, Yuchun","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yufeng","familyName":"Zhang","name":"Zhang, Yufeng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yuandong","familyName":"Liu","name":"Liu, Yuandong","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Rui","familyName":"Wang","name":"Wang, Rui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zisheng","familyName":"Lu","name":"Lu, Zisheng","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Han","familyName":"Hu","name":"Hu, Han","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhengyou","familyName":"Zhang","name":"Zhang, Zhengyou","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Artificial Intelligence (cs.AI)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-12T12:45:18Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-15T00:47:10Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-20T07:32:26Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:20:04Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.14409","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-15T01:54:04Z","registered":"2026-06-15T01:54:05Z","published":null,"updated":"2026-07-21T03:52:17Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.12995","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.12995","identifiers":[{"identifier":"2606.12995","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Zhihai","familyName":"Bi","name":"Bi, Zhihai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Qiang","familyName":"Zhang","name":"Zhang, Qiang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Guoyang","familyName":"Zhao","name":"Zhao, Guoyang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiahang","familyName":"Cao","name":"Cao, Jiahang","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xueyin","familyName":"Luo","name":"Luo, Xueyin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yushan","familyName":"Zhang","name":"Zhang, Yushan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jinglan","familyName":"Xu","name":"Xu, Jinglan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Ruoyu","familyName":"Geng","name":"Geng, Ruoyu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yulin","familyName":"Li","name":"Li, Yulin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Andrew F.","familyName":"Luo","name":"Luo, Andrew F.","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jun","familyName":"Ma","name":"Ma, Jun","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-11T07:31:05Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-12T00:32:17Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-19T05:39:03Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:50:00Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \\textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.12995","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-12T02:01:28Z","registered":"2026-06-12T02:01:28Z","published":null,"updated":"2026-07-21T03:52:13Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.08015","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.08015","identifiers":[{"identifier":"2606.08015","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Ziqian","familyName":"Wang","name":"Wang, Ziqian","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yitian","familyName":"Liu","name":"Liu, Yitian","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xingjian","familyName":"Mao","name":"Mao, Xingjian","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Minqian","familyName":"Wang","name":"Wang, Minqian","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yao","familyName":"Mu","name":"Mu, Yao","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-06T07:10:25Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-09T00:25:27Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-19T14:50:06Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:00:31Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"We propose Q-Guided Value-Gradient Matching (Q-VGM), an off-policy reinforcement learning method for a central difficulty in fine-tuning flow-matching vision-language-action (VLA) policies: improving an expressive flow-matching action expert with a learned $Q$-function. Effective improvement must exploit the critic's first-order signal $\\nabla_A Q$, yet flow policies make this hard: backpropagating values through the multi-step denoising chain is unstable at VLA scale, and the tractable action likelihoods required by policy gradients are unavailable under iterative denoising. Existing value-based methods therefore backpropagate through the full chain, use the critic only for test-time selection or guidance, or distill critic-improved actions as terminal labels that never supervise the velocity field. Q-VGM instead casts policy improvement as optimal control over the denoising dynamics, where the optimal residual velocity is the gradient of a denoising-time value function: clean-action estimates improved by iterative $Q$-gradient ascent with keep-best selection are converted into residual velocity targets that directly supervise the velocity field -- no action likelihoods, no backpropagation through the denoising chain, fully offline on a fixed replay buffer, and no critic at inference time. The critic is an action-sensitive stepwise IQL critic on compact latent states from the frozen VLA backbone. This enables a few-shot-initialization, learn-from-experience paradigm: starting from a few-shot-SFT $π_{0.5}$ policy, Q-VGM improves the policy from its own rollouts without additional expert supervision, raising the average LIBERO success rate from 79.0% to 92.5%, outperforming all same-backbone, same-critic baselines, and attaining high success rates on four real-robot manipulation tasks, including fine-grained plug insertion."},{"descriptionType":"Other","description":"13 pages, 3 figures, 4 tables"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.08015","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-09T03:19:16Z","registered":"2026-06-09T03:19:17Z","published":null,"updated":"2026-07-21T03:51:57Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2606.06491","type":"dois","attributes":{"doi":"10.48550/arxiv.2606.06491","identifiers":[{"identifier":"2606.06491","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Dong","familyName":"Jing","name":"Jing, Dong","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jingchen","familyName":"Nie","name":"Nie, Jingchen","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Tianqi","familyName":"Zhang","name":"Zhang, Tianqi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiaqi","familyName":"Liu","name":"Liu, Jiaqi","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Huaxiu","familyName":"Yao","name":"Yao, Huaxiu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zhiwu","familyName":"Lu","name":"Lu, Zhiwu","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Mingyu","familyName":"Ding","name":"Ding, Mingyu","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Artificial Intelligence (cs.AI)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-06-04T17:59:40Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-06-05T01:14:35Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-19T18:14:26Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:04:49Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-06","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by merging or splitting actions while preserving its motion semantics. (2) A model-side conditioning mechanism that feeds the speed to the policy. Statistics show that VSTA reaches the requested speed with negligible motion error. Experiments in simulation and on real-world tasks demonstrate that TempoVLA achieves flexible speed control in both directions, while VSTA additionally boosts the default $1\\times$ performance via better data utilization. Furthermore, by cooperating with a large multimodal model, TempoVLA realizes dynamic speed control, accelerating through low-risk phases and decelerating for high-risk ones."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2606.06491","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-06-05T02:24:13Z","registered":"2026-06-05T02:24:14Z","published":null,"updated":"2026-07-21T03:51:52Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2605.28634","type":"dois","attributes":{"doi":"10.48550/arxiv.2605.28634","identifiers":[{"identifier":"2605.28634","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Yutai","familyName":"Li","name":"Li, Yutai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Shaohui","familyName":"Peng","name":"Peng, Shaohui","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Jiaming","familyName":"Guo","name":"Guo, Jiaming","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Di","familyName":"Huang","name":"Huang, Di","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zihao","familyName":"Zhang","name":"Zhang, Zihao","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yuxuan","familyName":"Guo","name":"Guo, Yuxuan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yunkai","familyName":"Gao","name":"Gao, Yunkai","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Siming","familyName":"Lan","name":"Lan, Siming","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Ling","familyName":"Li","name":"Li, Ling","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Xing","familyName":"Hu","name":"Hu, Xing","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yunji","familyName":"Chen","name":"Chen, Yunji","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-05-27T15:41:18Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-05-28T01:17:09Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T09:11:07Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:26:24Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-05","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Vision-Language-Action (VLA) models offer a promising paradigm for generalist robotic policies, yet their adaptation is hindered by data inefficiency and poor generalization. We argue that these bottlenecks stem from the prevailing Direct Instruction-to-Control Mapping, which forces models to memorize monolithic trajectories rather than reusable motion patterns, i.e., primitives. We propose PrimitiveVLA, a framework that shifts this paradigm toward a Primitive-Centric Disassemble \u0026amp; Assemble paradigm. Supported by a shared Multimodal Canonical Representation (MCR), PrimitiveVLA unifies two phases: (1) Fine-tuning-phase Disassembly, which uses an automated pipeline to disassemble demonstrations into reusable primitives; and (2) Inference-phase Assembly, which employs a VLM-based planner and an LLM-generated switch module for robust closed-loop execution. By disassembling tasks into reusable primitives, PrimitiveVLA enables VLA models to learn invariant motion patterns instead of task-specific trajectories. Extensive experiments show that our framework improves data efficiency and achieves superior zero-shot generalization across unseen and long-horizon tasks."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2605.28634","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-05-28T02:41:57Z","registered":"2026-05-28T02:41:58Z","published":null,"updated":"2026-07-21T03:51:26Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2605.20209","type":"dois","attributes":{"doi":"10.48550/arxiv.2605.20209","identifiers":[{"identifier":"2605.20209","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Chia-Wen","familyName":"Chen","name":"Chen, Chia-Wen","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Yan","familyName":"Wu","name":"Wu, Yan","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Korrawe","familyName":"Karunratanakul","name":"Karunratanakul, Korrawe","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Siyu","familyName":"Tang","name":"Tang, Siyu","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Graphics (cs.GR)","lang":"en","subjectScheme":"arXiv"},{"subject":"Machine Learning (cs.LG)","lang":"en","subjectScheme":"arXiv"},{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-04-15T14:51:32Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-05-21T00:00:37Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-18T16:31:33Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T00:37:14Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-05","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsUri":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","rights":"arXiv.org perpetual, non-exclusive license"}],"descriptions":[{"descriptionType":"Abstract","description":"Achieving precise, versatile whole-body character control in physics-based animation remains challenging. Recent diffusion-based policies generate rich and expressive motions but typically rely on gradient-based test-time guidance to satisfy task objectives, which is slow and can reduce robustness. We introduce NaP-Control (Navigating Diffusion Prior for Versatile and Fast Character Control), abbreviated as NaP. Our method uses reinforcement learning to manipulate the latent noise of a task-agnostic diffusion policy prior, steering it toward task-specific behaviors for fast, robust control with high motion fidelity. In contrast to methods that rely solely on offline training, NaP interacts with the environment during training to correct motions and optimize task rewards, improving success rates and enabling adaptation to challenging scenarios. By directly predicting task-optimized diffusion noise, NaP eliminates iterative guidance during denoising and enables efficient inference. Experiments show that NaP attains higher success rates and faster inference while preserving natural motion across diverse tasks."},{"descriptionType":"Other","description":"ECCV 2026. Project page: https://chiawenchen.github.io/nap-control-project/"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2605.20209","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-05-21T01:57:36Z","registered":"2026-05-21T01:57:36Z","published":null,"updated":"2026-07-21T03:51:09Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2605.12071","type":"dois","attributes":{"doi":"10.48550/arxiv.2605.12071","identifiers":[{"identifier":"2605.12071","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Ali Sidar","familyName":"Yilmaz","name":"Yilmaz, Ali Sidar","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Buday","familyName":"Turan","name":"Turan, Buday","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Lukas","familyName":"Pries","name":"Pries, Lukas","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Markus","familyName":"Ryll","name":"Ryll, Markus","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Control of Fully Actuated Aerial Vehicles: A Comparison of Model-based and Sensor-based Dynamic Inversion"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"Systems and Control (eess.SY)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"FOS: Electrical engineering, electronic engineering, information engineering","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-05-12T12:58:03Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-07-21T00:28:35Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-05","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[{"relationType":"IsVersionOf","relatedIdentifier":"10.1109/icuas69441.2026.11598683","relatedIdentifierType":"DOI"}],"relatedItems":[],"sizes":[],"formats":[],"version":"1","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution Share Alike 4.0 International","rightsIdentifier":"cc-by-sa-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Fully actuated multirotor platforms decouple translational force generation from vehicle attitude, enabling independent control of position and orientation and shifting performance limitations from attitude authority to actuator dynamics and control effectiveness. This paper compares a model-based nonlinear dynamic inversion controller (geometric NDI) with a sensor-based incremental dynamic inversion controller (INDI) on a fixed-tilt fully actuated hexarotor. Both controllers share an identical outer-loop structure and are both executed at 500 Hz; therefore, performance differences can be attributed primarily to the inversion strategy. Controller performance is evaluated in five experiments covering attitude step tracking under nominal conditions and under a 50% mismatch in the rotor force coefficient, hover disturbance rejection under an external lateral load, waypoint tracking in the presence of wind gust disturbances, reduced control frequency, and injected sensor degradation. The results show that INDI offers clear advantages under parameter mismatch, gust disturbances, and sensor degradation, and maintains lower position errors across the controller-frequency sweep. However, its advantages are not universal: geometric NDI yields better attitude tracking at reduced control frequencies. To the authors' best knowledge, this work presents the first experimental validation of a full pose tracking INDI controller with decoupled translational and rotational dynamics. These findings highlight the trade-off between measurement-based and model-based inversion for robust control and rapid deployment of fully actuated UAVs."}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2605.12071","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-05-13T02:40:30Z","registered":"2026-05-13T02:40:31Z","published":null,"updated":"2026-07-21T03:50:45Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2605.10457","type":"dois","attributes":{"doi":"10.48550/arxiv.2605.10457","identifiers":[{"identifier":"2605.10457","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Rabin","familyName":"Gajmer","name":"Gajmer, Rabin","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Joonas","familyName":"Haapala","name":"Haapala, Joonas","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Zoltan","familyName":"Beck","name":"Beck, Zoltan","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Geometrically Approximated Modeling for Emitter-Centric Ray-Triangle Filtering in Arbitrarily Dynamic LiDAR Simulation"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Graphics (cs.GR)","lang":"en","subjectScheme":"arXiv"},{"subject":"Performance (cs.PF)","lang":"en","subjectScheme":"arXiv"},{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"I.3.7; I.3.5; I.6.3","lang":"en","subjectScheme":"ACM"},{"subject":"68U05","lang":"en","subjectScheme":"MSC"}],"contributors":[],"dates":[{"date":"2026-05-11T12:28:49Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-05-12T02:08:02Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-07-20T17:17:45Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-07-21T01:50:19Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-05","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"2","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"Real-time Light Detection And Ranging (LiDAR) simulation must find, per emitted ray, the closest intersecting triangle even in dynamic scenes containing large numbers of moving and deformable objects. Dominant acceleration-structure approaches require rebuilding each frame for dynamic geometry -- a cost that compounds directly with scene dynamics and cannot be amortized regardless of how little actually changed.\n This paper presents the Gajmer Ray-Casting Algorithm (GRCA), which inverts the question: instead of asking what does each ray hit? it asks which rays can each triangle possibly hit? GRCA geometrically models spinning LiDAR emitters as rotation-traced cones or planes and uses each triangle's emitter-centric apparent area to cull, per triangle, which channels and the rays within those channels can possibly reach it -- without any acceleration structure. GRCA is compute-based and vendor-agnostic by design, targeting highly dynamic, high-resolution simultaneous multi-sensor simulation. At its core, GRCA is a general-purpose ray-casting algorithm: the emitter-centric inversion applies to any setting where rays originate from a known position, not only LiDAR.\n Benchmarks evaluate 2-8 simultaneous 128x4096-ray LiDARs (360deg/180deg) over complex dynamic scenes -- with just two sensors casting ~1M rays per frame. With range culling inactive, GRCA reaches up to 7.97x over hardware-accelerated OptiX (GPU) and 14.55x over Embree (CPU).\n Two independent extensions further boost performance even in the most complex scene (~22M triangles, ~9M of which are dynamic, 8 LiDARs): range culling at realistic deployment ranges (10-100m) reaches up to 7.02x GPU and 9.33x CPU; a hybrid pipeline -- GRCA for dynamic geometry, OptiX/Embree for static -- reaches up to 10.5x GPU and 19.2x CPU."},{"descriptionType":"Other","description":"21 pages, 20 figures"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2605.10457","contentUrl":null,"metadataVersion":1,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-05-12T04:38:55Z","registered":"2026-05-12T04:38:55Z","published":null,"updated":"2026-07-21T03:50:42Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}},{"id":"10.48550/arxiv.2605.01395","type":"dois","attributes":{"doi":"10.48550/arxiv.2605.01395","identifiers":[{"identifier":"2605.01395","identifierType":"arXiv"}],"creators":[{"nameType":"Personal","givenName":"Srishti","familyName":"Siddharth","name":"Siddharth, Srishti","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Domenico","familyName":"Campolo","name":"Campolo, Domenico","affiliation":[],"nameIdentifiers":[]},{"nameType":"Personal","givenName":"Ravi","familyName":"Banavar","name":"Banavar, Ravi","affiliation":[],"nameIdentifiers":[]}],"titles":[{"title":"Quasi-Static Control of Discrete Cosserat Rod"}],"publisher":"arXiv","container":{},"publicationYear":2026,"subjects":[{"subject":"Systems and Control (eess.SY)","lang":"en","subjectScheme":"arXiv"},{"subject":"Robotics (cs.RO)","lang":"en","subjectScheme":"arXiv"},{"subject":"FOS: Electrical engineering, electronic engineering, information engineering","subjectScheme":"Fields of Science and Technology (FOS)"},{"subject":"FOS: Computer and information sciences","subjectScheme":"Fields of Science and Technology (FOS)"}],"contributors":[],"dates":[{"date":"2026-05-02T11:35:02Z","dateType":"Submitted","dateInformation":"v1"},{"date":"2026-05-05T00:30:36Z","dateType":"Updated","dateInformation":"v1"},{"date":"2026-05-28T13:19:06Z","dateType":"Submitted","dateInformation":"v2"},{"date":"2026-05-29T01:04:26Z","dateType":"Updated","dateInformation":"v2"},{"date":"2026-07-04T11:44:31Z","dateType":"Submitted","dateInformation":"v3"},{"date":"2026-07-07T01:13:48Z","dateType":"Updated","dateInformation":"v3"},{"date":"2026-07-19T13:41:12Z","dateType":"Submitted","dateInformation":"v4"},{"date":"2026-07-21T00:59:04Z","dateType":"Updated","dateInformation":"v4"},{"date":"2026-05","dateType":"Available","dateInformation":"v1"}],"language":null,"types":{"schemaOrg":"CreativeWork","resourceTypeGeneral":"Preprint","citeproc":"article","bibtex":"misc","ris":"GEN","resourceType":"Article"},"relatedIdentifiers":[],"relatedItems":[],"sizes":[],"formats":[],"version":"4","rightsList":[{"rightsIdentifierScheme":"SPDX","rightsUri":"https://creativecommons.org/licenses/by/4.0/legalcode","schemeUri":"https://spdx.org/licenses/","rights":"Creative Commons Attribution 4.0 International","rightsIdentifier":"cc-by-4.0"}],"descriptions":[{"descriptionType":"Abstract","description":"In this paper, we design feedback control laws for soft robots modelled using the Cosserat rod theory, which is spatially discretised using the Piecewise Constant Strain (PCS) approach. The PCS approach approximates the nonlinear PDEs describing the Cosserat rod by a finite-dimensional system of nonlinear ODEs. This simplification results in a model describing soft robots which is similar to the serial rigid-link manipulators. We design feedback control laws for the quasi-static PCS model by using external wrenches as control inputs. The control laws are designed based on feedback linearisation in strain and task spaces. An extensive set of numerical results demonstrates the performance of the control laws for end-effector trajectory tracking and shape control of soft robots."},{"descriptionType":"Other","description":"Accepted to 17th APCA International Conference on Automatic Control and Soft Computing (CONTROLO 2026)"}],"geoLocations":[],"fundingReferences":[],"url":"https://arxiv.org/abs/2605.01395","contentUrl":null,"metadataVersion":3,"schemaVersion":"http://datacite.org/schema/kernel-4","source":"mds","isActive":true,"state":"findable","reason":null,"viewCount":0,"downloadCount":0,"referenceCount":0,"citationCount":0,"partCount":0,"partOfCount":0,"versionCount":0,"versionOfCount":0,"created":"2026-05-05T03:12:28Z","registered":"2026-05-05T03:12:29Z","published":null,"updated":"2026-07-21T03:50:22Z"},"relationships":{"client":{"data":{"id":"arxiv.content","type":"clients"}}}}],"meta":{"total":98670,"totalPages":400,"page":1},"links":{"self":"https://api.datacite.org/dois?query=subjects.subject%3Arobot%2A","next":"https://api.datacite.org/dois?page%5Bnumber%5D=2\u0026page%5Bsize%5D=25\u0026query=subjects.subject%3Arobot%2A"}}