Cross-modal perception driven compliance control system for robot with body

By implementing a compliant control system for embodied robots driven by cross-modal perception, the technical problems caused by multi-source data bias and multimodal semantic gaps are solved. The system eliminates bias through a data acquisition module, maps data through a supervised training module, generates collision-free trajectories through a decision-making module, optimizes strategies through a control strategy module, constructs virtual scenes through a verification module, and shares results through a learning module. This enables embodied robots to achieve accurate perception and natural interaction in complex environments.

CN121018510APending Publication Date: 2025-11-28CHANGCHUN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511341544.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In existing technologies, embodied robots suffer from problems such as inaccurate perception, incoordination between intent recognition and interactive control, difficulty in optimizing and verifying control strategies, slow adaptation to new tasks, and difficulty in sharing the learning results of multiple robots in complex environments due to the large discrepancies in multi-source data and multimodal semantic differences.

Method used

A compliant control system for an embodied robot driven by cross-modal perception is adopted. The system eliminates multi-source data bias through a data acquisition module, maps multimodal data to a shared feature space using a supervised training module, analyzes intent distribution through a decision module, generates collision-free trajectories through a control strategy module, and includes a verification and optimization module, a learning module, and a learning module. The system constructs a virtual interactive scenario through the data acquisition module and the verification and optimization module, and shares the robot's learning results through the learning module.

Benefits of technology

It enables accurate perception, natural interaction, rapid adaptation to new tasks in complex environments, and effective sharing of learning results from multiple robots, thereby improving the compliance and adaptability of embodied robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121018510A_ABST
    Figure CN121018510A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, in particular to a cross-modal perceptual driving body robot compliance control system which comprises the steps that a sensor is adopted to synchronously collect environment information, multi-source data bias is eliminated through a space-time alignment algorithm, a radial basis function neural network is adopted to analyze multi-modal fusion features, and a multi-modal model is obtained; human operation intention probability distribution is extracted to decompose a task into a path planning layer and a motion control layer, a collision-free trajectory is generated through an RRT algorithm, a high-fidelity physical engine is adopted to construct a virtual interaction scene, and robot learning results are shared through federal learning. According to the method, the problems of inaccurate perception, incoordination between intention recognition and interaction control, difficulty in control strategy verification, slow new task adaptation and difficulty in multi-robot learning result sharing caused by multi-source data deviation and large multi-modal semantic difference of the body robot in a complex environment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of robot control, and particularly relates to a cross-modal perception driven embodied robot compliant control system. BACKGROUND

[0002] With the rapid development of robot technology, embodied robots show great application potential in many fields such as industrial production, medical care, home service, etc. However, to realize efficient, safe and natural human interaction and cooperation of embodied robots in complex dynamic environments, compliant control becomes a key problem. In traditional robot control systems, data acquisition is often single, and environment information is mostly obtained by relying on a single type of sensor, which leads to insufficient comprehensiveness and accuracy of environment perception. Moreover, there are time and space deviations in multi-source data collected by different sensors, and the existing technology lacks effective processing means, so that the data fusion effect is not good, and the real environment condition cannot be accurately reflected. For the processing of multi-modal data, the previous method is difficult to effectively map different modal data such as vision and force sensation to a shared feature space, the semantic gap is large, the complementary information of each modal data cannot be fully utilized, and the understanding ability of robots to complex scenes is limited. At the same time, in terms of decision and control, the traditional system is difficult to accurately analyze multi-modal fusion features to identify human operation intention, and cannot dynamically adjust the robot dominance according to the interaction force data and intention recognition result, resulting in unnatural and smooth human-robot cooperation. In terms of task execution and control strategy, the existing technology is not fine enough in task decomposition, and the path planning and motion control lack effective cooperation, it is difficult to generate a collision-free optimal trajectory, and the correlation between impedance control parameters and contact force is not strong, which affects the compliance and adaptability of the robot.

[0003] There are also deficiencies in the verification and optimization of control strategy and learning mechanism. The actual scene test is costly and time-consuming, and lacks efficient virtual verification methods; the adaptation speed is slow when learning new tasks, and it is difficult to share learning achievements among multiple robots, which cannot quickly improve the overall performance. Therefore, it is of great practical significance to develop a cross-modal perception driven embodied robot compliant control system. SUMMARY

[0004] In view of the above status, the present application provides a cross-modal perception driven embodied robot compliant control system, which can solve the problems of inaccurate perception, incoordination of intention recognition and interaction control, and difficult control strategy verification, slow adaptation to new tasks, and difficult sharing of learning achievements among multiple robots of embodied robots in complex environments due to multi-source data deviation and large multi-modal semantic gap. To achieve the above purpose, the present application adopts the following technical solutions:

[0005] The cross-modal perception driven embodied robot compliant control system comprises a data acquisition module for synchronously acquiring environmental information by sensors, eliminating multi-source data deviation through a space-time alignment algorithm, extracting modal feature vectors, and dynamically fusing by using an attention mechanism; a supervised training module for constructing a vision-force joint encoder, mapping different modal data to a shared feature space, adopting contrastive learning to narrow the semantic gap, and combining reinforcement learning to optimize the policy network; a decision module for radial basis neural network analysis of multi-modal fusion features, extraction of human operation intention probability distribution, combination of intention recognition results and interaction force data, and dynamic adjustment of robot dominance through an arbitration mechanism; a control strategy module for decomposing a task into two layers of path planning and motion control, generating a collision-free trajectory through an RRT algorithm, adjusting curvature in combination with force feedback, and associating impedance control parameters with contact force; a verification and optimization module for constructing a virtual interaction scene by a high-fidelity physical engine, mapping real sensor data to a digital twin model in real time, and testing a control strategy through virtual-real synchronization; and a learning module for continuously accumulating interaction data through an experience replay mechanism, accelerating new task adaptation through meta-learning, and sharing robot learning results through federated learning.

[0006] Further, the data acquisition module comprises a sensing array submodule for constructing a heterogeneous perception array by sensors to synchronously acquire multi-dimensional environmental information of vision, force sense, touch, and hearing; a space-time synchronization submodule for high-precision timestamp synchronization and space coordinate system calibration algorithm to eliminate data deviation caused by differences in sampling frequency and installation position of different sensors, and to ensure space-time consistency; a feature extraction submodule for extracting feature vectors from original data, including edge contours of a visual scene, fluctuation spectrum of force sense signals, texture distribution of touch feedback, and frequency energy of a hearing environment; and a modal fusion submodule for inputting the feature vectors to an attention fusion network based on a Transformer architecture, dynamically allocating modal weights through a self-attention mechanism, suppressing noise interference, and strengthening complementary information.

[0007] Further, the supervised training module comprises a dual-mode encoding submodule for a vision and force sense dual-channel encoder architecture, extracting image texture features and force signal time sequence features through independent branches; a feature mapping submodule for inputting original feature vectors of two modalities to a shared projection layer, and adopting a nonlinear transformation to map them to a unified dimension latent feature space; a contrastive learning submodule for calculating the similarity of cross-modal feature pairs through a contrastive learning mechanism, adopting an InfoNCE loss function to narrow the feature distance of semantically consistent samples, and expanding the distribution interval of irrelevant samples; and a strategy optimization submodule for inputting fusion features to a PPO reinforcement learning framework, and optimizing policy network parameters in combination with environmental interaction reward signals.

[0008] Further, the decision module comprises: an intention modeling submodule for constructing an intention analysis model by a radial basis neural network, and extracting high-order semantic information in the multi-modal fusion features through nonlinear mapping; an intention prediction submodule for inputting visual, force sense and tactile feature vectors into a hidden layer, calculating sample similarity by using a Gaussian kernel function, and outputting a probability distribution matrix of human operation intention; a decision fusion submodule for aligning intention recognition results and real-time interaction force data in a time sequence dimension, and generating a comprehensive decision index through a weighted fusion algorithm; and an authority arbitration submodule for a dynamic arbitration mechanism, dividing control dominance levels according to index threshold values, and smoothly transferring robot control authority to a passive following mode when human operation stability exceeds a set value.

[0009] Further, the control strategy module comprises: a hierarchical control submodule for adopting a hierarchical architecture design for the control strategy module, and decomposing a complex task into a path planning and motion control double-layer execution unit; a path planning submodule for quickly searching a feasible path in a three-dimensional space by using an improved RRT algorithm, and obtaining a collision-free initial trajectory by using a collision detection module to real-time eliminate obstacle regions; a curvature adjustment submodule for inputting contact force data collected by a force sense sensor into a curvature optimizer, dynamically adjusting local curvature parameters of a trajectory by using a force-deformation model, and ensuring compliant contact; and an impedance adjustment submodule for establishing a nonlinear mapping relationship between impedance control parameters and real-time contact forces, and automatically increasing a virtual damping coefficient to suppress impact when a force mutation is detected, and generating a dynamic control instruction sequence that takes into account accuracy and safety.

[0010] Further, the verification optimization module comprises: a twin mapping submodule for building a virtual interaction scene consistent with real environment parameters by using a high-fidelity physical engine, and real-time mapping visual images, force sense signals and tactile feedback of a real robot to a digital twin model through a data interface; a virtual-real synchronization submodule for ensuring that a motion state of a virtual model and a physical entity are kept in millisecond-level synchronization by using virtual-real synchronization technology, and extracting key indexes of contact force distribution and trajectory deviation in a virtual scene; and a strategy verification submodule for performing a thousand-level pressure test of a control strategy in a virtual environment, obtaining quantitative scores of strategy stability and response speed by using a performance evaluation algorithm, and obtaining an optimized control parameter set that has passed working condition verification.

[0011] Further, the learning module comprises a data playback submodule for constructing a dynamic data pool by an experience replay mechanism, continuously storing the interaction data stream between the robot and the environment through a circular buffer, extracting key state-action pairs and reward signals; a fast adaptation submodule for adopting a fast adaptation algorithm based on model-agnostic meta-learning, extracting cross-task common features, and compressing the new task parameter initialization time to one-fifth of that of traditional methods; and a federated learning submodule for building a federated learning framework, encrypting and uploading the local model parameters of the robot to the central server, generating a global optimization model through a secure aggregation algorithm, and distributing the updated strategy to individual robots.

[0012] Further, the policy optimization submodule is configured to input the fusion features into a PPO reinforcement learning framework, and optimize the policy network parameters in combination with the environment interaction reward signals, including: adopting a fusion feature processing method to extract rich and comprehensive feature representations through multi-modal information fusion technology; inputting the processed fusion features into the PPO reinforcement learning framework, which constructs a policy network for interaction with the environment, and obtains reward signals according to the environment feedback during the interaction process; through continuous iteration and interaction, the reward signals are used to update and optimize the policy network parameters, so that the policy network gradually learns a better behavior strategy.

[0013] Further, the permission arbitration submodule is configured to adopt a dynamic arbitration mechanism to divide control dominance levels according to index thresholds, and to smoothly transfer the robot control permission to a passive following mode when the human operation stability exceeds a set value, including: extracting human operation stability related index data through real-time monitoring of the system, and dividing different control dominance levels according to preset index thresholds; continuously obtaining real-time values of the human operation stability during system operation, and comparing the real-time values with the set value; when it is detected that the human operation stability exceeds the set value, the robot control permission is smoothly transferred to the passive following mode according to the dynamic arbitration mechanism and according to the predetermined rules, so that the control right is smoothly transferred in the human-robot collaboration process.

[0014] In the technical scheme provided by the present application, the data acquisition module is used for synchronously collecting environmental information by sensors, eliminating multi-source data deviation through a space-time alignment algorithm, dynamically fusing after extracting each modal feature vector by using an attention mechanism; the supervised training module is used for constructing a visual-force sensation joint encoder, mapping different modal data to a shared feature space, narrowing the semantic gap by using contrast learning, and combining reinforcement learning to optimize the policy network; the decision module is used for analyzing multi-modal fusion features by a radial basis neural network, extracting a human operation intention probability distribution, combining the intention recognition result with interactive force data, and dynamically adjusting the robot dominance through an arbitration mechanism; the control strategy module is used for decomposing a task into two layers of path planning and motion control, generating a collision-free trajectory through an RRT algorithm, adjusting the curvature in combination with force sensation feedback, and associating impedance control parameters with contact force; the verification and optimization module is used for constructing a virtual interactive scene by a high-fidelity physical engine, mapping real sensor data to a digital twin model in real time, and testing the control strategy through virtual-real synchronization; and the learning module is used for continuously accumulating interactive data by an experience replay mechanism, accelerating new task adaptation by using meta-learning, and sharing robot learning achievements through federated learning. The present application solves the problems of inaccurate perception, uncoordinated intention recognition and interactive control, difficult control strategy verification, slow new task adaptation, and difficult sharing of multi-robot learning achievements of an embodied robot in a complex environment due to multi-source data deviation and large multi-modal semantic gap. BRIEF DESCRIPTION OF DRAWINGS

[0015] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments with reference made to the accompanying drawings. The drawings are for purposes of illustration only and are not considered as limiting the application.

[0016] Figure 1 A first embodiment schematic diagram of the embodied robot compliant control system driven by cross-modal perception in the embodiments of the present application.

[0017] Figure 2 A second embodiment schematic diagram of the embodied robot compliant control system driven by cross-modal perception in the embodiments of the present application.

[0018] Figure 3 A third embodiment schematic diagram of the embodied robot compliant control system driven by cross-modal perception in the embodiments of the present application.

[0019] Figure 4 A fourth embodiment schematic diagram of the embodied robot compliant control system driven by cross-modal perception in the embodiments of the present application. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0021] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0022] Cross-modal perception-driven compliant control systems for embodied robots, such as Figure 1 As shown, the system includes: a data acquisition module for synchronously acquiring environmental information via sensors, eliminating multi-source data bias through a spatiotemporal alignment algorithm, extracting feature vectors for each modality, and dynamically fusing them using an attention mechanism; a supervised training module for constructing a vision-force joint encoder, mapping different modal data to a shared feature space, using contrastive learning to reduce semantic gaps, and combining reinforcement learning to optimize the strategy network; a decision-making module for analyzing multimodal fusion features using a radial basis function neural network, extracting the probability distribution of human operational intentions, combining the intention recognition results with interaction force data, and dynamically adjusting the robot's dominance through an arbitration mechanism; a control strategy module for decomposing the task into two layers: path planning and motion control, generating collision-free trajectories through the RRT algorithm, adjusting curvature through force feedback, and associating impedance control parameters with contact forces; a verification and optimization module for constructing virtual interactive scenarios using a high-fidelity physics engine, mapping real sensor data to a digital twin model in real time, and testing control strategies through virtual-real synchronization; and a learning module for continuously accumulating interaction data through an experience playback mechanism, accelerating adaptation to new tasks through meta-learning, and sharing robot learning results through federated learning.

[0023] like Figure 2 As shown, in this embodiment, the sensor array submodule is used to construct a heterogeneous sensing array of sensors to simultaneously collect multi-dimensional environmental information from vision, force, touch, and hearing; the spatiotemporal synchronization submodule is used for high-precision timestamp synchronization and spatial coordinate system calibration algorithms to eliminate data deviations caused by differences in sampling frequencies and installation positions of different sensors, ensuring spatiotemporal consistency; the feature extraction submodule is used to extract feature vectors from the raw data, including the edge contours of the visual scene, the fluctuation spectrum of the force signal, the texture distribution of the tactile feedback, and the frequency domain energy of the auditory environment; the modal fusion submodule is used to input the feature vectors into an attention fusion network based on the Transformer architecture, dynamically allocate modal weights through a self-attention mechanism, suppress noise interference, and enhance complementary information.

[0024] A heterogeneous perception array is constructed by using sensors to synchronously collect multi-dimensional environmental information in vision, force, touch, and hearing. Through high-precision timestamp synchronization and spatial coordinate system calibration algorithms, data deviations caused by differences in sampling frequency and installation position of different sensors are eliminated, and time-space consistent data is obtained. Feature vectors in the original data are extracted, including visual scene edge contours, force signal fluctuation frequency spectrum, touch feedback texture distribution, and hearing environment frequency energy. The feature vectors are input into an attention fusion network based on a Transformer architecture, and the modal weights are dynamically allocated through a self-attention mechanism to obtain a result that suppresses noise interference and enhances complementary information.

[0025] As shown in Figure 3 In this embodiment, a dual-modal encoding submodule is used for a vision and force dual-channel encoder architecture to extract image texture features and force signal time sequence features through independent branches. A feature mapping submodule is used to input the original feature vectors of the two modalities into a shared projection layer and map them to a unified dimension latent feature space through a nonlinear transformation. A contrast learning submodule is used to calculate the similarity of cross-modal feature pairs through a contrast learning mechanism, and an InfoNCE loss function is used to reduce the feature distance of semantically consistent samples and expand the distribution interval of unrelated samples. A strategy optimization submodule is used to input the fused features into a PPO reinforcement learning framework to optimize the strategy network parameters in combination with environmental interaction reward signals.

[0026] A dual-modal encoding submodule is constructed using a vision and force dual-channel encoder architecture to extract image texture features and force signal time sequence features through independent branches. The original feature vectors of the two modalities are input into a shared projection layer to obtain a unified dimension latent feature space through a nonlinear transformation. Through a contrast learning mechanism, the similarity of cross-modal feature pairs is calculated, and an InfoNCE loss function is used to reduce the feature distance of semantically consistent samples to obtain the effect of expanding the distribution interval of unrelated samples. The fused features are input into a PPO reinforcement learning framework to optimize the strategy network parameters in combination with environmental interaction reward signals, and finally an optimized strategy is obtained.

[0027] As shown in Figure 4As shown, in the embodiment, the intention modeling submodule is configured to construct an intention analysis model using a radial basis neural network to extract high-order semantic information in the multi-modal fusion features through nonlinear mapping; the intention prediction submodule is configured to input the visual, force sense, and tactile feature vectors to the hidden layer, calculate the sample similarity using a kernel function, and output a probability distribution matrix of the human operation intention; the decision fusion submodule is configured to align the intention recognition result and the real-time interaction force data in the time sequence dimension, and generate a comprehensive decision index through a weighted fusion algorithm; and the permission arbitration submodule is configured to implement a dynamic arbitration mechanism, divide the control dominance level according to an index threshold, and smoothly transfer the robot control permission to the passive following mode when the human operation stability exceeds a set value.

[0028] The intention modeling submodule is configured to construct an intention analysis model using a radial basis neural network to extract high-order semantic information in the multi-modal fusion features through nonlinear mapping. The visual, force sense, and tactile feature vectors are input to the hidden layer of the intention prediction submodule, the sample similarity is calculated using a kernel function, and a probability distribution matrix of the human operation intention is obtained. The intention recognition result and the real-time interaction force data are aligned in the time sequence dimension, and a comprehensive decision index is generated through a weighted fusion algorithm. The permission arbitration submodule is configured to implement a dynamic arbitration mechanism to divide the control dominance level according to an index threshold, and smoothly transfer the robot control permission when the human operation stability exceeds a set value to obtain the passive following mode.

[0029] In the embodiment, the hierarchical control submodule is configured to control the strategy module to adopt a hierarchical architecture design to decompose a complex task into a path planning and motion control double-layer execution unit. The path planning submodule is configured to quickly search for a feasible path in a three-dimensional space using an improved RRT algorithm, and use a collision detection module to real-time eliminate obstacle regions to obtain a collision-free initial trajectory. The curvature adjustment submodule is configured to input the contact force data collected by the force sense sensor to a curvature optimizer, dynamically adjust the local curvature parameters of the trajectory through a force-deformation model to ensure compliant contact. The impedance adjustment submodule is configured to establish a nonlinear mapping relationship between the impedance control parameters and the real-time contact force, automatically increase the virtual damping coefficient to suppress impact when a force mutation is detected, and generate a dynamic control instruction sequence that takes into account the accuracy and safety.

[0030] The control strategy module adopts a hierarchical architecture design to divide the control sub-modules, which is used to decompose complex tasks into path planning and motion control double-layer execution units. The path planning sub-module quickly searches for a feasible path in three-dimensional space through an improved RRT algorithm, and uses a collision detection module to remove obstacle areas in real time to obtain a collision-free initial trajectory. The contact force data collected by the force sensor is input into the curvature optimizer of the curvature adjustment sub-module, and the local curvature parameters of the trajectory are dynamically adjusted through the force-deformation model to obtain a compliant contact effect. The impedance adjustment sub-module establishes a nonlinear mapping relationship between the impedance control parameters and the real-time contact force, automatically increases the virtual damping coefficient when a force mutation is detected, and generates a dynamic control instruction sequence that takes into account accuracy and safety.

[0031] In this embodiment, the twin mapping sub-module is used to build a virtual interactive scene consistent with the real environment parameters by using a high-fidelity physical engine, and real-time mapping of the visual image, force signal and tactile feedback of the real robot to the digital twin model is realized through a data interface. The virtual-real synchronization sub-module is used to ensure that the motion state of the virtual model is synchronized with the physical entity at a millisecond level by using virtual-real synchronization technology, and key indicators such as contact force distribution and trajectory deviation in the virtual scene are extracted. The strategy verification sub-module is used to perform a thousand-level stress test of the control strategy in the virtual environment, and quantitative scores of strategy stability and response speed are obtained through a performance evaluation algorithm to obtain an optimized control parameter set verified by working conditions.

[0032] The twin mapping sub-module adopts a high-fidelity physical engine to build a virtual interactive scene consistent with the real environment parameters, and real-time mapping of the visual image, force signal and tactile feedback of the real robot to the digital twin model is realized through a data interface. The virtual-real synchronization sub-module uses virtual-real synchronization technology to ensure that the motion state of the virtual model is synchronized with the physical entity at a millisecond level, and key indicators such as contact force distribution and trajectory deviation in the virtual scene are extracted. The strategy verification sub-module is used to perform a thousand-level stress test of the control strategy in the virtual environment, and quantitative scores of strategy stability and response speed are obtained through a performance evaluation algorithm to obtain an optimized control parameter set verified by working conditions.

[0033] In this embodiment, the data playback sub-module is used to build a dynamic data pool by using an experience replay mechanism, and the interactive data stream of the robot and the environment is continuously stored through a circular buffer to extract key state-action pairs and reward signals. The fast adaptation sub-module is used to extract cross-task common features by using a fast adaptation algorithm based on model-agnostic meta-learning, and the initialization time of new task parameters is compressed to one-fifth of that of traditional methods. The federated learning sub-module is used to build a federated learning framework, encrypt the local model parameters of the robot and upload them to the central server, generate a global optimization model through a secure aggregation algorithm, and then distribute the updated strategy to individual robots.

[0034] The data playback submodule adopts an experience playback mechanism to construct a dynamic data pool, continuously stores the interaction data stream between the robot and the environment through a circular buffer, and extracts key state-action pairs and reward signals. The fast adaptation submodule adopts a fast adaptation algorithm based on model-agnostic meta-learning, extracts cross-task common features, and compresses the new task parameter initialization time to one-fifth of that of traditional methods. The federated learning submodule builds a federated learning framework, encrypts the local model parameters of the robot and uploads them to the central server, generates a globally optimized model through a secure aggregation algorithm, distributes the updated strategy to individual robots, and finally obtains a cross-modal perception driven embodied robot compliant control system.

[0035] In this embodiment, a fusion feature processing method is adopted to extract rich and comprehensive feature representations through multi-modal information fusion technology. The fusion features obtained by processing are input into the PPO reinforcement learning framework, which builds a policy network for interaction with the environment. During the interaction process, reward signals are obtained based on environmental feedback. Through continuous iterative interaction, the policy network parameters are updated and optimized using the reward signals, allowing the policy network to gradually learn better behavior strategies.

[0036] A fusion feature processing method is adopted to comprehensively extract rich features from multi-source information such as vision, hearing, and touch through multi-modal information fusion technology, obtaining fusion feature representations. The processed fusion features are input into a pre-built PPO reinforcement learning framework, which builds a policy network to interact with the environment. During the interaction process, reward signals are accurately obtained based on actual environmental feedback. Subsequently, through continuous iterative interaction, the policy network parameters are updated and optimized using the obtained reward signals, ultimately allowing the policy network to gradually learn and obtain better behavior strategies.

[0037] In this embodiment, the system extracts human operation stability related index data through real-time monitoring, and divides different control dominance levels according to preset index thresholds. During system operation, real-time values of human operation stability are continuously obtained and compared with set values. When the human operation stability exceeds the set value, the robot control authority is smoothly transferred to the passive following mode according to the dynamic arbitration mechanism, allowing smooth transition of control authority in human-robot collaboration.

[0038] The real-time monitoring system is used to extract the human operation stability related index data through sensors and other related means. According to the pre-set index threshold, the data is divided into different control right levels. During the system operation, the real-time values of the human operation stability are continuously obtained and compared with the set values. Once the human operation stability exceeds the set value, the robot control right is smoothly transferred to the passive following mode according to the dynamic arbitration mechanism and the established rules, so as to realize the smooth transition of the control right in the human-robot cooperation process and guarantee the smoothness and safety of the cooperation.

[0039] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.

Claims

1. A compliant control system for an embodied robot driven by cross-modal perception, characterized in that, The cross-modal perception-driven compliant robot control system includes: The data acquisition module is used to synchronously collect environmental information from sensors. It eliminates the bias of multi-source data through a spatiotemporal alignment algorithm, extracts the feature vectors of each modality, and then dynamically fuses them using an attention mechanism. The supervised training module is used to construct a visual-mechanical joint encoder, which maps data from different modalities to a shared feature space, uses contrastive learning to reduce semantic gaps, and combines reinforcement learning to optimize the policy network. The decision module is used to analyze multimodal fusion features using radial basis neural networks, extract the probability distribution of human operation intentions, combine the intention recognition results with interaction force data, and dynamically adjust the robot's dominance through an arbitration mechanism. The control strategy module is used to decompose the task into two layers: path planning and motion control. It generates a collision-free trajectory through the RRT algorithm, adjusts the curvature by combining force feedback, and associates the impedance control parameters with the contact force. The verification and optimization module is used to build virtual interactive scenarios with a high-fidelity physics engine, map real sensor data to the digital twin model in real time, and test control strategies through virtual-real synchronization. The learning module is used to continuously accumulate interactive data through the experience playback mechanism, accelerates adaptation to new tasks through meta-learning, and shares the robot's learning results through federated learning.

2. The cross-modal perception-driven compliant robot control system according to claim 1, characterized in that, The data acquisition module includes: The sensor array submodule is used to construct a heterogeneous sensing array of sensors to simultaneously collect multi-dimensional environmental information from vision, force, touch, and hearing. The spatiotemporal synchronization submodule is used for high-precision timestamp synchronization and spatial coordinate system calibration algorithms to eliminate data deviations caused by differences in sampling frequencies and installation locations of different sensors, ensuring spatiotemporal consistency. The feature extraction submodule is used to extract feature vectors from the raw data, including the edge contours of the visual scene, the fluctuation spectrum of the force signal, the texture distribution of the tactile feedback, and the frequency domain energy of the auditory environment. The modality fusion submodule is used to input feature vectors into an attention fusion network based on the Transformer architecture. It dynamically allocates modality weights through a self-attention mechanism to suppress noise interference and enhance complementary information.

3. The compliant control system for a cross-modal perception-driven embodied robot according to claim 1, characterized in that, The supervised training module includes: The dual-mode coding submodule is used for the vision and force dual-channel encoder architecture, and extracts image texture features and force signal temporal features through independent branches; The feature mapping submodule is used to input the original feature vectors of the two modalities into the shared projection layer and map them to a latent feature space of uniform dimension using a nonlinear transformation; The contrastive learning submodule is used to calculate the similarity of cross-modal feature pairs using the contrastive learning mechanism. It employs the InfoNCE loss function to reduce the feature distance between semantically consistent samples and expand the distribution interval of irrelevant samples. The policy optimization submodule is used to input fused features into the PPO reinforcement learning framework and optimize the policy network parameters by combining environmental interaction reward signals.

4. The compliant control system for a cross-modal perception-driven embodied robot according to claim 1, characterized in that, The decision-making module includes: The intent modeling submodule is used to construct an intent analysis model using a radial basis function neural network, and extracts high-order semantic information from multimodal fusion features through nonlinear mapping; The intent prediction submodule is used to input visual, force, and tactile feature vectors into the hidden layer, calculate sample similarity using a Gaussian kernel function, and output the probability distribution matrix of human operation intent. The decision fusion submodule is used to align the intent recognition results with real-time interaction force data in the time series dimension, and generate comprehensive decision indicators through a weighted fusion algorithm; The authorization arbitration submodule is used for dynamic arbitration mechanism. It divides the control authority level according to the indicator threshold. When the stability of human operation exceeds the set value, the robot control authority is smoothly transferred to passive follow mode.

5. The compliant control system for a cross-modal perception-driven embodied robot according to claim 1, characterized in that, The control strategy module includes: The hierarchical control submodule is used to control the strategy module. The hierarchical architecture design breaks down complex tasks into two-layer execution units: path planning and motion control. The path planning submodule is used to quickly search for feasible paths in three-dimensional space using the improved RRT algorithm, and uses the collision detection module to remove obstacle areas in real time to obtain a collision-free initial trajectory. The curvature adjustment submodule is used to input the contact force data collected by the force sensor into the curvature optimizer, and dynamically adjust the local curvature parameters of the trajectory through the force-deformation model to ensure compliant contact. The impedance adjustment submodule is used to establish a nonlinear mapping relationship between impedance control parameters and real-time contact force. When a sudden force change is detected, the virtual damping coefficient is automatically increased to suppress the impact, and a dynamic control command sequence that balances accuracy and safety is generated.

6. The compliant control system for a cross-modal perception-driven embodied robot according to claim 1, characterized in that, The verification optimization module includes: The twin mapping submodule is used to build a virtual interactive scene with parameters consistent with the real environment using a high-fidelity physics engine. It maps the visual images, force signals, and tactile feedback of the real robot to the digital twin model in real time through a data interface. The virtual-real synchronization submodule is used to ensure that the motion state of the virtual model is synchronized with the physical entity at the millisecond level using virtual-real synchronization technology, and to extract key indicators such as contact force distribution and trajectory deviation in the virtual scene. The strategy verification submodule is used to perform thousands of stress tests on the control strategy in a virtual environment. It obtains quantitative scores of strategy stability and response speed through performance evaluation algorithms, and obtains an optimized control parameter set that has been verified under operating conditions.

7. The cross-modal perception-driven compliant robot control system according to claim 1, characterized in that, The learning module includes: The data playback submodule is used to build a dynamic data pool for the experience playback mechanism. It continuously stores the interaction data stream between the robot and the environment through a circular buffer and extracts key state-action pairs and reward signals. The fast adaptation submodule is used to extract common features across tasks by using a fast adaptation algorithm based on model-independent meta-learning, and to compress the initialization time of new task parameters to one-fifth of that of traditional methods. The federated learning submodule is used to build a federated learning framework. It encrypts and uploads the robot's local model parameters to the central server, generates a globally optimized model through a secure aggregation algorithm, and then distributes the updated strategy to individual robots.

8. The cross-modal perception-driven compliant robot control system according to claim 3, characterized in that, The policy optimization submodule is used to input the fused features into the PPO reinforcement learning framework and optimize the policy network parameters by combining environmental interaction reward signals, including: A fusion feature processing approach is adopted, which extracts rich and comprehensive feature representations through multimodal information fusion technology; The fused features obtained from the processing are input into the PPO reinforcement learning framework, which constructs a policy network to interact with the environment. During the interaction, reward signals are obtained based on the feedback from the environment. Through continuous iterative interaction, the reward signals are used to update and optimize the parameters of the policy network, enabling the policy network to gradually learn better behavioral policies.

9. The cross-modal perception-driven compliant robot control system according to claim 4, characterized in that, The aforementioned authority arbitration submodule is used for a dynamic arbitration mechanism. It classifies control dominance levels based on indicator thresholds. When human operational stability exceeds a set value, the robot's control authority is smoothly transferred to a passive follower mode, including: The system extracts data on human operational stability-related indicators through a real-time monitoring system and classifies different levels of control authority based on preset indicator thresholds. During system operation, real-time values ​​of human operational stability are continuously acquired and compared with set values. When the stability of human operation is detected to exceed the set value, the robot's control authority is smoothly transferred to the passive follow mode according to the established rules based on the dynamic arbitration mechanism, so as to ensure a smooth transition of control during human-machine collaboration.

Citation Information

Cited By

  • Remote operation data acquisition and data closed-loop system for intelligent humanoid robot with body

    CN121589818A

  • A path coding-based multi-modal trajectory prediction and path tracking control method

    CN122379576A