Robot System and Autonomous Navigation Method Based on Multimodal Perception Collaboration
Through the combination of multimodal perception module and flexible joint robotic arms, the lack of perception and control imbalance of traditional robot systems in complex environments is solved, and autonomous navigation with high adaptability and high reliability is achieved, suitable for industrial automation and dangerous environment inspections.
Patent Information
- Application Number
- CN202510445257.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional robot systems have limited range of motion during movement, lack adaptive feedback ability to collision or external pressure, and lack perception ability in complex environments, imbalance in rigid and flexible control of robotic arms and lag in response to dynamic scenes.
The multimodal perception module (binocular vision camera, pressure-sensitive haptic array, microphone array and gas sensor) is used to combine with the flexible joint robotic arm (composed of bionic joints, shape memory alloy drivers and air pressure feedback devices) to realize real-time multimodal data processing and action command generation. Through reinforcement learning, optimized path planning is supported, and self-healing coatings and multimodal interactions are supported.
It improves the robot's perception robustness and motion freedom in complex environments, achieves high adaptability and high reliability autonomous navigation, and is suitable for industrial automation and dangerous environment inspections.
Smart Images

Figure CN119952734B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent robots, and in particular, to a robot system and an autonomous navigation method based on multi-modal perception collaboration. Background Art
[0002] Currently, during the movement of robots, traditional rigid joints limit the range of motion and lack the ability of adaptive feedback to collisions or external pressures. Summary of the Invention
[0003] In view of this, this application provides a robot system and an autonomous navigation method based on multi-modal perception collaboration, which can improve the motion control ability of the robot.
[0004] A robot system based on multi-modal perception collaboration includes:
[0005] A main body frame, with a central processor integrated inside the main body frame;
[0006] A multi-modal perception module, including a binocular vision camera, a pressure-sensitive tactile array, a microphone array, and a gas sensor, for collecting environmental data;
[0007] A flexible joint robotic arm, composed of at least 3 series-connected bionic joints, with a shape memory alloy driver and a pneumatic feedback device built in each joint;
[0008] An edge computing unit, establishing communication connections with the central processor and the multi-modal perception module respectively, for real-time processing of perception data and generating action instructions.
[0009] In one embodiment, the pressure-sensitive tactile array is provided with a graphene film and a micro-current sensor.
[0010] In one embodiment, the driving mode of the bionic joint is a composite control of shape memory alloy and pneumatic muscle.
[0011] In one embodiment, the edge computing unit has a reinforcement learning model built in, and optimizes path planning through the Q-learning algorithm.
[0012] In one embodiment, the above-mentioned robot system further includes:
[0013] A human-computer interaction module, provided with voice command recognition, gesture control, and electroencephalogram signal interface.
[0014] In one embodiment, the gas sensor is used to detect the concentration of volatile organic compounds.
[0015] In one embodiment, the central processor supports the federated learning framework, allowing multiple robots to share local model parameters.
[0016] In one embodiment, the above-mentioned robot system further includes:
[0017] A self-healing coating covering the surface of the robotic arm, and the materials in the self-healing coating activate the repair function at a preset temperature.
[0018] In addition, an autonomous navigation method is also provided, which uses the above-mentioned robot system. The autonomous navigation method includes:
[0019] S1: Construct a three-dimensional environmental map through a multi-modal perception module and mark dynamic obstacles;
[0020] S2: Use the improved A* algorithm to calculate the initial path and optimize the path weight in combination with a reinforcement learning model;
[0021] S3: Adjust the movement speed and angle of the robotic arm according to the real-time stress feedback of the flexible joint.
[0022] In one embodiment, an obstacle movement prediction model is introduced for path weight optimization.
[0023] The above-mentioned robot system and autonomous navigation method based on multi-modal perception collaboration. The robot system includes: a main body frame with a central processor integrated inside; a multi-modal perception module including a binocular vision camera, a pressure-sensitive tactile array, a microphone array, and a gas sensor for collecting environmental data; a flexible joint robotic arm composed of at least 3 series-connected bionic joints, with a shape memory alloy driver and a pneumatic feedback device built in each joint; an edge computing unit that establishes communication connections with the central processor and the multi-modal perception module respectively, for real-time processing of perception data and generating action instructions. The central processor is integrated in the main body frame, supporting modular expansion and a compact design: the multi-modal sensors, drive unit, and computing module are highly integrated, reducing external wiring interference and enhancing the movement freedom of the robotic arm. Through the collaborative design of multi-modal perception fusion, bionic flexible drive, and edge-side real-time computing, the core problems of traditional robot systems such as insufficient complex environment perception ability, imbalance between rigid and flexible control of the robotic arm, and lag in dynamic scene response are solved, providing a highly adaptable and highly reliable solution for scenarios such as industrial automation and inspection in dangerous environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a structural schematic block diagram of a robot system based on multi-modal perception collaboration provided by an embodiment of the present application;
[0025] Figure 2 It is a structural schematic block diagram of a robot system based on multi-modal perception collaboration provided by another embodiment of the present application;
[0026] Figure 3Schematic flowchart of an autonomous navigation method provided by an embodiment of the present application.
[0027] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0028] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0029] In addition, if the description in the present application involves "first", "second", etc., it is only for descriptive purposes (such as to distinguish the same or similar elements), and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of the technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0030] As Figure 1 shown, a robot system based on multi-modal perception collaboration is provided. The robot system includes:
[0031] A main body frame 10, with a central processor 11 integrated inside the main body frame 10;
[0032] A multi-modal perception module 20, including a binocular vision camera 21, a pressure-sensitive tactile array 22, a microphone array 23 and a gas sensor 24, for collecting environmental data;
[0033] A flexible joint robotic arm 30, composed of at least 3 series-connected bionic joints, with a shape memory alloy driver and a pneumatic feedback device built in each joint;
[0034] An edge computing unit 40, which establishes communication connections with the central processor 11 and the multi-modal perception module 20 respectively, for real-time processing of perception data and generating action instructions.
[0035] In this embodiment, by integrating a binocular vision camera 21, a pressure-sensitive tactile array 22, a microphone array 23, and a gas sensor 24, through multi-modal data fusion of vision (3D modeling), touch (contact force detection), hearing (sound source localization), and gas detection (VOC concentration), the limitation of a single sensor is broken through, and the perception robustness in complex environments (such as low light, high noise, and toxic gas leakage) is significantly improved, realizing omni-directional environmental perception. Moreover, the binocular vision and the pressure-sensitive tactile array 22 cooperate to mark moving obstacles (such as pedestrians and vehicles), and the visual misjudgment (such as transparent objects or reflective surfaces) can be corrected in real time through tactile feedback.
[0036] In this embodiment, the flexible joint manipulator 30 is composed of a shape memory alloy (SMA) actuator and a pneumatic feedback device. The SMA actuator provides a fast response (±90° deflection, response time < 0.5 s), which is suitable for precision operations (such as grasping fragile items); the pneumatic feedback device monitors the external force impact in real time, and buffers the instantaneous pressure received by the manipulator through reverse contraction, avoiding damage caused by rigid collisions, and having the effect of balancing high-precision motion and anti-impact; the micro-current sensor of the pressure-sensitive tactile array 22 is linked with the flexible joint, and dynamically adjusts the grasping force according to the object material (such as hardness and surface texture), reducing the risk of slipping and having the ability of adaptive grasping. The flexible joint manipulator 30 realizes the rigid-flexible controllability of the bionic flexible manipulator as a whole.
[0037] In this embodiment, the edge computing unit 40 and the central processor 11 cooperate to process the perception data. On the one hand, through local computing (instead of relying on the cloud), the latency from perception data processing to action instruction generation is reduced to the millisecond level, meeting the real-time requirements of dynamic scenarios (such as obstacle avoidance and emergency steering), and realizing low-latency path planning; on the other hand, the edge computing unit independently runs the reinforcement learning model, reducing the computing power burden on the central processor, ensuring that the system still maintains stable performance during multi-task parallel operation, and realizing efficient resource utilization.
[0038] In this embodiment, the binocular vision camera usually transmits high-resolution image data through the MIPI-CSI interface, the pressure-sensitive tactile array 22 transmits tactile signals (such as the pressure distribution matrix) through the I2C / SPI protocol, the microphone array uses the I2S protocol for audio stream transmission, and the gas sensor sends concentration data through the UART or CAN bus; the vision and tactile data volumes are large, and data compression technologies (such as JPEG-LS and point cloud compression) need to be used to reduce the transmission burden, while low-bandwidth sensors (such as gas sensors) use low-power protocols (such as LoRa) to reduce energy consumption.
[0039] In this embodiment, the PTP protocol (IEEE 1588) is adopted to achieve microsecond-level time synchronization between binocular vision and the microphone array, ensuring the time alignment of audio and video data; the piezoresistive tactile array 22 and the gas sensor synchronize sampling through a hardware trigger signal to reduce timing errors; a unified timestamp is embedded in the data packet, and the transmission delay is compensated by a dynamic calibration algorithm (such as Kalman filtering) of the edge computing unit, ensuring the timing consistency of multimodal data at the central processor end and achieving software timestamp alignment.
[0040] In this embodiment, the edge computing unit (such as the Xinchi X9CC chip) is responsible for real-time preprocessing (such as stereo matching of binocular vision and noise filtering of the piezoresistive tactile array 22), and only transmits key feature data (such as target position, tactile mode) to the central processor, reducing the CPU load. The central processor (such as the ARM Cortex-A / R series heterogeneous architecture) fuses multimodal data and uses deep learning models such as Transformer for cross-modal feature alignment (such as spatial mapping of vision-tactile), and finally generates robotic arm action instructions. The key sensor (such as the piezoresistive tactile array 22) adopts a dual-channel CAN bus design to ensure the reliability of data transmission; the visual data is directly connected to the central processor through a PCIe interface to reduce interference in the middle link. The data link layer adopts AES-256 encryption and digital signature technology (such as HMAC-SHA256) to prevent data tampering; the central processor verifies the sensor identity through the HSM security module to ensure the system security.
[0041] In one embodiment, when grasping a fragile object, the binocular vision camera 21 quickly locates the target position, and the piezoresistive tactile array 22 real-time feeds back the contact force distribution data. The contact force distribution data is compressed by the edge computing unit 40 and then transmitted to the central processor through the PCIe bus interface; the CPU fuses the visual coordinates and the tactile pressure gradient, and adjusts the grasping force of the robotic arm through a shape memory alloy actuator, finally achieving safe operation.
[0042] In this embodiment, the central processor 11 dynamically switches the sensor power supply mode according to the task requirements (such as the visual camera entering the low-power state during inactive periods), and the piezoresistive tactile array 22 adopts distributed PMU power management to reduce the impact of power noise on the signal; the interface is extended through FPGA or CPLD to support the plug-and-play of future newly added sensors (such as infrared thermal imagers), and is compatible with multiple communication protocols (such as USB 3.0, Gigabit Ethernet) to achieve an expandable interface.
[0043] In this embodiment, a central processor is integrated within the main frame 10, which supports modular expansion and has a compact design. The multi-modal sensors, drive units, and computing modules are highly integrated, reducing external wiring interference and enhancing the degrees of freedom of the robotic arm's movement. Through the collaborative design of multi-modal perception fusion, bionic flexible drive, and edge-side real-time computing, the core problems of traditional robotic systems, such as insufficient complex environment perception capabilities, imbalance between the rigid and flexible control of the robotic arm, and lag in dynamic scene response, are solved, providing a highly adaptable and reliable solution for scenarios such as industrial automation and inspection in dangerous environments.
[0044] In one embodiment, the piezoresistive tactile array 22 is provided with a graphene film and a micro-current sensor.
[0045] In this embodiment, the high electrical conductivity of the graphene film (resistivity about 1×10⁻ 6 Ω·m) combined with the micro-current sensor (detection accuracy up to 0.01 μA) can identify changes in contact force below 0.05 N (superior to the 0.1 N threshold of traditional piezoresistive sensors), is suitable for the delicate operation of fragile tissues (such as blood vessels, nerves) by medical robots, and has the ability to detect micro-forces. By analyzing the difference in micro-current distribution, the pressure gradient of the contact surface can be resolved (such as distinguishing the local forces on the fingertips and the eggshell when grasping an egg), preventing object breakage caused by uneven force, and realizing multi-dimensional force feedback. The graphene film has ultra-high sensitivity and resolution.
[0046] In one embodiment, graphene has strong chemical stability, and the circuit design of the micro-current sensor has low noise. In extreme environments with humidity > 90% or temperature -20°C to 80°C, the drift rate of the tactile signal is < 3% (the drift rate of traditional piezoresistive sensors > 15%), ensuring reliable use in scenarios such as chemical industry and polar regions, and having humidity / temperature adaptability. The graphene layer can absorb high-frequency electromagnetic interference (such as near a 5G base station), avoiding distortion of the tactile signal and playing an electromagnetic shielding role.
[0047] In one embodiment, due to the flexibility of the graphene film (bending radius < 1 mm) and the embedded packaging of the micro-current sensor, the piezoresistive tactile array 22 can closely adhere to the curved surface of the robotic arm (such as bionic finger joints), eliminating the measurement blind area between the traditional rigid sensor and the curved surface, and achieving curved surface fitting. The fracture strength of graphene (130 GPa) enables it to maintain a sensitivity of > 95% after 100,000 bending cycles, extending the life of the tactile module and being able to resist mechanical fatigue.
[0048] In one embodiment, the low driving voltage of graphene (1~3V) and the power consumption control of the microcurrent sensor (single channel <0.1mW), wherein the overall power consumption of the pressure-sensitive tactile array 22 is reduced by 60% compared with the traditional solution, is suitable for battery-powered mobile robots; through the piezoelectric effect of the graphene film (theoretical output is about 5V / MPa), mechanical energy collection can be partially realized, reducing dependence on external power supply, and has the potential for self-powering.
[0049] In this embodiment, through the collaborative innovation of graphene materials and microcurrent sensing technology, the pressure-sensitive tactile array 22 has ultra-high sensitivity, strong environmental resistance, flexible fit and low power consumption, which solves the problems of traditional tactile sensors being prone to failure, short life and insufficient precision in complex scenarios. It is particularly suitable for fields with strict requirements on tactile feedback, such as medical minimally invasive surgery and precision electronic assembly.
[0050] In one embodiment, the bionic joint is driven by a composite control of shape memory alloy and pneumatic muscle.
[0051] In this embodiment, SMA (shape memory alloys) achieves micro-displacement control of ±0.1 mm (response time <0.5 s) through current heating, is suitable for precision operations (such as circuit board welding), and can achieve high-precision positioning.
[0052] In one embodiment, the pneumatic muscle can output a pulling force of >200N at an air pressure of 3-5bar, withstand sudden external force impact (such as instantaneous shaking when carrying heavy objects), avoid mechanical damage, and withstand large loads and impacts.
[0053] In this embodiment, the robot dynamically switches between "rigid mode" (dominated by SMA) and "flexible mode" (dominated by pneumatic muscles) to adapt to the different needs of grasping fragile objects (glass cups) and carrying heavy objects (metal boxes).
[0054] In one embodiment, SMA consumes energy only when deformation is triggered, which is more than 40% lower than the energy consumption of traditional motor drive. Pneumatic muscles use compressed air circulation for energy, and the gas released by the contraction of pneumatic muscles can be recycled into the gas tank, achieving more than 70% reuse of air pressure energy. When SMA temporarily fails due to overheating, pneumatic muscles can independently maintain basic joint movement functions (such as emergency obstacle avoidance), thereby improving the system's fault tolerance.
[0055] In this embodiment, when the end of the robotic arm is impacted, the pneumatic muscle triggers buffer contraction within 10ms, and at the same time the SMA fine-tunes the joint angle to disperse the impact force to the overall structure (vibration attenuation rate > 60%), achieving real-time shock absorption adjustment.
[0056] In one embodiment, according to the feedback of the pressure-sensitive tactile array 22, the driving ratio of the SMA and the pneumatic muscle is dynamically adjusted. For example, when grasping a sponge: the pneumatic muscle dominates (low stiffness to prevent squeezing deformation); when grasping a bolt: the SMA dominates (high stiffness to ensure the accuracy of the tightening torque).
[0057] In this embodiment, through the combined drive of the SMA and the pneumatic muscle, the performance boundary of a single drive technology is broken through, and a dynamic balance of high precision and high load, low energy consumption and high reliability, rigidity and flexibility is achieved.
[0058] In one embodiment, the edge computing unit 40 has a built-in reinforcement learning model for optimizing path planning through the Q-learning algorithm.
[0059] In this embodiment, the Q-learning algorithm does not require a pre-set complete environmental map. It autonomously generates a feasible path in an unknown scenario (such as an earthquake ruins) through trial-and-error learning. The exploration efficiency is increased by 60% compared with the random search algorithm, and there is no dependence on prior knowledge.
[0060] In one embodiment, the Q-learning algorithm only densely updates the Q values in the high-frequency access areas (such as corridor intersections) in the state-action space, reducing the computational amount by 75% compared with the deep reinforcement learning (DRL) model with global updates, achieving low computing power consumption; using a hash table to compress and store the Q value matrix, making the memory occupancy of path planning for a 1 km² scenario < 50 MB, adapting to the resource limitations of edge devices.
[0061] In one embodiment, the improved Q-learning algorithm integrates an obstacle movement prediction model (such as a Kalman filter or an LSTM network). The prediction error of the dynamic obstacle trajectory is < 0.2 m (such as a pedestrian suddenly turning), and the path safety redundancy distance is reduced by 30%, avoiding overly conservative detours and being able to predictively avoid obstacles; through the long-term reward accumulation mechanism of Q-learning, suppressing the interference of sensor instantaneous noise (such as visual misdetection of obstacles) on the path, reducing the mis-avoidance rate to < 5%, and achieving anti-sensing noise.
[0062] In this embodiment, through the marginalized deployment of the Q-learning algorithm and the reinforcement learning model, the dependence of traditional path planning on static environments and high computing power is broken through, achieving real-time response in dynamic scenarios, efficient learning in resource-constrained environments, and multi-objective collaborative optimization capabilities, providing a highly adaptable and low-latency autonomous navigation solution for scenarios such as warehousing logistics (high-dynamic cargo handling) and disaster rescue (unknown ruins exploration).
[0063] In one embodiment, as Figure 2 shown, the above robot system further includes:
[0064] The human-machine interaction module 50 is provided with a voice command recognition, gesture control, and electroencephalogram signal interface.
[0065] In this embodiment, the triple interaction interfaces of voice, gesture, and electroencephalogram signals can achieve scene adaptive control. On the one hand, for example, in an open environment (such as a factory workshop), natural language commands (such as "move to area A and grab the red part") are supported, and the recognition accuracy rate is >95% (when the noise <70 dB), realizing voice command recognition; on the other hand, dynamic gestures (such as fist clenching - opening to control the robotic arm to grab) are captured through binocular vision 21, which is applicable to scenarios with high noise or requiring silent operation (such as a laboratory), realizing gesture control; furthermore, based on a non-invasive EEG headset (such as the P300 signal), the user's intention is analyzed to help users with physical disabilities or in high-risk environments (such as nuclear power plants) achieve "contactless control". In addition, when a certain modality fails (such as voice being interfered by high-frequency noise), the system automatically switches to other interaction methods to ensure the continuity of commands, having a redundant fault tolerance mechanism.
[0066] In one embodiment, the human-machine interaction module 50 establishes a communication connection with the edge computing unit 40. The edge computing unit 40 locally processes the interaction signals. The voice command recognition delay <200 ms (lightweight deployment of the end-to-end model), the gesture trajectory tracking frequency ≥30 Hz, and the electroencephalogram signal analysis period <1 s, meeting the real-time control requirements; the interaction module supports multi-system protocols such as ROS / Android / IOS and can be quickly connected to third-party devices (such as AR glasses, smart gloves) through the API.
[0067] When the voice command is ambiguous (such as "avoid that thing"), the real-time object positioning by visual perception ("that thing" = dynamically marked obstacle) is combined to correct the path planning, realizing intention error correction; through voice intonation analysis (such as a hasty command triggering an emergency mode) and electroencephalogram emotion recognition (such as an anxiety signal activating a safety lock), the safety of human-machine collaboration is improved, realizing emotional interaction.
[0068] In this embodiment, through the multi-modal interaction fusion of voice - gesture - electroencephalogram, the limitation of traditional robots relying on a single control method is broken through, realizing full-scene adaptability, highly inclusive operation (covering healthy users and disabled groups), and accurate intention analysis.
[0069] In one embodiment, the gas sensor 24 is used to detect the concentration of volatile organic compounds (VOC).
[0070] In one embodiment, the high sensitivity (detection lower limit ≤ 1 ppm) and wide detection range (0 - 1000 ppm) of the VOC sensor are used to detect dangerous gases such as benzene, formaldehyde, and methane in real time during chemical leaks, the initial stage of a fire, or in confined spaces (such as underground pipelines). It can trigger an alarm 3 - 5 minutes earlier than traditional smoke sensors, enabling the identification of hidden risks; through the VOC concentration distribution map (fused with a 3D map), the leakage source (such as a pipeline rupture point) can be located with an error < 0.5 m, guiding robots or personnel to quickly intervene and dispose of the situation, and realizing concentration gradient analysis.
[0071] In one embodiment, the edge computing unit 40 is built - in with a reinforcement learning model for optimizing path planning through the Q - learning algorithm. At this time, when the VOC concentration exceeds a threshold (such as 50 ppm), the edge computing unit 40 automatically generates a detour path to avoid high - pollution areas. At the same time, it combines wind speed data to predict the gas diffusion direction, preventing the robot from entering the diffusion path, and realizing a dynamic poison - avoiding path; it triggers the mechanical arm sealing mode (such as enabling a self - repairing coating to isolate VOC corrosion) or starts the air purification module (optional), extending the operation time of the device in a toxic environment and realizing the switching of the protection level.
[0072] In one embodiment, when visual recognition detects "transparent liquid leakage", VOC detection is used to distinguish whether it is water (harmless) or acetone (dangerous), avoiding misjudgment and eliminating perceptual ambiguity; when the pressure - sensitive tactile array 22 detects an unknown viscous substance, its volatile characteristics (such as the smell of ethanol) are combined to speculate that it is alcohol gel, and the grasping force is optimized.
[0073] In this embodiment, through the special strengthening of the VOC concentration detection ability, the perception ability of the robot system in a chemical - dangerous environment is upgraded from "passive obstacle avoidance" to a full - chain safety mechanism of "active early warning - dynamic protection - intelligent decision - making", significantly improving the operation safety and task reliability in scenarios such as chemical inspection, post - disaster rescue (such as exploring toxic ruins), and laboratory automation.
[0074] In one embodiment, the central processor 11 supports the federated learning framework, allowing multiple robots to share local model parameters.
[0075] The central processor 11 aggregates the local model parameters of multiple robots (such as the path - planning Q - value table, VOC recognition model) through an encrypted channel. Different robots (such as factory inspection machines, medical delivery machines) share the feature - layer parameters, enabling a single robot to obtain prior knowledge of unexperienced scenarios (such as hospital corridors). The convergence speed of path planning in the new environment is increased by 50%, realizing cross - scenario generalization ability; the decoupling of the local model update frequency (such as every 10 minutes) and the global model synchronization period (such as every 2 hours) avoids computing power fluctuations caused by frequent communication, and realizes incremental model optimization.
[0076] In one embodiment, multiple robots share the local parameters of the Q-value table through federated learning, reducing the path planning time by 40% when a new robot enters a similar environment.
[0077] In this embodiment, the original data in federated learning is stored locally, and only the model gradient difference parameters are uploaded. Sensitive data such as the patient position data of medical robots and the production line layout information of industrial robots do not leave the local area, meeting the requirements of regulations such as GDPR / HIPAA and having privacy compliance; homomorphic encryption and gradient noise injection are used to prevent the original data from being reverse-derived during the parameter transmission process, reducing the model leakage risk by 90% and resisting malicious attacks.
[0078] In one embodiment, the parameter compression rate > 70%. Synchronizing a 1GB model for 100 robots only requires 10MB of bandwidth (traditional centralized training requires 100GB), saving bandwidth; the frequency of participating in federated learning is dynamically adjusted according to the remaining battery power of the robots (high-power nodes undertake more training tasks), extending the overall battery life of the cluster by 20% and achieving energy efficiency balance.
[0079] In one embodiment, the federated learning framework integrates environment difference-aware weights (for example, the weights of robots in chemical industrial areas are higher than those in office areas). In the event of a sudden leakage, the weight of the robot model near the accident point is increased to 80%, quickly optimizing the risk handling strategy of the global model and achieving hotspot area focus; when the environment distribution changes (such as the adjustment of warehouse shelf layout), by detecting the difference between the local and global model losses, incremental learning is automatically triggered, and the model iteration cycle is shortened to 30 minutes, achieving concept drift response.
[0080] In this embodiment, through the distributed cooperation mechanism of the federated learning framework, knowledge sharing and efficient model evolution of multiple robots under data privacy protection are realized, overcoming the bottlenecks of single-agent learning such as slow cross-scenario adaptation, low data utilization efficiency, and weak security compliance, providing an extensible and highly robust collaborative intelligent base for large-scale robot cluster applications such as smart cities (cross-regional patrols) and logistics warehousing (multi-warehouse linkage).
[0081] In one embodiment, as Figure 2 shown, the above robot system further includes:
[0082] A self-healing coating 60, covering the surface of the robotic arm, and the materials in the self-healing coating 60 activate the repair function at a preset temperature.
[0083] In one embodiment, the microcapsules in the self-healing coating encapsulate a healing agent (such as a polyurethane prepolymer), and a preset temperature (60 °C) triggers the capsules to rupture and release the healing agent: for scratches or cracks with a width ≤ 50 μm, the repair is completed within 5 minutes after the temperature trigger, and the tensile strength recovery rate of the coating after repair > 90%, achieving the repair of microscopic cracks; after the healing agent fills the cracks, it isolates external corrosive media (such as acid rain, VOC gases), extending the life of the robotic arm in a chemical environment by 2 - 3 times and enhancing corrosion resistance.
[0084] In one embodiment, the temperature trigger mechanism can be linked with the working conditions of the robotic arm. When the robotic arm moves at high speed, the surface friction temperature can reach 60 - 80 °C, automatically activating the repair function and repairing micro-damages caused by wear in real time; in a low-temperature environment (such as polar operations), it is locally heated to the preset temperature through built-in heating wires to ensure the normal operation of the coating repair function, realizing external heat source-assisted repair.
[0085] In one embodiment, the coating adopts a flexible match between an elastomeric substrate (such as silicone rubber) and the healing agent. When the robotic arm joint bends, the coating can withstand a tensile deformation > 200% without peeling, and the flexibility retention rate after repair > 95%, achieving the technical effect of no breakage during bending; the healing agent is chemically bonded to the substrate, and there is no delamination phenomenon in the repaired area after 100,000 bending cycles, generally realizing the strengthening of interface adhesion.
[0086] In one embodiment, self-repair replaces manual maintenance. For minor damages (accounting for more than 70% of the robotic arm failures), there is no need to stop the machine for disassembly and repair, reducing the annual maintenance cost by 50% and reducing manual intervention; in scenarios such as disaster rescue, the robotic arm can continuously operate > 500 hours without surface maintenance, and the risk of task interruption drops by 80%, achieving continuous operation guarantee.
[0087] In this embodiment, through the design of the temperature-responsive self-healing coating, the real-time repair of damages, corrosion resistance enhancement, and long-life operation of the robotic arm in complex environments are realized, overcoming the bottlenecks such as performance degradation and frequent maintenance caused by surface damages of traditional robotic arms, providing a highly reliable and low-maintenance robot protection solution for scenarios such as chemical inspection, space exploration (extreme temperature fluctuations), and deep-sea operations (high-salt corrosion).
[0088] In addition, as Figure 3 shown, a method for autonomous navigation is also provided. Using the above robotic system, the method for autonomous navigation includes:
[0089] S1: Construct a three-dimensional map of the environment through the multi-modal perception module and mark dynamic obstacles.
[0090] In one embodiment, centimeter-level map accuracy: the binocular vision point cloud resolution reaches ±2 cm, the pressure-sensitive tactile array 22 supplements the visual blind area (such as transparent glass), the dynamic obstacle marking error < 5 cm, achieving centimeter-level map accuracy; the dynamic obstacle tracking frequency ≥ 10 Hz (such as moving pedestrians and vehicles), supporting obstacle avoidance response in high-speed motion scenarios (robot moving speed ≤ 2 m / s), and real-time environment update.
[0091] S2: Calculate the initial path using the improved A* algorithm and optimize the path weight in combination with the reinforcement learning model.
[0092] In one embodiment, the improved A* algorithm and the Q-learning model cooperate to optimize the path weight, comprehensively optimizing the path length, energy consumption, and safety factor, improving the "safety-efficiency" balance of the inspection path in the chemical industrial park by 40%; integrating the LSTM obstacle trajectory prediction model (prediction duration 3 - 5 seconds), reducing the path safety redundancy distance by 50%, avoiding excessive detours, and realizing dynamic obstacle prediction.
[0093] S3: Adjust the movement speed and angle of the robotic arm according to the real-time stress feedback of the flexible joint.
[0094] In one embodiment, the real-time control closed-loop of the flexible joint stress feedback and the edge computing unit. When the joint stress exceeds the threshold (such as 20 MPa), the movement speed of the robotic arm automatically decreases by 50% to prevent structural damage and achieve overload protection; dynamically adjust the joint angle according to the stress distribution, improving the movement smoothness of the robotic arm in a narrow space (such as inside a pipeline) by 60%, reducing the risk of jamming, and realizing bionic motion optimization.
[0095] In this embodiment, through the full-link fusion of environment perception - algorithm decision - execution control, high-precision modeling, intelligent path optimization, and safe motion control in dynamic and complex scenarios are achieved, solving the fragmentation problem of traditional navigation methods in the "perception blind area - planning rigidity - control lag" chain, and providing a real-time and robust autonomous navigation solution for scenarios such as intelligent manufacturing (flexible production line logistics) and disaster rescue (ruins search and rescue).
[0096] In one embodiment, an obstacle movement prediction model is introduced for path weight optimization.
[0097] In one embodiment, an obstacle movement prediction model (such as Kalman filter, LSTM time series network) is fused with an improved A* algorithm, and the position prediction error of dynamic obstacles such as pedestrians and AGVs is < 0.3 m (within 1 second), the speed prediction error is < 15%, the path avoidance distance is reduced by 40%, the detour redundancy is reduced, and the trajectory prediction accuracy is improved; it supports tracking more than 10 dynamic obstacles simultaneously, predicting the time-to-collision (TTC), and preferentially avoiding high-risk targets (such as fast-moving objects with TTC < 2 seconds), realizing multi-obstacle collaborative prediction.
[0098] Among them, the LSTM time series network is Long Short-Term Memory.
[0099] In one embodiment, the prediction model outputs a probability distribution map of the future trajectory of the obstacle, quantifies the dynamic risk weight, and in the heuristic function of the A* algorithm, the path cost in the high-probability conflict area (such as the corridor intersection) is increased by 3 to 5 times, guiding the robot to select a path with a low conflict probability and generating a risk-sensitive path; when the obstacle suddenly turns or accelerates (such as a pedestrian suddenly running), the model updates the weight matrix every 0.5 seconds, and the path adjustment response time < 200 ms, realizing real-time weight update.
[0100] In one embodiment, the lightweight prediction model is deployed on the edge computing unit 40. The single-obstacle prediction calculation amount is < 10 MFLOPS, and the CPU occupancy rate is < 25% when 10 obstacles are predicted in parallel, meeting the real-time requirements of edge devices and realizing low computing power consumption; through path cost smoothing filtering (such as exponential weighted moving average), the path oscillation caused by instantaneous prediction noise is suppressed, and the smoothness of the planned trajectory is improved by 70%, realizing anti-prediction jitter.
[0101] In one embodiment, the prediction model is linked with multi-modal perception (vision, sound source localization) data. The position of the obstacle is detected by binocular vision, and the predicted trajectory is corrected by combining the sound source direction of the microphone array 23 (such as vehicle honking), solving the problem of prediction failure caused by visual occlusion; a dedicated prediction model is pre-trained for typical scenarios to predict the fixed route movement mode of AGVs and plan the avoidance strategy at the intersection in advance; the crowd flow trend (such as the tidal flow of people at the subway exit) is recognized, and the globally optimal detour path is generated.
[0102] In this embodiment, through the deep binding of the dynamic obstacle prediction model and the path weight, the traditional path planning is upgraded from "passive obstacle avoidance" to an intelligent decision-making mode of "active prediction - dynamic optimization", significantly improving the navigation efficiency and safety in high-dynamic scenarios (such as transportation hubs, flexible production lines), while ensuring the stable operation of the algorithm in resource-constrained environments, providing core technical support for applications such as autonomous mobile robots (AMRs) and unmanned delivery vehicles.
[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0104] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method including such element.
[0105] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A robot system based on multi-modal perception collaboration, characterized in that Comprising: A main body framework, with a central processing unit integrated inside the main body framework; A multi-modal perception module, including a binocular vision camera, a pressure-sensitive tactile array, a microphone array and a gas sensor, for collecting environmental data; A flexible joint robotic arm, composed of at least 3 serially-connected bionic joints, with a shape memory alloy actuator and a pneumatic feedback device built in each joint; An edge computing unit, establishing communication connections with the central processing unit and the multi-modal perception module respectively, for real-time processing of perception data and generating action instructions; The pressure-sensitive tactile array is provided with a graphene film and a micro-current sensor, and the micro-current sensor is embedded and encapsulated so that the pressure-sensitive tactile array can be closely attached to the curved surface of the robotic arm; the graphene film collects mechanical energy based on the piezoelectric effect, enabling the pressure-sensitive tactile array to achieve partial self-power supply and perform electromagnetic shielding; The driving mode of the bionic joint is the composite control of a shape memory alloy and a pneumatic muscle, and according to the feedback of the pressure-sensitive tactile array, the driving ratio of the shape memory alloy and the pneumatic muscle is dynamically adjusted; among them, the pneumatic muscle uses compressed air for cyclic energy supply, and the gas released by the contraction and relaxation of the pneumatic muscle can be recycled to the air storage tank; when the shape memory alloy fails temporarily due to overheating, the basic motion function of the joint is independently maintained based on the pneumatic muscle; when the end of the robotic arm is impacted, the pneumatic muscle triggers a buffer contraction within 10 ms, and at the same time the shape memory alloy adjusts the joint angle to disperse the impact force to the overall structure.
2. The robot system according to claim 1, characterized in that The edge computing unit is built with a reinforcement learning model, and optimizes the path planning through the Q-learning algorithm.
3. The robot system according to claim 1, wherein, Also including: A human-computer interaction module, provided with voice command recognition, gesture control and brain wave signal interfaces.
4. The robot system according to claim 1, characterized in that The gas sensor is used to detect the concentration of volatile organic compounds.
5. The robot system according to claim 1, characterized in that The central processing unit supports a federated learning framework, allowing multiple robots to share local model parameters.
6. The robot system according to claim 1, characterized in that Also including: A self-healing coating, covering the surface of the robotic arm, and the materials in the self-healing coating activate the repair function at a preset temperature.
7. An autonomous navigation method, using the robot system according to any one of claims 1 to 6, the autonomous navigation method comprising: S1: Construct a three-dimensional environmental map through the multi-modal perception module and mark dynamic obstacles; S2: Calculate an initial path using the improved A* algorithm and optimize the path weight in combination with the reinforcement learning model; S3: Adjust the movement speed and angle of the robotic arm according to the real-time stress feedback of the flexible joint.
8. The autonomous navigation method according to claim 7, characterized in that, In step S2, the obstacle movement prediction model is introduced for path weight optimization.
Citation Information
Patent Citations
Variable-stiffness self-repairing material containing metal-sulfydryl coordinate bond as well as preparation and application thereof
CN113817173A
Flexible mechanical arm based on coupling of memory alloy and pneumatic artificial muscle
CN116038679A
Multi-perception fusion bionic search robot with man-machine bidirectional interaction function
CN117021112A
Path navigation control method and system of intelligent mechanical arm
CN118848989A
Large model and small model collaborative robot operation action real-time control method and system
CN119567267A