Robot system based on multi-mode perception cooperation and autonomous navigation method
By integrating multimodal perception modules, flexible joint robot arms and edge computing units in the robot system, the problems of insufficient perception capabilities and imbalance in motion control in complex environments are solved, and high perception and flexible movement are achieved, providing high adaptability and high reliability solutions for industrial automation and dangerous environment inspections.
Patent Information
- Application Number
- CN202510445257.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional robot systems have problems such as insufficient perception capabilities, imbalance in rigid and flexible control of robotic arms, and lag in response to dynamic scenes in complex environments.
Using a robot system based on multimodal perception collaboration, combining flexible joint robot arms and edge computing units, the robot's high perceptual and flexible movement in complex environments is achieved through multimodal perception data fusion, bionic flexible driving and real-time edge computing.
It significantly improves the perceived robustness and freedom of movement of robots in complex environments, solves the lack of perception and control of traditional robot systems, and provides high adaptability and high reliability solutions for industrial automation and dangerous environment inspections.
Smart Images

Figure CN119952734A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent robots, and in particular to a robot system and an autonomous navigation method based on multimodal perception collaboration. Background Art
[0002] When current robots move, traditional rigid joints limit their range of motion and lack the ability to adaptively respond to collisions or external pressure. Summary of the invention
[0003] In view of this, the present application provides a robot system and an autonomous navigation method based on multimodal perception collaboration, which can improve the robot's motion control capability.
[0004] A robot system based on multimodal perception and collaboration, comprising: A main frame, wherein a central processing unit is integrated inside the main frame; A multimodal perception module, which includes a binocular vision camera, a pressure-sensitive tactile array, a microphone array, and a gas sensor, for collecting environmental data; The flexible joint robot arm is composed of at least three bionic joints connected in series, each of which has a built-in shape memory alloy driver and a pneumatic feedback device; The edge computing unit establishes communication connections with the central processing unit and the multimodal perception module respectively, and is used to process the perception data in real time and generate action instructions.
[0005] In one embodiment, the pressure-sensitive tactile array is provided with a graphene film and a micro-current sensor.
[0006] In one embodiment, the bionic joint is driven by a composite control of shape memory alloy and pneumatic muscle.
[0007] In one embodiment, the edge computing unit has a built-in reinforcement learning model and optimizes path planning through a Q-learning algorithm.
[0008] In one embodiment, the above-mentioned robot system further comprises: The human-computer interaction module is equipped with voice command recognition, gesture control and brain wave signal interface.
[0009] In one embodiment, the gas sensor is used to detect the concentration of volatile organic compounds.
[0010] In one embodiment, the central processor supports a federated learning framework, allowing multiple robots to share local model parameters.
[0011] In one embodiment, the above-mentioned robot system further comprises: The self-healing coating is covered on the surface of the robot arm. The material in the self-healing coating activates the repair function at a preset temperature.
[0012] In addition, an autonomous navigation method is also provided, using the above robot system, the autonomous navigation method comprising: S1: Build a 3D map of the environment through a multimodal perception module and mark dynamic obstacles; S2: Use the improved A* algorithm to calculate the initial path and optimize the path weight in combination with the reinforcement learning model; S3: Adjust the movement speed and angle of the robotic arm according to the real-time stress feedback of the flexible joint.
[0013] In one embodiment, path weight optimization introduces an obstacle movement prediction model.
[0014] The above-mentioned robot system and autonomous navigation method based on multimodal perception collaboration, the robot system includes: a main frame, the main frame has a central processor integrated inside; a multimodal perception module, including a binocular vision camera, a pressure-sensitive tactile array, a microphone array and a gas sensor, for collecting environmental data; a flexible joint robot arm, composed of at least 3 bionic joints in series, each joint has a built-in shape memory alloy driver and a gas pressure feedback device; an edge computing unit, respectively, establishes a communication connection with the central processor and the multimodal perception module, for real-time processing of perception data and generating action instructions, the main frame is integrated with a central processor, supports modular expansion, compact design: multimodal sensors, drive units and computing modules are highly integrated, reducing external wiring interference and improving the freedom of movement of the robot arm. Through the collaborative design of multimodal perception fusion, bionic flexible drive and edge real-time computing, the core problems of traditional robot systems in complex environment perception, imbalance of rigid and flexible control of robot arms and delayed response to dynamic scenes are solved, providing highly adaptable and reliable solutions for scenes such as industrial automation and hazardous environment inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic block diagram of the structure of a robot system based on multimodal perception collaboration provided in one embodiment of the present application; Figure 2 A schematic block diagram of a multi-modal sensing and collaborative robot system according to another embodiment of the present application; Figure 3 A flowchart of an autonomous navigation method provided in accordance with an embodiment of the present application.
[0016] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0017] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0018] In addition, if the descriptions of "first", "second", etc. are involved in this application, they are only used for descriptive purposes (such as for distinguishing the same or similar elements), and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0019] like Figure 1 As shown, a robot system based on multimodal perception collaboration is provided, and the robot system includes: A main frame 10, wherein a central processing unit 11 is integrated inside the main frame 10; The multimodal sensing module 20 includes a binocular vision camera 21, a pressure-sensitive tactile array 22, a microphone array 23, and a gas sensor 24, which are used to collect environmental data; The flexible joint robot arm 30 is composed of at least three bionic joints connected in series, each joint having a built-in shape memory alloy driver and a pneumatic feedback device; The edge computing unit 40 establishes communication connections with the central processing unit 11 and the multimodal perception module 20 respectively, and is used to process the perception data in real time and generate action instructions.
[0020] In this embodiment, by integrating a binocular camera 21, a pressure-sensitive tactile array 22, a microphone array 23 and a gas sensor 24, through the multimodal data fusion of vision (3D modeling), touch (contact force detection), hearing (sound source localization) and gas detection (VOC concentration), the limitation of a single sensor is broken, and the perception robustness in complex environments (such as low light, high noise and toxic gas leakage) is significantly improved, and all-round environmental perception is achieved. In addition, the binocular vision and the pressure-sensitive tactile array 22 work together to mark moving obstacles (such as pedestrians and vehicles), and correct visual misjudgments (such as transparent objects or reflective surfaces) in real time through tactile feedback.
[0021] In this embodiment, the flexible joint mechanical arm 30 is composed of a shape memory alloy (SMA) driver and a pneumatic feedback device. The SMA driver provides a fast response (±90° deflection, response time <0.5s), which is suitable for precision operations (such as grasping fragile objects); the pneumatic feedback device monitors the impact of external forces in real time, and buffers the instantaneous pressure on the mechanical arm by reverse contraction to avoid damage caused by rigid collision, and has the effect of high-precision movement and impact resistance balance; the microcurrent sensor of the pressure-sensitive tactile array 22 is linked with the flexible joint, and dynamically adjusts the grasping force according to the material of the object (such as hardness, surface texture), reduces the risk of slippage, and has adaptive grasping ability. The flexible joint mechanical arm 30 realizes the rigidity and flexibility controllability of the bionic flexible mechanical arm as a whole.
[0022] In this embodiment, the edge computing unit 40 and the central processor 11 cooperate to process the perception data. On the one hand, through localized computing (rather than relying on the cloud), the delay from processing the perception data to generating action instructions is reduced to milliseconds, meeting the real-time requirements of dynamic scenarios (such as obstacle avoidance and emergency turns) and realizing low-latency path planning. On the other hand, the edge computing unit independently runs the reinforcement learning model to reduce the computing power burden of the central processor, ensuring that the system maintains stable performance when multiple tasks are running in parallel, thereby realizing efficient resource utilization.
[0023] In this embodiment, the binocular vision camera usually transmits high-resolution image data through the MIPI-CSI interface, the pressure-sensitive tactile array 22 transmits tactile signals (such as pressure distribution matrix) through the I2C / SPI protocol, the microphone array uses the I2S protocol for audio stream transmission, and the gas sensor sends concentration data through the UART or CAN bus; the amount of visual and tactile data is large, and data compression technology (such as JPEG-LS, point cloud compression) is required to reduce the transmission burden, while low-bandwidth sensors (such as gas sensors) use low-power protocols (such as LoRa) to reduce energy consumption.
[0024] In this embodiment, the PTP protocol (IEEE 1588) is used to achieve microsecond-level time synchronization between binocular vision and the microphone array to ensure the time alignment of audio and video data; the pressure-sensitive tactile array 22 and the gas sensor are synchronously sampled through hardware trigger signals to reduce timing errors; a unified timestamp is embedded in the data packet, and the transmission delay is compensated by the dynamic calibration algorithm of the edge computing unit (such as Kalman filtering) to ensure the timing consistency of multimodal data on the central processor side and achieve software timestamp alignment.
[0025] In this embodiment, the edge computing unit (such as the Xinchi X9CC chip) is responsible for real-time preprocessing (such as stereo matching of binocular vision and noise filtering of the pressure-sensitive tactile array 22), and only transmits key feature data (such as target position and tactile pattern) to the central processing unit to reduce the CPU load. The central processing unit (such as ARM Cortex-A / R series heterogeneous architecture) fuses multimodal data and uses deep learning models such as Transformer to perform cross-modal feature alignment (such as visual-tactile spatial mapping), and finally generates robotic arm motion instructions. Key sensors (such as the pressure-sensitive tactile array 22) adopt a dual-channel CAN bus design to ensure the reliability of data transmission; visual data is directly connected to the central processing unit through the PCIe interface to reduce interference in the intermediate links, and the data link layer adopts AES-256 encryption and digital signature technology (such as HMAC-SHA256) to prevent data tampering; the central processing unit verifies the identity of the sensor through the HSM security module to ensure system security.
[0026] In one embodiment, when grasping fragile objects, the binocular vision camera 21 quickly locates the target position, and the pressure-sensitive tactile array 22 provides real-time feedback of the contact force distribution data. The touch force distribution data is compressed by the edge computing unit 40 and transmitted to the central processing unit through the PCIe bus interface; the CPU integrates the visual coordinates with the tactile pressure gradient, and adjusts the gripping force of the robotic arm through the shape memory alloy driver, ultimately achieving safe operation.
[0027] In this embodiment, the central processing unit 11 dynamically switches the sensor power supply mode according to task requirements (such as the visual camera enters a low-power state during inactive periods), and the pressure-sensitive tactile array 22 adopts distributed PMU power management to reduce the impact of power supply noise on the signal; interface expansion is achieved through FPGA or CPLD, supporting plug-and-play of new sensors (such as infrared thermal imagers) in the future, and is compatible with multiple communication protocols (such as USB 3.0, Gigabit Ethernet), realizing an extensible interface.
[0028] In this embodiment, the main frame 10 integrates a central processor, supports modular expansion, and has a compact design; the multimodal sensor, drive unit, and computing module are highly integrated to reduce external wiring interference and improve the freedom of movement of the robot arm. Through the collaborative design of multimodal perception fusion, bionic flexible drive, and edge-end real-time computing, the core problems of traditional robot systems such as insufficient perception of complex environments, imbalance of rigid-flexible control of the robot arm, and delayed response to dynamic scenarios are solved, providing highly adaptable and reliable solutions for scenarios such as industrial automation and inspections in hazardous environments.
[0029] In one embodiment, the pressure-sensitive tactile array 22 is provided with a graphene film and a micro-current sensor.
[0030] In this embodiment, the high conductivity of the graphene film (resistivity of about 1×10⁻ 6 The piezoresistive sensor can detect contact forces below 0.05 N (better than the 0.1 N threshold of traditional piezoresistive sensors), which is suitable for medical robots to perform delicate operations on fragile tissues (such as blood vessels and nerves) and has micro-force detection capabilities. The micro-current distribution difference can be used to analyze the pressure gradient of the contact surface (such as distinguishing the local force between the fingertips and the eggshell when grabbing an egg), preventing damage to objects caused by uneven force, and realizing multi-dimensional force feedback. The graphene film has ultra-high sensitivity and resolution.
[0031] In one embodiment, graphene has strong chemical stability, and the circuit design of the microcurrent sensor has low noise. In extreme environments with humidity > 90% or temperature -20°C ~ 80°C, the tactile signal drift rate is < 3% (the drift rate of traditional piezoresistive sensors is > 15%), ensuring reliable use in chemical, polar and other scenarios, and has humidity / temperature adaptability: the graphene layer can absorb high-frequency electromagnetic interference (such as near 5G base stations), avoid tactile signal distortion, and play an electromagnetic shielding role.
[0032] In one embodiment, due to the flexibility of the graphene film (bending radius <1mm) and the embedded packaging of the microcurrent sensor, the pressure-sensitive tactile array 22 can fit tightly to the curved surface of the robotic arm (such as a bionic finger joint), eliminating the measurement blind spot between the traditional rigid sensor and the curved surface and achieving curved surface fitting; the fracture strength of graphene (130GPa) enables it to maintain a sensitivity of >95% after 100,000 bending cycles, thereby extending the life of the tactile module and being able to resist mechanical fatigue.
[0033] In one embodiment, the low driving voltage of graphene (1~3V) and the power consumption control of the microcurrent sensor (single channel <0.1mW), wherein the overall power consumption of the pressure-sensitive tactile array 22 is reduced by 60% compared with the traditional solution, is suitable for battery-powered mobile robots; through the piezoelectric effect of the graphene film (theoretical output is about 5V / MPa), mechanical energy collection can be partially realized, reducing dependence on external power supply, and has the potential for self-powering.
[0034] In this embodiment, through the collaborative innovation of graphene materials and microcurrent sensing technology, the pressure-sensitive tactile array 22 has ultra-high sensitivity, strong environmental resistance, flexible fit and low power consumption, which solves the problems of traditional tactile sensors being prone to failure, short life and insufficient precision in complex scenarios. It is particularly suitable for fields with strict requirements on tactile feedback, such as medical minimally invasive surgery and precision electronic assembly.
[0035] In one embodiment, the bionic joint is driven by a composite control of shape memory alloy and pneumatic muscle.
[0036] In this embodiment, SMA (shape memory alloys) achieves micro-displacement control of ±0.1 mm (response time <0.5 s) through current heating, is suitable for precision operations (such as circuit board welding), and can achieve high-precision positioning.
[0037] In one embodiment, the pneumatic muscle can output a pulling force of >200N at an air pressure of 3-5bar, withstand sudden external force impact (such as instantaneous shaking when carrying heavy objects), avoid mechanical damage, and withstand large loads and impacts.
[0038] In this embodiment, the robot dynamically switches between "rigid mode" (dominated by SMA) and "flexible mode" (dominated by pneumatic muscles) to adapt to the different needs of grasping fragile objects (glass cups) and carrying heavy objects (metal boxes).
[0039] In one embodiment, SMA consumes energy only when deformation is triggered, which is more than 40% lower than the energy consumption of traditional motor drive. Pneumatic muscles use compressed air circulation for energy, and the gas released by the contraction of pneumatic muscles can be recycled into the gas tank, achieving more than 70% reuse of air pressure energy. When SMA temporarily fails due to overheating, pneumatic muscles can independently maintain basic joint movement functions (such as emergency obstacle avoidance), thereby improving the system's fault tolerance.
[0040] In this embodiment, when the end of the robotic arm is impacted, the pneumatic muscle triggers buffer contraction within 10ms, and at the same time the SMA fine-tunes the joint angle to disperse the impact force to the overall structure (vibration attenuation rate > 60%), achieving real-time shock absorption adjustment.
[0041] In one embodiment, the driving ratio of SMA and pneumatic muscle is dynamically adjusted according to the feedback of the pressure-sensitive tactile array 22. For example, when grabbing a sponge, pneumatic muscle is dominant (low stiffness to prevent squeezing and deformation); when grabbing a bolt, SMA is dominant (high stiffness to ensure the accuracy of tightening torque).
[0042] In this embodiment, the composite drive of SMA and pneumatic muscle breaks through the performance boundary of single drive technology, and achieves a dynamic balance between high precision and high load, low energy consumption and high reliability, and rigidity and flexibility.
[0043] In one embodiment, the edge computing unit 40 has a built-in reinforcement learning model for optimizing path planning through a Q-learning algorithm.
[0044] In this embodiment, the Q-learning algorithm does not need to pre-set a complete environmental map. It can autonomously generate feasible paths in unknown scenarios (such as earthquake ruins) through trial and error learning. The exploration efficiency is improved by 60% compared with the random search algorithm, and there is no reliance on prior knowledge.
[0045] In one embodiment, the Q-learning algorithm only densely updates the Q value for high-frequency access areas in the state-action space (such as corridor intersections), which reduces the amount of computation by 75% compared to the global update deep reinforcement learning (DRL) model and achieves low computing power consumption. The Q value matrix is compressed and stored in a hash table, so that the path planning memory usage of a 1km² scene is less than 50MB, which adapts to the resource limitations of edge devices.
[0046] In one embodiment, an improved Q-learning algorithm is integrated with an obstacle movement prediction model (such as a Kalman filter or an LSTM network), and the prediction error of the dynamic obstacle trajectory is less than 0.2m (such as a pedestrian suddenly turning), and the path safety redundant distance is reduced by 30%, thereby avoiding overly conservative detours and enabling predictive obstacle avoidance. Through the long-term reward accumulation mechanism of Q-learning, the interference of sensor instantaneous noise (such as visual misdetection of obstacles) on the path is suppressed, and the false obstacle avoidance rate is reduced to less than 5%, achieving resistance to perceptual noise.
[0047] In this embodiment, through the edge-deployed Q-learning algorithm and reinforcement learning model, the traditional path planning’s dependence on static environment and high computing power is broken through, and real-time response in dynamic scenarios, efficient learning in resource-constrained environments, and multi-objective collaborative optimization capabilities are achieved. It provides highly adaptable and low-latency autonomous navigation solutions for scenarios such as warehousing logistics (highly dynamic cargo handling) and disaster rescue (exploration of unknown ruins).
[0048] In one embodiment, Figure 2 As shown, the above robot system also includes: The human-computer interaction module 50 is provided with voice command recognition, gesture control and brain wave signal interface.
[0049] In this embodiment, the triple interactive interface of voice, gesture and brain wave signal can realize scene adaptive control. On the one hand, for example, in an open environment (such as a factory workshop), natural language instructions (such as "move to area A and grab the red part") are supported, and the recognition accuracy is greater than 95% (when the noise is less than 70dB), realizing voice command recognition; on the other hand, dynamic gestures (such as clenching fists and opening to control the robot arm to grab) are captured through binocular vision 21, which is suitable for scenes with high noise or silent operation (such as laboratories) to realize gesture control; furthermore, based on the non-invasive EEG headband (such as P300 signal), the user's intention is analyzed to help users with physical disabilities or high-risk environments (such as nuclear power plants) to achieve "contactless control". In addition, when a certain mode fails (such as voice interference by high-frequency noise), the system automatically switches to other interaction methods to ensure the continuity of instructions and has a redundant fault-tolerant mechanism.
[0050] In one embodiment, the human-computer interaction module 50 establishes a communication connection with the edge computing unit 40, the edge computing unit 40 locally processes the interaction signal, the voice command recognition delay is less than 200ms (lightweight deployment of end-to-end model), the gesture trajectory tracking frequency is ≥30Hz, and the brain wave signal analysis cycle is less than 1s, meeting the real-time control requirements; the interaction module supports multiple system protocols such as ROS / Android / IOS, and can quickly access third-party devices (such as AR glasses, smart gloves) through API.
[0051] When the voice command is ambiguous (such as "avoid that thing"), the real-time object positioning based on visual perception ("that thing" = dynamically marked obstacle) is used to correct the path planning and realize intention correction. Through voice intonation analysis (such as rapid commands triggering emergency mode) and brain wave emotion recognition (such as anxiety signals activating the safety lock), the safety of human-machine collaboration is improved and emotional interaction is realized.
[0052] In this embodiment, through the multimodal interactive fusion of voice, gesture and brain waves, the limitation of traditional robots relying on a single control method is broken through, and full-scene adaptability, highly inclusive operation (covering healthy users and disabled groups) and accurate intention analysis are achieved.
[0053] In one embodiment, the gas sensor 24 is used to detect the concentration of volatile organic compounds (VOCs).
[0054] In one embodiment, the VOC sensor has high sensitivity (lower limit of detection ≤ 1ppm) and wide detection range (0~1000ppm), and can detect dangerous gases such as benzene, formaldehyde, and methane in real time during chemical leaks, early stages of fires, or in confined spaces (such as underground pipelines), triggering an alarm 3 to 5 minutes earlier than traditional smoke sensors to identify hidden risks. Through the VOC concentration distribution map (fused with the three-dimensional map), the source of the leak (such as the rupture point of the pipeline) is located with an error of less than 0.5m, guiding robots or personnel to quickly intervene and deal with the problem, thereby realizing concentration gradient analysis.
[0055] In one embodiment, the edge computing unit 40 has a built-in reinforcement learning model for optimizing path planning through a Q-learning algorithm. At this time, when the VOC concentration exceeds a threshold (such as 50ppm), the edge computing unit 40 automatically generates a detour path to avoid high-pollution areas, and at the same time combines wind speed data to predict the direction of gas diffusion to prevent the robot from entering the diffusion path, thereby achieving a dynamic poison avoidance path; triggering the robotic arm sealing mode (such as enabling a self-healing coating to isolate VOC corrosion) or starting an air purification module (optional) to extend the operating time of the equipment in a toxic environment and achieve protection level switching.
[0056] In one embodiment, when visual identification of "transparent liquid leakage" is performed, VOC detection is used to distinguish whether it is water (harmless) or acetone (dangerous), thereby avoiding misjudgment and eliminating perceptual ambiguity; when the pressure-sensitive tactile array 22 detects an unknown viscous substance, it is inferred to be an alcohol gel based on the volatility characteristics of VOC (such as the smell of ethanol), thereby optimizing the grasping force.
[0057] In this embodiment, through the special enhancement of VOC concentration detection capability, the perception capability of the robot system in chemical hazardous environments is upgraded from "passive obstacle avoidance" to a full-chain safety mechanism of "active warning-dynamic protection-intelligent decision-making", which significantly improves the operational safety and task reliability in scenarios such as chemical inspection, post-disaster rescue (such as toxic ruins exploration) and laboratory automation.
[0058] In one embodiment, the central processor 11 supports a federated learning framework, allowing multiple robots to share local model parameters.
[0059] The central processor 11 aggregates the local model parameters of multiple robots (such as path planning Q value table, VOC recognition model) through encrypted channels. Different robots (such as factory inspection machines, medical delivery machines) share feature layer parameters, so that a single robot can obtain prior knowledge of scenes that have not been experienced (such as hospital corridors). The convergence speed of path planning in new environments is increased by 50%, and cross-scenario generalization capabilities are achieved. The local model update frequency (such as every 10 minutes) is decoupled from the global model synchronization period (such as every 2 hours) to avoid computing power fluctuations caused by frequent communications and achieve incremental model optimization.
[0060] In one embodiment, multiple robots share local parameters of the Q-value table through federated learning, which reduces the path planning time of a new robot entering a similar environment by 40%.
[0061] In this embodiment, the original data in federated learning is stored locally, and only the model gradient differential parameters are uploaded. Sensitive data such as medical robot patient location data and industrial robot production line layout information are not stored locally, meeting the requirements of regulations such as GDPR / HIPAA and having privacy compliance. Homomorphic encryption and gradient noise injection are used to prevent the original data from being reversely deduced during parameter transmission, reducing the risk of model leakage by 90% and resisting malicious attacks.
[0062] In one embodiment, the parameter compression rate is greater than 70%, and 100 robots synchronizing a 1GB model only requires 10MB of bandwidth (traditional centralized training requires 100GB), saving bandwidth. The frequency of participating in federated learning is dynamically adjusted according to the remaining power of the robots (high-power nodes take on more training tasks), and the overall battery life of the cluster is extended by 20%, achieving energy efficiency balance.
[0063] In one embodiment, the federated learning framework integrates environmental difference perception weights (for example, the weight of robots in chemical areas is higher than that in office areas). In the event of a sudden leakage, the update weight of the robot model near the accident point is increased to 80%, which quickly optimizes the hazard disposal strategy of the global model and achieves focus on hot spots. When the environmental distribution changes (such as warehouse shelf layout adjustment), incremental learning is automatically triggered by detecting the difference in local-global model losses, and the model iteration cycle is shortened to 30 minutes to cope with concept drift.
[0064] In this embodiment, through the distributed collaboration mechanism of the federated learning framework, knowledge sharing and efficient model evolution of multiple robots are achieved under data privacy protection, and bottlenecks such as slow cross-scenario adaptation, low data utilization efficiency, and weak security compliance of single learning are overcome. It provides a scalable and highly robust collaborative intelligent base for large-scale robot cluster applications such as smart cities (cross-regional inspections) and logistics warehousing (multi-warehouse linkage).
[0065] In one embodiment, Figure 2 As shown, the above robot system also includes: The self-repairing coating 60 covers the surface of the robot arm, and the material in the self-repairing coating 60 activates the repair function at a preset temperature.
[0066] In one embodiment, the microcapsules in the self-healing coating encapsulate the repair agent (such as polyurethane prepolymer), and the preset temperature (60°C) triggers the capsule to rupture and release the repair agent: for scratches or cracks with a width of ≤50μm, the repair is completed within 5 minutes after the temperature is triggered, and the tensile strength recovery rate of the coating after repair is greater than 90%, achieving micro crack repair; after the cracks are filled with the repair agent, external corrosive media (such as acid rain, VOC gas) are isolated, so that the life of the robot arm in a chemical environment is extended by 2~3 times, and the corrosion resistance is enhanced.
[0067] In one embodiment, the temperature trigger mechanism and the working condition of the robotic arm can be linked. When the robotic arm moves at high speed, the surface friction temperature can reach 60~80℃, automatically activating the repair function to repair micro-damages caused by wear in real time. In a low-temperature environment (such as polar operations), local heating to the preset temperature is achieved through the built-in heating wire to ensure the normal operation of the coating repair function and achieve external heat source assisted repair.
[0068] In one embodiment, the coating uses an elastomeric substrate (such as silicone rubber) and a flexible match with the repair agent. When the robotic arm joint is bent, the coating can withstand a tensile deformation of >200% without peeling off, and the flexibility retention rate after repair is >95%, achieving the technical effect of bending without breaking; the repair agent is chemically bonded to the substrate, and the repaired area still has no stratification after 100,000 bending cycles, thereby achieving overall interface adhesion enhancement.
[0069] In one embodiment, autonomous repair replaces manual maintenance. For minor damage (accounting for more than 70% of robot arm failures), there is no need to shut down for disassembly and repair, reducing annual maintenance costs by 50% and reducing manual intervention. In scenarios such as disaster relief, the robot arm can operate continuously for more than 500 hours without surface maintenance, reducing the risk of mission interruption by 80%, and ensuring continuous operation.
[0070] In this embodiment, through the design of a temperature-responsive self-healing coating, real-time damage repair, corrosion resistance enhancement and long-life operation of the robotic arm in complex environments are achieved, which overcomes the bottlenecks of performance degradation and frequent maintenance caused by surface damage of traditional robotic arms, and provides a highly reliable and low-maintenance robot protection solution for scenarios such as chemical inspection, space exploration (extreme temperature fluctuations) and deep-sea operations (high salt corrosion).
[0071] In addition, if Figure 3 As shown, an autonomous navigation method is also provided, using the above robot system, the autonomous navigation method includes: S1: Build a 3D map of the environment through a multimodal perception module and mark dynamic obstacles.
[0072] In one embodiment, centimeter-level map accuracy: binocular vision point cloud resolution reaches ±2cm, pressure-sensitive tactile array 22 supplements visual blind spots (such as transparent glass), dynamic obstacle marking error is <5cm, achieving centimeter-level map accuracy; dynamic obstacle tracking frequency is ≥10Hz (such as moving pedestrians and vehicles), supporting obstacle avoidance response in high-speed motion scenes (robot moving speed ≤2m / s), and real-time environmental updates.
[0073] S2: Use the improved A* algorithm to calculate the initial path and optimize the path weight in combination with the reinforcement learning model.
[0074] In one embodiment, the improved A* algorithm and Q-learning model are used to collaboratively optimize the path weight, and the comprehensive optimization of the path length, energy consumption and safety factor improves the "safety-efficiency" balance of the chemical park inspection path by 40%. The LSTM obstacle trajectory prediction model is integrated (prediction time 3-5 seconds), and the path safety redundant distance is reduced by 50%, avoiding excessive detours and realizing dynamic obstacle prediction.
[0075] S3: Adjust the movement speed and angle of the robotic arm according to the real-time stress feedback of the flexible joint.
[0076] In one embodiment, the flexible joint stress feedback and the real-time control closed loop of the edge computing unit are implemented. When the joint stress exceeds a threshold (such as 20MPa), the robot arm's movement speed is automatically reduced by 50% to prevent structural damage and achieve overload protection. The joint angle is dynamically adjusted according to the stress distribution to increase the robot arm's movement smoothness by 60% in narrow spaces (such as pipes), reduce the risk of jamming, and achieve bionic motion optimization.
[0077] In this embodiment, through the full-link fusion of environmental perception-algorithm decision-making-execution control, high-precision modeling, intelligent path optimization and safe motion control in dynamic and complex scenarios are achieved, which solves the problem of the traditional navigation method being separated from the chain of "perception blind spots-planning rigidity-control lag", and provides an autonomous navigation solution with both real-time and robustness for scenarios such as smart manufacturing (flexible production line logistics) and disaster rescue (ruin search and rescue).
[0078] In one embodiment, path weight optimization introduces an obstacle movement prediction model.
[0079] In one embodiment, the obstacle movement prediction model (such as Kalman filter, LSTM time series network) is integrated with the improved A* algorithm, and the position prediction error of dynamic obstacles such as pedestrians and AGVs is less than 0.3m (within 1 second), the speed prediction error is less than 15%, the path avoidance distance is reduced by 40%, the detour redundancy is reduced, and the trajectory prediction accuracy is improved; it supports tracking more than 10 dynamic obstacles at the same time, predicts the conflict time window (TTC, Time to Collision), and prioritizes avoiding high-risk targets (such as fast-moving objects with TTC < 2 seconds, realizing collaborative prediction of multiple obstacles.
[0080] Among them, the LSTM timing network is Long Short-Term Memory (Long Short-Term Memory Network).
[0081] In one embodiment, the prediction model outputs a probability distribution map of the future trajectory of the obstacle, quantifies the dynamic risk weight, and in the heuristic function of the A* algorithm, increases the path cost of high-probability conflict areas (such as corridor intersections) by 3 to 5 times, guides the robot to choose a low-conflict probability path, and generates a risk-sensitive path; when the obstacle suddenly turns or accelerates (such as a pedestrian suddenly running), the model updates the weight matrix every 0.5 seconds, and the path adjustment response time is less than 200ms, realizing real-time weight update.
[0082] In one embodiment, the lightweight prediction model is deployed on the edge computing unit 40. The calculation amount of single obstacle prediction is less than 10 MFLOPS, and the CPU occupancy rate is less than 25% when 10 obstacles are predicted in parallel, which meets the real-time requirements of edge devices and achieves low computing power consumption; through path cost smoothing filtering (such as exponentially weighted moving average), the path oscillation caused by instantaneous prediction noise is suppressed, and the smoothness of the planned trajectory is improved by 70%, achieving anti-prediction jitter.
[0083] In one embodiment, the prediction model is linked with multimodal perception (vision, sound source positioning) data, and the obstacle position is detected through binocular vision. The predicted trajectory is corrected in combination with the sound source direction of the microphone array 23 (such as vehicle horns), so as to solve the prediction failure problem caused by visual occlusion; a dedicated prediction model is pre-trained for typical scenarios to predict the fixed route movement mode of the AGV and plan the intersection avoidance strategy in advance; the crowd flow trend (such as tidal flow of people at subway exits) is identified to generate the global optimal detour path.
[0084] In this embodiment, through the deep binding of the dynamic obstacle prediction model and the path weight, the traditional path planning is upgraded from "passive obstacle avoidance" to an intelligent decision-making mode of "active prediction-dynamic optimization", which significantly improves the navigation efficiency and safety in high-dynamic scenarios (such as transportation hubs and flexible production lines). At the same time, it ensures the stable operation of the algorithm in resource-constrained environments, and provides core technical support for applications such as autonomous mobile robots (AMRs) and unmanned delivery vehicles.
[0085] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0086] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0087] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A robot system based on multimodal perception and collaboration, characterized in that: include: A main frame, wherein a central processing unit is integrated inside the main frame; A multimodal perception module, which includes a binocular vision camera, a pressure-sensitive tactile array, a microphone array, and a gas sensor, for collecting environmental data; The flexible joint robot arm is composed of at least three bionic joints connected in series, each of which has a built-in shape memory alloy driver and a pneumatic feedback device; The edge computing unit establishes communication connections with the central processing unit and the multimodal perception module respectively, and is used to process the perception data in real time and generate action instructions.
2. The robot system according to claim 1, characterized in that: The pressure-sensitive tactile array is provided with a graphene film and a micro-current sensor.
3. The robot system according to claim 1, characterized in that: The driving mode of the bionic joint is the composite control of shape memory alloy and pneumatic muscle.
4. The robot system according to claim 1, characterized in that: The edge computing unit has a built-in reinforcement learning model and optimizes path planning through a Q-learning algorithm.
5. The robot system according to claim 1, characterized in that: Also includes: The human-computer interaction module is equipped with voice command recognition, gesture control and brain wave signal interface.
6. The robot system according to claim 1, characterized in that: The gas sensor is used to detect the concentration of volatile organic matter.
7. The robot system according to claim 1, characterized in that: The central processor supports a federated learning framework, allowing multiple robots to share local model parameters.
8. The robot system according to claim 1, characterized in that: Also includes: The self-healing coating is covered on the surface of the robot arm. The material in the self-healing coating activates the repair function at a preset temperature.
9. An autonomous navigation method, using the robot system according to any one of claims 1 to 8, the autonomous navigation method comprising: S1: Build a 3D map of the environment through a multimodal perception module and mark dynamic obstacles; S2: Use the improved A* algorithm to calculate the initial path and optimize the path weight in combination with the reinforcement learning model; S3: Adjust the movement speed and angle of the robotic arm according to the real-time stress feedback of the flexible joint.
10. The autonomous navigation method according to claim 9, characterized in that: The path weight optimization in S2 introduces an obstacle movement prediction model.
Citation Information
Patent Citations
Variable-stiffness self-repairing material containing metal-sulfydryl coordinate bond as well as preparation and application thereof
CN113817173A
Flexible mechanical arm based on coupling of memory alloy and pneumatic artificial muscle
CN116038679A
Scheduling method, device and system for mobile edge computing and storage medium
CN116737361A
Multi-perception fusion bionic search robot with man-machine bidirectional interaction function
CN117021112A
Self-adaptive multi-mode sensor fusion method and system for robot navigation and obstacle avoidance
CN118443000A
Cited By
Patrol robot path optimization method and system based on dynamic obstacle avoidance
CN120255529A
Industrial robot safety cooperation method fusing multi-modal data and reinforcement learning
CN120941418A
Industrial robot remote debugging system and method based on edge computing
CN121340291A
An edge-computing-based industrial robot remote debugging system and method
CN121340291B