Industrial robot autonomous collaborative decision-making method and system based on multi-modal perception and medium

By synchronously collecting multimodal information in industrial robots, cross-modal space-time alignment is achieved, and combined with dynamic weight adjustment and deep learning optimization frameworks, the accuracy and reliability of multimodal perception decisions in the existing technology are solved, and more efficient and stable autonomous collaborative decision-making is achieved.

CN120190832AActive Publication Date: 2025-06-24SHENZHEN HUAZHONG NUMERICAL CONTROL

Patent Information

Application Number
CN202510672293.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing multimodal perception technology of industrial robots has problems such as space-time dislocation of data, low resource utilization efficiency, single optimization dimensions and weak system fault tolerance, making it difficult to achieve accurate and reliable autonomous collaborative decision-making.

Method used

By synchronously collecting visual, tactile and auditory information, time synchronization and coordinate conversion technology are used to achieve cross-modal space-time alignment; task-driven dynamic weight adjustment mechanism and attention calculation model are introduced to realize adaptive generation of feature fusion parameters; deep learning is used to build a multi-objective optimization framework, comprehensively considering indicators such as energy consumption, accuracy, and safety; and setting up a double-layer fault tolerance mechanism to ensure system stability.

Benefits of technology

It realizes the spatial and temporal integration of multimodal data, dynamically adjusts sensor weights, and comprehensively optimizes multi-objective decisions, which significantly improves the accuracy and reliability of autonomous collaborative decision-making of industrial robots, and improves the stability and fault tolerance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120190832A_ABST
    Figure CN120190832A_ABST
Patent Text Reader

Abstract

The invention provides an industrial robot autonomous collaborative decision-making method and system based on multi-modal sensing and a medium, and belongs to the technical field of industrial robot intelligent control. The method comprises the steps of performing cross-modal space-time alignment to eliminate data space-time differences by collecting visual, tactile and auditory information, realizing multi-modal feature fusion in combination with a dynamic weight adjustment mechanism and an attention calculation model, and generating a joint decision strategy through a deep learning optimization model. The system dynamically allocates sensor weights according to task types, introduces a multi-objective optimization mechanism of energy consumption, precision and safety, and sets a fault-tolerant rule to automatically recover the weights or recalibrate the sensors. According to the method, the problems of rigid data fusion and single optimization dimension in traditional multi-modal decision making are solved, and the precision, the response speed and the environmental adaptability of collaborative operation of the industrial robot are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent control of industrial robots. Specifically, it involves an autonomous collaborative decision-making method, system, and medium for industrial robots based on multi-modal perception. Background Art

[0002] In the field of multi-modal perception of industrial robots, there are significant deficiencies in the existing technologies. First, the fusion of multi-source data is insufficient. Due to the differences in the acquisition frequencies and spatial coordinate systems of visual, tactile, and auditory sensors, data spatio-temporal misalignment occurs, seriously affecting the accuracy of decision-making. Second, existing methods mostly adopt fixed weight allocation strategies and cannot dynamically adjust the priorities of various sensors according to the task type, resulting in low resource utilization efficiency. In addition, traditional optimization models often focus on a single indicator and ignore the multi-objective collaborative optimization of energy consumption, safety, etc., making it difficult to adapt to complex and changing industrial scenarios. Finally, the system fault tolerance ability is weak, lacking a real-time verification mechanism for decision-making effects, and unable to quickly backtrack or calibrate sensors when the task fails. These problems jointly restrict the performance improvement of autonomous collaborative decision-making of industrial robots. As the application scenarios of robots develop towards complexity, dynamics, and diversification, single-modal perception is difficult to meet the needs of autonomous collaborative decision-making of robots. There is an urgent need for a method and system that can fuse multiple modal information to achieve more accurate and reliable decision-making. Summary of the Invention

[0003] The purpose of this application is to provide an autonomous collaborative decision-making method, system, and medium for industrial robots based on multi-modal perception. By synchronously collecting visual, tactile, and auditory information and using time synchronization and coordinate transformation technologies to achieve cross-modal spatio-temporal alignment, the problem of data misalignment is solved. For the weight allocation problem, a task-driven dynamic weight adjustment mechanism is introduced, combined with an attention calculation model, to realize the adaptive generation of feature fusion parameters. A multi-objective optimization framework is constructed using deep learning, comprehensively considering indicators such as energy consumption, accuracy, and safety, to generate joint decisions and iteratively optimize. At the same time, a two-layer fault tolerance mechanism is set up to ensure the stability of the system and promote the intelligent upgrade of industrial robots.

[0004] This application provides an autonomous collaborative decision-making method for industrial robots based on multi-modal perception, including the following steps: Collect the perception information within the preset period of the autonomous collaborative system, including visual information, tactile information, and auditory information; Perform cross-modal spatio-temporal alignment on the visual information, tactile information, and auditory information through a preset method to generate first perception feature information; Extract visual feature data, tactile feature data, and auditory feature data from the first perception feature information through a preset mechanism and perform dynamic weight adjustment to generate feature fusion parameters; Generate a multi-modal joint decision-making strategy based on the feature fusion parameters, and iteratively optimize the decision-making parameters through a preset optimization model; Obtain the decision parameter optimization effect data and compare it with the preset effect threshold to determine whether the optimization effect meets the requirements.

[0005] Among them, in the multi-modal perception-based industrial robot autonomous collaborative decision-making method described in this application, the cross-modal spatio-temporal alignment includes: Precisely align the images captured by the camera, the force changes detected by the tactile sensor, and the sounds collected by the microphone according to the acquisition time through the time synchronization module; Unify the position information detected by different sensors into the three-dimensional space coordinate system where the robot is located through coordinate transformation.

[0006] Among them, in the multi-modal perception-based industrial robot autonomous collaborative decision-making method described in this application, the weight dynamic adjustment includes: Automatically set the importance ratio of each sensor according to the current task type. In the fine assembly task, the weight of the tactile sensor is set to the first weight ratio; When it is necessary to detect machine abnormalities, the weight of the sound sensor is increased to the second weight ratio, and at the same time, the weight of the camera image is reduced to the third weight ratio.

[0007] Among them, in the multi-modal perception-based industrial robot autonomous collaborative decision-making method described in this application, the generation process of the feature fusion parameters is as follows: Analyze the correlation degree between visual, tactile, and auditory data through an attention calculation model; Automatically allocate the comprehensive influence values of each sensor according to the correlation degree, ensuring that the sum of the influence values of vision, touch, and hearing is 1.

[0008] Among them, in the multi-modal perception-based industrial robot autonomous collaborative decision-making method described in this application, the preset optimization model uses a deep learning model, where: The moving state of the robot, the surrounding environment conditions, and the task objectives together constitute the decision-making basis; The output control instructions include moving route parameters, operation force parameters, and cooperation timing parameters; Continuously optimize the decision-making results through the comprehensive scores of energy consumption, operation accuracy, and safety.

[0009] Among them, in the multi-modal perception-based industrial robot autonomous collaborative decision-making method described in this application, when comparing the optimization effects: When the task completion accuracy rate is lower than the preset first completion rate or the response delay exceeds the preset time, automatically restore the previous weight settings; If the decision fails to meet the standard three times in a row, all sensors are recalibrated and the data fusion rules are updated.

[0010] In a second aspect, the present application provides an industrial robot autonomous collaborative decision-making system based on multi-modal perception. The system includes: a memory and a processor. The memory includes a program of an industrial robot autonomous collaborative decision-making method based on multi-modal perception. When the program of the industrial robot autonomous collaborative decision-making method based on multi-modal perception is executed by the processor, the following steps are implemented: Collect the perception information within a preset period of the autonomous collaborative system, including visual information, tactile information, and auditory information; Perform cross-modal spatio-temporal alignment on the visual information, tactile information, and auditory information through a preset method to generate first perception feature information; Extract visual feature data, tactile feature data, and auditory feature data based on the first perception feature information through a preset mechanism and perform dynamic weight adjustment to generate feature fusion parameters; Generate a multi-modal joint decision-making strategy based on the feature fusion parameters, and iteratively optimize the decision parameters through a preset optimization model; Obtain the decision parameter optimization effect data and compare it with a preset effect threshold to determine whether the optimization effect meets the requirements.

[0011] Among them, in the industrial robot autonomous collaborative decision-making system based on multi-modal perception described in the present application, the cross-modal spatio-temporal alignment includes: Precisely align the images captured by the camera, the force changes detected by the tactile sensor, and the sounds collected by the microphone according to the acquisition time through the time synchronization module; Unify the position information detected by different sensors into the three-dimensional space coordinate system where the robot is located through coordinate transformation.

[0012] Among them, in the industrial robot autonomous collaborative decision-making system based on multi-modal perception described in the present application, the dynamic weight adjustment includes: Automatically set the importance ratio of each sensor according to the current task type. In the fine assembly task, the weight of the tactile sensor is set to the first weight ratio; When it is necessary to detect machine abnormalities, the weight of the sound sensor is increased to the second weight ratio, and at the same time, the weight of the camera image is reduced to the third weight ratio.

[0013] In a third aspect, the present application also provides a computer-readable storage medium, which includes a program for the method of autonomous collaborative decision-making of industrial robots based on multi-modal perception. When the program for the method of autonomous collaborative decision-making of industrial robots based on multi-modal perception is executed by a processor, the steps of the method of autonomous collaborative decision-making of industrial robots based on multi-modal perception as described in any one of the above are implemented.

[0014] As can be seen from the above, the method, system, and medium for autonomous collaborative decision-making of industrial robots based on multi-modal perception provided by the embodiments of the present application synchronously collect visual, tactile, and auditory information, and use a time synchronization module and coordinate transformation technology to achieve cross-modal spatio-temporal alignment, fundamentally solving the problem of data spatio-temporal misalignment. Aiming at the defect of rigid weight allocation, the system introduces a task-driven dynamic weight adjustment mechanism, preferentially uses tactile data in fine assembly tasks, enhances the weight of auditory sensors during anomaly detection, and combines an attention calculation model to analyze the multi-modal correlation degree to realize the adaptive generation of feature fusion parameters. To solve the problem of single optimization dimension, a deep learning model is used to construct a multi-objective optimization framework, comprehensively considering energy consumption, accuracy, and safety indicators to generate joint decisions on movement routes, operation forces, and collaboration timing, and continuously improving decision parameters through iterative optimization. The system also innovatively sets a double-layer fault tolerance mechanism: when the task accuracy rate is lower than the threshold or the response times out, it automatically reverts to the historical weight configuration; after three consecutive decision failures, a full calibration of sensors and rule update are triggered, significantly improving the system stability and effectively promoting the intelligent upgrade of industrial robots.

[0015] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the method for autonomous collaborative decision-making of industrial robots based on multi-modal perception provided by the embodiments of the present application; Figure 2 It is a flowchart of cross-modal spatio-temporal alignment of the method for autonomous collaborative decision-making of industrial robots based on multi-modal perception provided by the embodiments of the present application; Figure 3Flowchart of weight dynamic adjustment for the multi-modal perception-based industrial robot autonomous collaborative decision-making method provided by the embodiments of the present application; Figure 4 Flowchart of generating feature fusion parameters for the multi-modal perception-based industrial robot autonomous collaborative decision-making method provided by the embodiments of the present application. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0019] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0020] Please refer to Figure 1 , Figure 1 which is a flowchart of the multi-modal perception-based industrial robot autonomous collaborative decision-making method in some embodiments of the present application. The multi-modal perception-based industrial robot autonomous collaborative decision-making method is used in terminal devices such as computers and mobile phone terminals. The multi-modal perception-based industrial robot autonomous collaborative decision-making method includes the following steps: S101. Collect the perception information within the preset period of the autonomous collaborative system, including visual information, tactile information, and auditory information; S102. Align the visual information, tactile information, and auditory information in a cross-modal spatio-temporal manner through a preset method to generate first perception feature information; S103. Extract visual feature data, tactile feature data, and auditory feature data based on the first perception feature information through a preset mechanism, and perform dynamic weight adjustment to generate feature fusion parameters; S104. Generate a multi-modal joint decision-making strategy based on the feature fusion parameters, and iteratively optimize the decision-making parameters through a preset optimization model; S105. Obtain the data of the optimization effect of the decision-making parameters and compare it with a preset effect threshold to determine whether the optimization effect meets the requirements.

[0021] Among them, first, multi-modal perception information within a preset period is synchronously collected through multi-source sensors, including visual information (such as the workpiece pose and environmental images captured by a high-resolution camera), tactile information (such as the contact force / moment of the end effector detected by a six-axis force sensor), and auditory information (such as the abnormal mechanical transmission audio collected by a directional microphone). Then, cross-modal spatio-temporal alignment technology is used to uniformly process heterogeneous data to achieve timestamp synchronization of sensor data and eliminate the timing deviation caused by sampling frequency differences. At the same time, based on the hand-eye calibration matrix and the robot kinematic model, the visual image coordinate system and the tactile sensor coordinate system are uniformly transformed into the robot base coordinate system to construct a three-dimensional spatial data mapping relationship, solve the problem of inconsistent spatial references, and generate the first perception feature information that is spatio-temporally consistent. Subsequently, a dynamic weight adaptive adjustment mechanism is introduced. High-order features of each modality are extracted using a deep neural network, and combined with the real-time task type (such as assembly, inspection, anomaly handling) and scene complexity (light intensity, noise level), the cross-modal correlation weight is calculated through an attention mechanism, and finally the optimized feature fusion parameters are generated. Based on the fusion parameters, a multi-objective joint decision-making model is constructed. The robot motion state, environmental constraints, and task objectives are encoded into the state space using a deep reinforcement learning framework, and a decision instruction set including path planning parameters, operation force thresholds, and cooperation timings is output, and iterative optimization is performed through a multi-objective optimization function until the decision score reaches the convergence condition. Finally, the system monitors the decision execution effect in real time (such as task completion time, positioning deviation, energy consumption data), and compares and analyzes it with preset thresholds (such as accuracy threshold ±0.1mm, response delay upper limit 200ms): If the optimization effect does not meet the standard, the weight backtracking mechanism is triggered or the sensor online calibration process is started to form a perception-decision-validation closed-loop control to ensure the robustness and self-adaptability of the system in a dynamic industrial environment.

[0022] Please refer to Figure 2 , Figure 2 is a flowchart of cross-modal spatio-temporal alignment of an industrial robot autonomous collaborative decision-making method based on multi-modal perception in some embodiments of the present application. According to an embodiment of the present invention, the cross-modal spatio-temporal alignment includes: S201. Precisely align the images captured by the camera, the force changes detected by the tactile sensor, and the sounds collected by the microphone according to the acquisition time through the time synchronization module; S202. Unify the position information detected by different sensors into the three-dimensional space coordinate system where the robot is located through coordinate transformation.

[0023] Among them, at the time synchronization level, a hardware clock synchronization module using the Precision Time Protocol is adopted to align the data acquisition timestamps of the camera, the tactile sensor (such as a six-axis force sensor), and the microphone, and an interpolation compensation algorithm is used to eliminate the timing deviation caused by the difference in the sensor sampling frequencies, ensuring that the visual images, tactile force feedback, and collision sound data of the same event (such as the moment when the robotic arm touches the workpiece) are strictly synchronized. At the space alignment level, a multi-level coordinate system transformation model is established. First, the external parameter matrix of the camera is obtained through hand-eye calibration to map the image pixel coordinates to the robot base coordinate system; at the same time, the detected contact force / moment is transformed from the tool coordinate system to the base coordinate system by using the installation pose matrix of the tactile sensor; for audio localization, the azimuth angle of the sound source is estimated by the direction of arrival of the microphone array, and the three-dimensional coordinates of the sound source are calculated by combining the robot kinematic model. Finally, all sensor data are unified into the three-dimensional space coordinate system for robot motion control through homogeneous coordinate transformation, forming a spatio-temporal consistent multi-modal perception benchmark. Through the above dual alignment mechanism, the problem of data fusion distortion caused by spatio-temporal reference differences in traditional multi-sensor systems is completely solved, providing high-precision input for subsequent dynamic weight allocation and collaborative decision-making.

[0024] Please refer to Figure 3 , Figure 3 which is the flowchart of the dynamic weight adjustment of the autonomous collaborative decision-making method of the industrial robot based on multi-modal perception in some embodiments of this application. According to the embodiments of the present invention, the dynamic weight adjustment includes: S301. Automatically set the importance ratio of each sensor according to the current task type, and set the weight of the tactile sensor to the first weight ratio in the fine assembly task; S302. When it is necessary to detect machine anomalies, the weight of the sound sensor is increased to the second weight ratio, and at the same time, the weight of the camera image is reduced to the third weight ratio.

[0025] Among them, the dynamic weight allocation mechanism realizes the intelligent allocation of multi-modal perception resources through task type recognition and sensor effectiveness evaluation. The system is built with a task semantic parsing module that extracts keywords and classifies the intent of task instructions based on natural language processing, and generates task type tags in combination with scene perception data (such as workpiece size tolerance ±0.05mm, ambient noise level >75dB). For precision assembly tasks, the system calls a preset priority rule library to increase the weight of the tactile sensor to the first weight ratio (such as 60%-70%), preferentially uses six-axis force sensor data to analyze the contact force feedback, and at the same time reduces the visual weight to the auxiliary level, only retaining the key pose verification function; when switching to the equipment anomaly diagnosis mode, the frequency domain analysis module is used to identify the acoustic signal characteristics, dynamically increase the weight of the sound sensor to the second weight ratio (such as 65%-75%), and uses wavelet denoising and MFCC feature extraction to enhance the acoustic fingerprint recognition accuracy, while reducing the visual weight to the third weight ratio (such as 10%-15%) to reduce the computational overhead. The weight allocation process is continuously optimized through a reinforcement learning model, and the weight boundary values are dynamically corrected in combination with historical operation data (such as assembly success rate, anomaly detection accuracy) to ensure the precise matching of sensor resource allocation and task requirements, which can improve the task execution efficiency compared with the traditional fixed weight scheme.

[0026] Please refer to Figure 4 , Figure 4 is a flowchart for generating feature fusion parameters of the industrial robot autonomous collaborative decision-making method based on multi-modal perception in some embodiments of this application. According to the embodiments of the present invention, the generation process of the feature fusion parameters is as follows: S401. Analyze the correlation degree between visual, tactile, and auditory data through an attention calculation model; S402. Automatically allocate the comprehensive influence values of each sensor according to the correlation degree, ensuring that the sum of the influence values of vision, touch, and hearing is 1.

[0027] Among them, the multi-modal correlation analysis and weight allocation method of the present invention realizes cross-modal data collaborative optimization based on an improved multi-head attention mechanism. Specifically, first, a cross-modal feature interaction matrix is constructed. Visual data extracts spatial features through a CNN, tactile data encodes temporal mechanical features through an LSTM, and auditory data extracts frequency domain features using a Mel spectrogram convolutional network. Subsequently, the self-attention calculation module models the internal feature correlation degree of a single modality, and uses the cross-attention branch to calculate the cross-modal interaction weights of vision-tactile, vision-auditory, and tactile-auditory to form a correlation score matrix. Based on this, a dynamic normalization weight allocator is designed, and finally, the softmax function is used to constrain the weights of vision, touch, and hearing to satisfy the sum of the influence values to be 1. Compared with the traditional weighted average strategy, this method can improve the multi-modal data fusion efficiency and avoid the problem of weight overfitting through preset normalization constraints.

[0028] According to an embodiment of the present invention, the preset optimization model adopts a deep learning model, where: The moving state of the robot, the surrounding environment conditions, and the task objectives together constitute the decision-making basis; The output control instructions include moving route parameters, operation force parameters, and cooperation timing parameters; The decision-making results are continuously optimized through the comprehensive evaluation of energy consumption, operation accuracy, and safety.

[0029] Among them, the preset optimization model of the present invention is constructed based on the deep reinforcement learning framework, and realizes intelligent decision-making through multi-dimensional state perception and multi-objective collaborative optimization. This model uses the real-time moving state of the robot, the surrounding environment information, and the task objective requirements as the core decision-making basis. The moving state of the robot covers dynamic data such as its current position, speed, and posture; the surrounding environment conditions include information such as the distribution of obstacles, the positions of other devices, and environmental changes; the task objective clarifies the specific operation requirements and the expected achieved effects. Based on these comprehensive decision-making bases, the model outputs control instructions including moving route parameters, operation force parameters, and cooperation timing parameters to accurately control the actions of the robot. At the same time, to ensure the scientificity and efficiency of decision-making, the model constructs a comprehensive evaluation system for energy consumption, operation accuracy, and safety, conducts multi-dimensional quantitative evaluation on each decision-making result, and uses the evaluation feedback to optimize the decision-making parameters to achieve continuous iterative improvement of the decision-making results, enabling the industrial robot to make more energy-saving, accurate, and safe decisions in a complex and changeable working environment.

[0030] According to an embodiment of the present invention, when comparing the optimization effects: When the task completion accuracy rate is lower than the preset first completion rate or the response delay exceeds the preset time, automatically restore the previous weight setting; If the decision-making fails to meet the standard three times in a row, recalibrate all sensors and update the data fusion rules.

[0031] Among them, when evaluating and comparing the optimization effect of decision-making parameters, the present invention sets up a dual safeguard mechanism to ensure the stability and reliability of the decision-making system of industrial robots. When the task completion accuracy rate is lower than the preset first completion rate standard, or the system response delay exceeds the preset time threshold, it indicates that the current decision-making parameters and weight configuration fail to meet the task requirements. At this time, the system will automatically restore to the weight setting that has been verified effective in the previous stage, quickly correct the decision-making deviation, and maintain the operation stability. If the decision-making results fail to reach the preset standard for three consecutive times, it is determined that there may be problems with sensor data deviation or data fusion rule failure in the system. The system will trigger a comprehensive calibration program to recalibrate all sensors and update the data fusion rules simultaneously. By deeply optimizing the perception and decision-making links, the problem of unqualified decision-making is solved from the root, thereby constructing a dynamic feedback and self-repair mechanism, effectively improving the fault tolerance ability and continuous operation efficiency of the autonomous collaborative decision-making system of industrial robots under complex working conditions.

[0032] According to an embodiment of the present invention, the optimization process of the comprehensive score further includes: Real-time monitoring of the environmental noise level and light intensity. When the environmental noise exceeds the preset noise threshold, dynamically increase the band-pass filtering intensity of the auditory sensor and reduce the visual weight; When the light intensity is lower than the preset intensity threshold, activate the infrared supplementary light module and increase the depth of the visual feature extraction network, and at the same time increase the tactile weight to compensate for the impact of the decline in visual perception ability.

[0033] Among them, in the optimization process of the comprehensive score, the system real-time monitors the environmental noise level and light intensity parameters through the environmental perception module. When it is detected that the environmental noise exceeds the preset noise threshold (such as 70 dB), the system automatically triggers a dynamic noise reduction mechanism: on the one hand, by enhancing the band-pass filtering intensity of the auditory sensor, signal enhancement processing is performed on specific noise frequency bands to improve the recognition accuracy of abnormal sounds; on the other hand, based on the multi-modal weight allocation strategy, the decision-making weight of the visual sensor is synchronously reduced to reduce the misjudgment impact of visual information under noise interference. When it is monitored that the light intensity is lower than the preset intensity threshold (such as 150 lux), the system starts the environmental adaptability compensation mechanism. First, activate the infrared supplementary light module to perform spectral compensation on the visual acquisition area, and at the same time increase the depth of the convolutional layer of the visual feature extraction network to improve the feature analysis ability of low-illumination images; at the same time, dynamically increase the decision-making weight of the tactile sensor to 1.2 - 1.5 times the reference value, and compensate for the information loss caused by the decline in visual perception ability by strengthening the decision-making participation of tactile feedback. This multi-modal dynamic adjustment mechanism based on environmental parameter adaptation effectively guarantees the perception robustness and decision-making reliability of industrial robots under complex working conditions.

[0034] According to an embodiment of the present invention, it further includes: Collect the working environment temperature of the robot during the dynamic adjustment of the weights; When the working environment temperature is greater than the first temperature threshold, increase the sampling frequency of the tactile sensor to compensate for the material deformation error caused by high temperature, and reduce the visual weight to avoid image distortion interference; When the working environment temperature is less than the second temperature threshold, enable the low-frequency enhancement mode of the auditory sensor to improve the recognition of abnormal metal sounds, and increase the tactile weight to enhance the force control accuracy at low temperatures.

[0035] Among them, during the dynamic adjustment of the weights, the system collects the working environment temperature data in real time through the temperature sensor integrated in the robot body, and triggers the differential multi-modal perception optimization strategy according to the preset first temperature threshold (such as 45 °C) and the second temperature threshold (such as 5 °C). When it is detected that the environmental temperature is higher than 45 °C, the system starts the high-temperature compensation mechanism, increases the sampling frequency of the tactile sensor from the reference value, for example, from 100 Hz to 200 Hz, and compensates for the deviation of the contact force measurement caused by the thermal expansion of the metal material through high-frequency data acquisition; at the same time, the visual weight is reduced to reduce the interference of image distortion caused by high-temperature air waves on the workpiece positioning accuracy. When the environmental temperature is lower than 5 °C, the system activates the low-temperature enhancement mode, enables low-frequency band-pass filtering for the auditory sensor, strengthens the feature extraction of abnormal metal cold brittle fracture sounds, and at the same time increases the reference value of the tactile weight, and offsets the mechanical transmission hysteresis effect caused by low temperature by increasing the decision-making ratio of force feedback. This temperature adaptive adjustment mechanism keeps the comprehensive perception accuracy of the robot meeting the requirements in extreme working conditions such as foundry workshops (high temperature) and cold chain warehouses (low temperature) by dynamically balancing the performance attenuation of multi-modal sensors and scene requirements, which is significantly better than the level of traditional fixed parameter schemes.

[0036] According to the embodiment of the present invention, when generating the multi-modal joint decision-making strategy, it further includes: If the task completion degree of a certain robot is lower than the preset progress threshold ratio, automatically increase its path planning priority by 2 levels, and other robots avoid its working area; When the target positions of multiple robots overlap, dynamically allocate the passing permissions according to the task urgency; Calculate the remaining power of each robot in real time, and reduce the motion speed weight of the robot with the power lower than the preset power threshold to extend the cooperation cycle.

[0037] Among them, when generating a multi-modal joint decision-making strategy, the system evaluates the task completion status of each robot in real time through a task progress monitoring module. When it is detected that the task completion degree of a certain robot is lower than the preset progress threshold (such as 30%), the lag compensation mechanism is automatically triggered, the path planning priority of this robot is increased by 2 levels (for example, from the default priority 3 to priority 1), and its working area is marked as a high-priority passage area. The remaining robots adjust their paths in real time based on the obstacle avoidance planning model and detour around the lagging robot's working area at a safe distance. For the conflict event of overlapping target positions of multiple robots, the system calls the task urgency evaluation model and dynamically allocates passage rights according to the preset rule base (such as the task urgency weight for dangerous goods handling is 0.9 and for ordinary assembly tasks is 0.4). The high-urgency robot obtains the right of way and demarcates an exclusive path through the laser navigation module. At the same time, the system integrates a power monitoring unit to calculate the remaining power of each robot in real time. When it is detected that the power is lower than the preset threshold (such as 20%), the energy consumption balancing strategy is started, the reference value of the motion speed weight coefficient of the low-power robot is reduced, its moving speed is re-planned through the kinematic model, and the collaborative task cycle is extended (such as adjusting the handling cycle from 5 minutes per time to 5.5 - 5.8 minutes per time), which extends the continuous working time of the robot cluster by more than 2.3 hours while avoiding task interruption. This multi-dimensional collaborative optimization mechanism improves the overall production capacity of the multi-robot assembly line, reduces the task conflict rate, and decreases the low-power shutdown events.

[0038] The present invention also discloses an industrial robot autonomous collaborative decision-making system based on multi-modal perception, including a memory and a processor. The memory includes an industrial robot autonomous collaborative decision-making method program based on multi-modal perception. When the industrial robot autonomous collaborative decision-making method program based on multi-modal perception is executed by the processor, the following steps are implemented: Collect the perception information within the preset period of the autonomous collaborative system, including visual information, tactile information, and auditory information; Align the visual information, tactile information, and auditory information across modalities in space and time through a preset method to generate first perception feature information; Extract visual feature data, tactile feature data, and auditory feature data according to the first perception feature information through a preset mechanism and perform dynamic weight adjustment to generate feature fusion parameters; Generate a multi-modal joint decision-making strategy according to the feature fusion parameters and iteratively optimize the decision parameters through a preset optimization model; Obtain the decision parameter optimization effect data and compare it with the preset effect threshold to judge whether the optimization effect meets the requirements.

[0039] Among them, first, multi-modal perception information within a preset period is synchronously collected by multi-source sensors, including visual information (such as the workpiece pose and environmental images captured by a high-resolution camera), tactile information (such as the contact force / moment of the end effector detected by a six-axis force sensor), and auditory information (such as the abnormal mechanical transmission audio collected by a directional microphone); then, a cross-modal spatio-temporal alignment technology is adopted to uniformly process heterogeneous data, realize the timestamp synchronization of sensor data, and eliminate the timing deviation caused by the sampling frequency difference; at the same time, based on the hand-eye calibration matrix and the robot kinematic model, the visual image coordinate system and the tactile sensor coordinate system are uniformly transformed into the robot base coordinate system to construct a three-dimensional space data mapping relationship, solve the problem of inconsistent spatial benchmarks, and generate the first perception feature information with consistent spatio-temporal. Subsequently, a dynamic weight adaptive adjustment mechanism is introduced. High-order features of each modality are extracted using a deep neural network, and combined with the real-time task type (such as assembly, detection, abnormal handling) and scene complexity (light intensity, noise level), the cross-modal correlation weight is calculated through an attention mechanism, and finally the optimized feature fusion parameters are generated. Based on the fusion parameters, a multi-objective joint decision-making model is constructed, and a deep reinforcement learning framework is used to encode the robot motion state, environmental constraints, and task objectives into the state space, output a decision instruction set including path planning parameters, operation force thresholds, and cooperation timings, and perform iterative optimization through a multi-objective optimization function until the decision score reaches the convergence condition. Finally, the system monitors the decision execution effect (such as task completion time, positioning deviation, energy consumption data) in real time, and compares it with the preset thresholds (such as the accuracy threshold of ±0.1 mm, the response delay upper limit of 200 ms) for analysis: if the optimization effect does not meet the standard, the weight backtracking mechanism is triggered or the sensor online calibration process is started to form a perception-decision-validation closed-loop control to ensure the robustness and self-adaptability of the system in a dynamic industrial environment.

[0040] According to an embodiment of the present invention, the cross-modal spatio-temporal alignment includes: Precisely align the images captured by the camera, the force changes detected by the tactile sensor, and the sounds collected by the microphone according to the acquisition time through a time synchronization module; Unify the position information detected by different sensors into the three-dimensional space coordinate system where the robot is located through coordinate transformation.

[0041] Among them, at the time synchronization level, a hardware clock synchronization module adopting the Precision Time Protocol is used to align the data acquisition timestamps of cameras, tactile sensors (such as six-axis force sensors), and microphones, and an interpolation compensation algorithm is used to eliminate the timing deviation caused by the difference in sensor sampling frequencies, ensuring that the visual images, tactile force feedback, and collision sound data of the same event (such as the moment when the robotic arm touches the workpiece) are strictly synchronized. At the spatial alignment level, a multi-level coordinate system transformation model is established. First, the external parameter matrix of the camera is obtained through hand-eye calibration to map the image pixel coordinates to the robot base coordinate system; at the same time, the detected contact force / moment is transformed from the tool coordinate system to the base coordinate system by using the installation pose matrix of the tactile sensor; for audio localization, the azimuth angle of the sound source is estimated by the direction of arrival of the microphone array, and the three-dimensional coordinates of the sound source are calculated in combination with the robot kinematic model. Finally, all sensor data are unified into the three-dimensional space coordinate system of robot motion control through homogeneous coordinate transformation to form a spatio-temporally consistent multi-modal perception benchmark. Through the above dual alignment mechanism, the problem of data fusion distortion caused by spatio-temporal reference differences in traditional multi-sensor systems is completely solved, providing high-precision input for subsequent dynamic weight allocation and collaborative decision-making.

[0042] According to an embodiment of the present invention, the dynamic weight adjustment includes: Automatically setting the importance ratio of each sensor according to the current task type, and setting the weight of the tactile sensor to the first weight ratio in the fine assembly task; When it is necessary to detect machine anomalies, the weight of the sound sensor is increased to the second weight ratio, and at the same time, the weight of the camera image is decreased to the third weight ratio.

[0043] Among them, the dynamic weight allocation mechanism realizes the intelligent allocation of multi-modal perception resources through task type recognition and sensor efficiency evaluation. The system is built with a task semantic parsing module that extracts keywords and classifies intentions from task instructions based on natural language processing, and generates task type tags in combination with scene perception data (such as workpiece size tolerance ±0.05mm, ambient noise level >75dB). For precision assembly tasks, the system calls a preset priority rule library to increase the weight of the tactile sensor to the first weight ratio (such as 60%-70%), preferentially uses six-axis force sensor data to analyze contact force feedback, and at the same time reduces the visual weight to an auxiliary level, only retaining the key pose verification function; when switching to the device anomaly diagnosis mode, the frequency domain analysis module is used to identify the acoustic signal characteristics, dynamically increase the weight of the sound sensor to the second weight ratio (such as 65%-75%), and uses wavelet denoising and MFCC feature extraction to enhance the accuracy of voiceprint recognition, while reducing the visual weight to the third weight ratio (such as 10%-15%) to reduce the computational overhead. The weight allocation process is continuously optimized through a reinforcement learning model, and the weight boundary values are dynamically corrected in combination with historical operation data (such as assembly success rate, anomaly detection accuracy) to ensure the precise matching of sensor resource allocation and task requirements, which can improve the task execution efficiency compared with the traditional fixed weight scheme.

[0044] According to an embodiment of the present invention, the generation process of the feature fusion parameter is as follows: Analyze the correlation degree between visual, tactile, and auditory data through an attention calculation model; Automatically allocate the comprehensive influence values of each sensor according to the correlation degree, ensuring that the sum of the influence values of vision, touch, and hearing is 1.

[0045] Among them, the multi-modal correlation analysis and weight allocation method of the present invention is based on an improved multi-head attention mechanism to achieve cross-modal data collaborative optimization. Specifically, first, a cross-modal feature interaction matrix is constructed. Visual data extracts spatial features through a CNN, tactile data encodes temporal mechanical features through an LSTM, and auditory data extracts frequency domain features using a Mel spectrogram convolutional network. Subsequently, the self-attention calculation module models the internal feature correlation degree of a single modality, and uses the cross-attention branch to calculate the cross-modal interaction weights of vision-tactile, vision-auditory, and tactile-auditory to form a correlation score matrix. Based on this, a dynamic normalization weight allocator is designed, and finally, the softmax function is used to constrain the vision, touch, and hearing weights to satisfy the sum of the influence values to be 1. This method can improve the efficiency of multi-modal data fusion compared with the traditional weighted average strategy, and avoids the problem of weight overfitting through preset normalization constraints.

[0046] According to an embodiment of the present invention, the preset optimization model uses a deep learning model, where: The moving state of the robot, the surrounding environment conditions, and the task objectives together constitute the decision-making basis; The output control instructions include movement route parameters, operation force parameters, and cooperation timing parameters; The decision-making results are continuously optimized through the comprehensive evaluation of energy consumption, operation accuracy, and safety.

[0047] Among them, the preset optimization model of the present invention is constructed based on the deep reinforcement learning framework, and intelligent decision-making is realized through multi-dimensional state perception and multi-objective collaborative optimization. This model uses the real-time movement state of the robot, surrounding environment information, and task objective requirements as the core decision-making basis. The robot's movement state includes dynamic data such as its current position, speed, and posture; the surrounding environment conditions include information such as obstacle distribution, positions of other devices, and environmental changes; the task objective clearly defines specific operation requirements and expected achieved effects. Based on these comprehensive decision-making bases, the model outputs control instructions including movement route parameters, operation force parameters, and cooperation timing parameters to accurately control the robot's actions. At the same time, to ensure the scientificity and efficiency of decision-making, the model constructs a comprehensive evaluation system for energy consumption, operation accuracy, and safety. By quantitatively evaluating each decision-making result in multiple dimensions and using the evaluation feedback to optimize decision-making parameters, continuous iterative improvement of the decision-making results is realized, enabling the industrial robot to make more energy-saving, accurate, and safe decisions in a complex and changeable working environment.

[0048] According to the embodiments of the present invention, when comparing the optimization effects: When the task completion accuracy rate is lower than the preset first completion rate or the response delay exceeds the preset time, the previous weight setting is automatically restored; If the decision-making fails to meet the standard for three consecutive times, all sensors are recalibrated and the data fusion rule is updated.

[0049] Among them, when evaluating and comparing the optimization effects of decision-making parameters, the present invention sets up a dual safeguard mechanism to ensure the stability and reliability of the industrial robot decision-making system. When the task completion accuracy rate is lower than the preset first completion rate standard, or the system response delay exceeds the preset time threshold, it indicates that the current decision-making parameters and weight configuration fail to meet the task requirements. At this time, the system will automatically restore the weight setting that has been verified effective in the previous stage, quickly correct the decision-making deviation, and maintain the operation stability. If the decision-making results of three consecutive times do not meet the preset standard, it is determined that there may be problems with sensor data deviation or data fusion rule failure in the system. The system will trigger a comprehensive calibration program to recalibrate all sensors and update the data fusion rule at the same time. By deeply optimizing the perception and decision-making links, the problem of unqualified decision-making is solved from the root, thereby constructing a dynamic feedback and self-repair mechanism, effectively improving the fault tolerance and continuous operation efficiency of the industrial robot autonomous collaborative decision-making system under complex working conditions.

[0050] According to the embodiments of the present invention, the optimization process of the comprehensive evaluation also includes: Real-time monitor the environmental noise level and light intensity. When the environmental noise exceeds the preset noise threshold, dynamically increase the band filtering intensity of the auditory sensor and reduce the visual weight; When the light intensity is lower than the preset intensity threshold, activate the infrared supplementary light module and increase the depth of the visual feature extraction network. At the same time, increase the tactile weight to compensate for the impact of the decline in visual perception ability.

[0051] Among them, in the optimization process of the comprehensive score, the system uses the environmental perception module to real-time monitor the environmental noise level and light intensity parameters. When it is detected that the environmental noise exceeds the preset noise threshold (such as 70 dB), the system automatically triggers the dynamic noise reduction mechanism: on the one hand, by enhancing the band filtering intensity of the auditory sensor, signal enhancement processing is performed on specific noise frequency bands to improve the recognition accuracy of abnormal sounds; on the other hand, based on the multi-modal weight allocation strategy, the decision-making weight of the visual sensor is synchronously reduced to reduce the misjudgment impact of visual information under noise interference. When it is monitored that the light intensity is lower than the preset intensity threshold (such as 150 lux), the system starts the environmental adaptability compensation mechanism. First, activate the infrared supplementary light module to perform spectral compensation on the visual acquisition area, and at the same time increase the depth of the convolutional layer of the visual feature extraction network to improve the feature analysis ability of low-light images; at the same time, dynamically increase the decision-making weight of the tactile sensor to 1.2 - 1.5 times the reference value, and compensate for the information loss caused by the decline in visual perception ability by strengthening the decision-making participation of tactile feedback. This multi-modal dynamic adjustment mechanism based on environmental parameter adaptation effectively guarantees the perception robustness and decision reliability of the industrial robot under complex working conditions.

[0052] According to an embodiment of the present invention, it further includes: Collect the working environment temperature of the robot during the dynamic adjustment of the weight; When the working environment temperature is greater than the first temperature threshold, increase the sampling frequency of the tactile sensor to compensate for the material deformation error caused by high temperature, and reduce the visual weight to avoid image distortion interference; When the working environment temperature is less than the second temperature threshold, enable the low-frequency enhancement mode of the auditory sensor to improve the recognition of metal abnormal sounds, and increase the tactile weight to strengthen the force control accuracy at low temperature.

[0053] Among them, during the dynamic adjustment of the weights, the system collects the working environment temperature data in real time through a temperature sensor integrated in the robot body, and triggers a differential multi-modal perception optimization strategy according to a preset first temperature threshold (such as 45 °C) and a second temperature threshold (such as 5 °C). When it is detected that the ambient temperature is higher than 45 °C, the system activates a high-temperature compensation mechanism to increase the sampling frequency of the tactile sensor from the reference value, for example, from 100 Hz to 200 Hz, and compensates for the measurement deviation of the contact force caused by the thermal expansion of the metal material through high-frequency data acquisition; at the same time, the visual weight is down-regulated to reduce the interference of image distortion caused by high-temperature air waves on the workpiece positioning accuracy. When the ambient temperature is lower than 5 °C, the system activates a low-temperature enhancement mode, enables low-frequency band-pass filtering for the auditory sensor to strengthen the feature extraction of abnormal sounds of metal cold brittle fracture, and at the same time up-regulates the reference value of the tactile weight to offset the influence of mechanical transmission hysteresis caused by low temperature by increasing the decision-making proportion of force feedback. This temperature adaptive adjustment mechanism maintains the robot's comprehensive perception accuracy to meet the requirements in extreme working conditions such as foundry workshops (high temperature) and cold chain warehouses (low temperature) by dynamically balancing the performance attenuation of multi-modal sensors and scene requirements, which is significantly better than the level of traditional fixed parameter solutions.

[0054] According to an embodiment of the present invention, when generating a multi-modal joint decision-making strategy, it further includes: If the task completion degree of a certain robot is lower than the preset progress threshold ratio, automatically increase its path planning priority by 2 levels, and other robots avoid its working area; When the target positions of multiple robots overlap, dynamically allocate access permissions according to the task urgency; Calculate the remaining power of each robot in real time, and reduce the motion speed weight of the robot with the power lower than the preset power threshold to extend the cooperation cycle.

[0055] Among them, when generating a multi-modal joint decision-making strategy, the system uses the task progress monitoring module to evaluate the task completion status of each robot in real time. When it is detected that the task completion degree of a certain robot is lower than the preset progress threshold (such as 30%), the lag compensation mechanism is automatically triggered, the path planning priority of this robot is increased by 2 levels (for example, from the default priority 3 to priority 1), and its working area is marked as a high-priority passage area. The remaining robots adjust their paths in real time based on the obstacle avoidance planning model and detour around the lagging robot's working area at a safe distance. For the conflict event of overlapping target positions of multiple robots, the system calls the task urgency evaluation model and dynamically allocates passage rights according to the preset rule base (such as the task urgency weight of dangerous goods handling is 0.9 and that of ordinary assembly tasks is 0.4). The high-urgency robot obtains the right of priority passage and demarcates an exclusive path through the laser navigation module. At the same time, the system integrates a power consumption monitoring unit to calculate the remaining power of each robot in real time. When it is detected that the power is lower than the preset threshold (such as 20%), the energy consumption balancing strategy is started, the reference value of the motion speed weight coefficient of the low-power robot is reduced, its moving speed is re-planned through the kinematic model, and the collaborative task cycle is extended (such as adjusting the handling cycle from 5 minutes per time to 5.5 - 5.8 minutes per time). While avoiding task interruption, the continuous working time of the robot cluster is extended by more than 2.3 hours. This multi-dimensional collaborative optimization mechanism improves the overall production capacity of the multi-robot assembly line, reduces the task conflict rate, and decreases the low-power shutdown events.

[0056] In the third aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a program for the autonomous collaborative decision-making method of an industrial robot based on multi-modal perception. When the program for the autonomous collaborative decision-making method of an industrial robot based on multi-modal perception is executed by a processor, the steps of the autonomous collaborative decision-making method of an industrial robot based on multi-modal perception as described in any one of the above are implemented.

[0057] The industrial robot autonomous collaborative decision-making method, system and medium based on multi-modal perception disclosed by the present invention achieve high-precision autonomous operation in industrial scenarios through a multi-source sensor collaborative perception and intelligent decision-making mechanism. Specifically, the system first synchronously collects visual, tactile and auditory information through a high-resolution camera, a six-axis force sensor and a directional microphone, and uses a cross-modal spatio-temporal alignment technology to solve the problem of multi-source data fusion: in the time dimension, timestamp alignment is achieved based on a precise time protocol hardware clock synchronization module, and the interpolation compensation algorithm is combined to eliminate the sampling frequency difference; in the space dimension, the visual pixel coordinates are mapped to the robot base coordinate system through a hand-eye calibration matrix, the force / torque data is converted by the tactile sensor installation pose matrix, and the three-dimensional coordinates of the sound source are estimated by the microphone array direction of arrival, and finally a spatio-temporally unified multi-modal perception benchmark is constructed. Subsequently, the system realizes the optimized allocation of multi-modal perception resources based on a dynamic weight allocation mechanism: the current operation type is identified through a task semantic parsing module, the sensor weight ratio is dynamically adjusted in combination with environmental parameters, and an improved multi-head attention mechanism is used to calculate the cross-modal correlation degree, and the feature fusion parameters are generated through softmax normalization. The decision optimization layer adopts a deep reinforcement learning framework, takes the robot motion state, environmental constraints and task objectives as inputs, outputs a control instruction set including path planning parameters, operation force thresholds and collaboration time sequences, and continuously optimizes the decision-making model through multi-objective scoring of energy consumption, accuracy and safety. The system also integrates an environmental adaptive compensation mechanism: when the temperature exceeds the set threshold, the tactile sampling frequency and visual weight are adjusted, and when the temperature is lower than the set threshold, the auditory low-frequency enhancement is enabled and the tactile weight is increased; in the task execution layer, efficient cooperation of multiple robots is achieved through a lag compensation mechanism, an urgency arbitration model and a power balance strategy. Finally, an adaptive optimization of perception-decision-execution is realized through a closed-loop verification mechanism, significantly improving the comprehensive accuracy and cooperation efficiency in complex industrial scenarios.

[0058] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0059] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0060] In addition, each functional unit in the embodiments of the present invention may be fully integrated in a processing unit, or each unit may be separately regarded as a unit, or two or more units may be integrated in one unit; the above-mentioned integrated units may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0061] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memories, random access memories, magnetic disks, or optical disks and other various media that can store program codes.

[0062] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

Claims

1. An industrial robot autonomous collaborative decision-making method based on multimodal perception, characterized in that, It includes the following steps: Collect the perception information within the preset period of the autonomous collaborative system, including visual information, tactile information, and auditory information; Perform cross-modal spatio-temporal alignment on the visual information, tactile information, and auditory information through a preset method to generate first perception feature information; Extract visual feature data, tactile feature data, and auditory feature data according to the first perception feature information through a preset mechanism and perform dynamic weight adjustment to generate feature fusion parameters; Generate a multi-modal joint decision-making strategy according to the feature fusion parameters, and iteratively optimize the decision-making parameters through a preset optimization model; Obtain the decision parameter optimization effect data and compare it with the preset effect threshold to determine whether the optimization effect meets the requirements.

2. The method for autonomous collaborative decision-making of an industrial robot based on multi-modal perception according to claim 1, wherein The cross-modal spatio-temporal alignment includes: Precisely align the images captured by the camera, the force changes detected by the tactile sensor, and the sounds collected by the microphone according to the acquisition time through the time synchronization module; Unify the position information detected by different sensors into the three-dimensional space coordinate system where the robot is located through coordinate transformation.

3. The method for autonomous collaborative decision-making of an industrial robot based on multi-modal perception according to claim 1, wherein The dynamic weight adjustment includes: Automatically set the importance ratio of each sensor according to the current task type. In the fine assembly task, the weight of the tactile sensor is set to the first weight ratio; When it is necessary to detect machine abnormalities, the weight of the sound sensor is increased to the second weight ratio, and at the same time, the weight of the camera image is reduced to the third weight ratio.

4. The industrial robot autonomous collaborative decision-making method based on multi-modal perception according to claim 1, wherein, The generation process of the feature fusion parameters is: Analyze the correlation degree between visual, tactile, and auditory data through an attention calculation model; Automatically allocate the comprehensive influence values of each sensor according to the correlation degree, ensuring that the sum of the influence values of vision, touch, and hearing is 1.

5. The method for autonomous collaborative decision-making of an industrial robot based on multi-modal perception according to claim 1, wherein The preset optimization model uses a deep learning model, where: The movement state of the robot, the surrounding environment conditions, and the task objectives together constitute the decision-making basis; The output control instructions include movement route parameters, operation force parameters, and cooperation timing parameters; Continuously optimize the decision-making results through the comprehensive evaluation of energy consumption, operation accuracy, and safety.

6. The industrial robot autonomous collaborative decision-making method based on multi-modal perception according to claim 1, characterized in that When comparing the optimization effects: When the task completion accuracy rate is lower than the preset first completion rate or the response delay exceeds the preset time, automatically restore the previous weight settings; If the decision-making fails to meet the standard three times in a row, recalibrate all sensors and update the data fusion rules.

7. An industrial robot autonomous collaborative decision-making system based on multimodal perception, characterized in that It includes a memory and a processor. The memory includes a program for the autonomous collaborative decision-making method of an industrial robot based on multi-modal perception. When the program for the autonomous collaborative decision-making method of an industrial robot based on multi-modal perception is executed by the processor, the following steps are implemented. Specifically: Collect the perception information within the preset period of the autonomous collaborative system, including visual information, tactile information, and auditory information; Perform cross-modal spatio-temporal alignment on the visual information, tactile information, and auditory information through a preset method to generate first perception feature information; Extract visual feature data, tactile feature data, and auditory feature data according to the first perception feature information through a preset mechanism and perform dynamic weight adjustment to generate feature fusion parameters; Generate a multi-modal joint decision-making strategy according to the feature fusion parameters, and iteratively optimize the decision-making parameters through a preset optimization model; Obtain the optimization effect data of the decision-making parameters and compare it with the preset effect threshold to determine whether the optimization effect meets the requirements.

8. The industrial robot autonomous collaborative decision-making system based on multi-modal perception according to claim 7, characterized in that The cross-modal spatio-temporal alignment includes: Precisely align the images captured by the camera, the force changes detected by the tactile sensor, and the sounds collected by the microphone according to the acquisition time through the time synchronization module; Unify the position information detected by different sensors into the three-dimensional space coordinate system where the robot is located through coordinate transformation.

9. The industrial robot autonomous collaborative decision-making system based on multi-modal perception according to claim 8, characterized in that, The weight dynamic adjustment includes: Automatically set the importance ratio of each sensor according to the current task type. In the fine assembly task, the weight of the tactile sensor is set to the first weight ratio; When it is necessary to detect machine anomalies, the weight of the sound sensor is increased to the second weight ratio, and at the same time, the weight of the camera image is decreased to the third weight ratio.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes an industrial robot autonomous collaborative decision-making method, system, and medium program based on multi-modal perception. When the industrial robot autonomous collaborative decision-making method, system, and medium program based on multi-modal perception are executed by a processor, the steps of the industrial robot autonomous collaborative decision-making method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Industrial intelligent detection method and system based on multi-modal large model

    CN118503832A

  • Multi-modal sensing integrated robot cooperative operation method

    CN119871451A

  • Interactive Tactile Perception Method for Classification and Recognition of Object Instances

    US20220126453A1

Cited By

  • Robot body intelligent control method and system based on multi-mode perception

    CN120516727A

  • Multi-mode adaptive industrial instrument intelligent control system based on deep learning

    CN120630717A

  • Robot dynamic risk assessment and decision-making system and method based on multi-modal perception

    CN120680531A

  • Wafer component driving detection method and system based on perforated plate and multiple sensors

    CN120740686A

  • Robot control method and related equipment

    CN120773043A