Robot multi-mode sensing and motion cooperative control method and device

Dynamically adjusting the parameters of the robot arm through multimodal perceptual data, the problem that traditional sorting robots cannot adapt to the dynamic changes of conveyor belts and cargoes is solved, and accurate and efficient sorting operations and production process stability is achieved.

CN120347724AActive Publication Date: 2025-07-22HANGZHOU FANJIA TECH CO LTD

Patent Information

Application Number
CN202510814097.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Due to the dependence of fixed control parameters, traditional sorting robots cannot sense and adapt to the dynamic changes in conveyor belt speed and cargo position in real time, resulting in insufficiency of sorting, increasing maintenance costs and affecting the stability of production processes.

Method used

Multimodal sensing data is used to obtain conveyor belt and cargo information in real time through lidar, infrared sensors and vision sensors, dynamically adjust the historical action parameters of the robotic arm, and build control instructions for the optimal motion coordination parameters.

Benefits of technology

Accurate and efficient sorting operations are achieved, sorting efficiency is improved, cargo damage and sorting errors are reduced, and the stability of the production process is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347724A_ABST
    Figure CN120347724A_ABST
Patent Text Reader

Abstract

The invention discloses a robot multi-modal sensing and motion cooperative control method and device, and the method comprises the steps: obtaining the information of a plurality of types of sensors according to a preset period through the plurality of types of sensors disposed in a robot operation region, and taking the information as multi-modal sensing data, the multi-mode sensing data comprises laser radar data, infrared sensor data and visual sensor data, and a conveying device is arranged in the robot operation area; analyzing conveying parameters of a conveying belt of the conveying equipment and cargo parameters of cargoes on the conveying belt based on the multi-mode sensing data; according to the conveying parameters and the cargo parameters, historical motion parameters of a mechanical arm of the robot are dynamically adjusted, and optimal motion cooperation parameters suitable for the conveying equipment are obtained; and constructing and executing a motion control instruction corresponding to the optimal motion cooperation parameter so as to carry out operation on the goods on the conveyor belt. Therefore, by adopting the embodiment of the invention, the sorting efficiency can be improved, the maintenance cost can be reduced, and the stability of the whole production process can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot technology, and particularly to a method and device for multi-modal perception and motion cooperative control of a robot. Background Art

[0002] In modern industrial production, automated assembly lines are widely used. For example, in a logistics sorting center, sorting robots are widely used to quickly sort the goods on the conveyor belt according to categories. These robots accurately operate the robotic arms to grab the goods from the conveyor belt and place them at the designated positions, thus realizing an efficient and accurate logistics sorting process.

[0003] Currently, traditional sorting robots usually execute tasks with fixed control parameters. These parameters include the motion trajectory of the robotic arm, the timing of the grasping action, and the synchronization with the conveyor belt speed, etc. In practical applications, the robot executes sorting actions according to a pre-set program based on the fixed positions of the goods on the conveyor belt and the fixed speed of the conveyor belt. This control method can work effectively when the conveyor belt speed and the goods positions are relatively stable.

[0004] However, as the usage time increases, the conveying equipment may have its operating parameters changed due to aging, wear or other factors, such as the fluctuation of the conveyor belt speed or the deviation of the goods positions on the conveyor belt. In this case, since the sorting robot depends on fixed parameters and cannot perceive and adapt to these dynamic changes in real time, it may not be able to accurately grab the goods, and may even damage the goods or its own equipment, thus reducing the sorting efficiency, increasing the maintenance cost, and affecting the stability of the entire production process. Summary of the Invention

[0005] The embodiments of this application provide a method and device for multi-modal perception and motion cooperative control of a robot. To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the subsequent detailed description.

[0006] In a first aspect, the embodiments of this application provide a method for multi-modal perception and motion cooperative control of a robot, which is applied to a robot. The method includes: Acquiring multi-class sensor information at a preset period through multi-class sensors arranged in the robot working area as multi-modal perception data. The multi-modal perception data includes lidar data, infrared sensor data, and visual sensor data. A conveying device is provided in the robot working area; Analyzing the conveying parameters of the conveyor belt of the conveying device and the goods parameters of the goods on the conveyor belt based on the multi-modal perception data; Dynamically adjust the historical motion parameters of the robotic arm according to the transfer parameters and the cargo parameters to obtain the optimal motion coordination parameters applicable to the transfer device; Construct and execute the motion control instructions corresponding to the optimal motion coordination parameters to perform operations on the cargo on the conveyor belt.

[0007] In a second aspect, an embodiment of the present application provides a robotic multi-modal perception and motion coordination control device, which includes: A sensor information acquisition module, configured to acquire multi-class sensor information at a preset period through multiple types of sensors arranged in the robot operation area as multi-modal perception data, where the multi-modal perception data includes lidar data, infrared sensor data, and visual sensor data, and a transfer device is provided in the robot operation area; A parameter analysis module, configured to analyze the transfer parameters of the conveyor belt of the transfer device and the cargo parameters of the cargo on the conveyor belt based on the multi-modal perception data; An action parameter adjustment module, configured to dynamically adjust the historical motion parameters of the robotic arm according to the transfer parameters and the cargo parameters to obtain the optimal motion coordination parameters applicable to the transfer device; An instruction execution module, configured to construct and execute the motion control instructions corresponding to the optimal motion coordination parameters to perform operations on the cargo on the conveyor belt.

[0008] The technical solutions provided by the embodiments of the present application may include the following beneficial effects: In the embodiments of the present application, on the one hand, through the multi-modal perception data, the robot can obtain the running state of the conveyor belt and the detailed information of the cargo in real time, so as to dynamically adjust the historical motion parameters of the robotic arm, realizing accurate and efficient sorting operations. The accurate sorting operations can significantly improve the sorting efficiency, reduce the damage or sorting errors of the cargo caused by misoperations, and further reduce the costs generated by frequent maintenance and repair of the equipment. On the other hand, by dynamically adjusting the historical motion parameters of the robotic arm to adapt to the real-time running state of the conveyor belt and the characteristics of the cargo, the robot can maintain a stable operation performance under different working conditions. This adaptive ability can effectively avoid operation interruptions or incorrect operations caused by problems such as changes in the conveyor belt speed, offset of the cargo position, or wear of the conveyor belt, thereby ensuring the stability of the entire production process.

[0009] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0011] Figure 1 It is a schematic flowchart of a method for multi-modal perception and motion cooperative control of a robot provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a multi-modal perception and motion cooperative control system of a robot provided by an embodiment of the present application; Figure 3 It is a signal diagram between a time-domain signal and a frequency-domain signal provided by an embodiment of the present application; Figure 4 It is a model architecture diagram of a pre-trained action parameter analysis model provided by an embodiment of the present application; Figure 5 It is a schematic flowchart of a model training method for an action parameter analysis model provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of a multi-modal perception and motion cooperative control device of a robot provided by an embodiment of the present application; Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0012] The following description and drawings fully disclose specific implementation manners of the present application, enabling those skilled in the art to practice them.

[0013] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0014] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0015] In the description of the present application, it should be understood that terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations. In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0016] Currently, traditional sorting robots usually execute tasks with fixed control parameters. These parameters include the movement trajectory of the robotic arm, the timing of the grasping action, and the synchronization with the conveyor belt speed, etc. In practical applications, the robot executes sorting actions according to a pre-set program based on the fixed positions of the goods on the conveyor belt and the fixed speed of the conveyor belt. This control method can work effectively when the conveyor belt speed and the positions of the goods are relatively stable.

[0017] The inventors have realized that as the usage time increases, the conveying equipment may have its operating parameters changed due to aging, wear, or other factors, such as fluctuations in the conveyor belt speed or the deviation of the goods' positions on the conveyor belt. In such cases, since the sorting robot relies on fixed parameters and cannot perceive and adapt to these dynamic changes in real time, it may fail to accurately grasp the goods, and may even damage the goods or its own equipment, thereby reducing the sorting efficiency, increasing the maintenance cost, and affecting the stability of the entire production process.

[0018] To solve the above problems, the present application provides a method and device for multi-modal perception and motion collaborative control of a robot to solve the problems existing in the above related technical problems. In the embodiments of the present application, on the one hand, through multi-modal perception data, the robot can obtain the operating state of the conveyor belt and the detailed information of the goods in real time, thereby dynamically adjusting the historical action parameters of the robotic arm to achieve precise and efficient sorting operations. Precise sorting operations can significantly improve the sorting efficiency, reduce the damage to the goods or sorting errors caused by misoperations, and further reduce the costs generated due to frequent maintenance and repair of the equipment. On the other hand, by dynamically adjusting the historical action parameters of the robotic arm to adapt to the real-time operating state of the conveyor belt and the characteristics of the goods, the robot can maintain a stable operation performance under different working conditions. This adaptive ability can effectively avoid operation interruptions or incorrect operations caused by problems such as changes in the conveyor belt speed, deviation of the goods' positions, or wear of the conveyor belt, thereby ensuring the stability of the entire production process. The following will be described in detail with exemplary embodiments.

[0019] The following will be combined with the attached Figure 1 - attached Figure 5 to introduce in detail the method for multi-modal perception and motion collaborative control of a robot provided by the embodiments of the present application. This method can be implemented depending on a computer program and can run on a robot multi-modal perception and motion collaborative control device based on the von Neumann architecture. This computer program can be integrated into an application or run as an independent tool-type application.

[0020] Please refer to Figure 1 , which is a schematic flowchart of a method for multi-modal perception and motion collaborative control of a robot provided by the embodiments of the present application and is applied to a robot. As Figure 1As shown in the figure, the method of the embodiment of the present application includes the following steps: S101, obtain multi-class sensor information at a preset period through multi-class sensors arranged in the robot working area as multi-modal perception data, where the multi-modal perception data includes LiDAR data, infrared sensor data, and visual sensor data, and a conveying device is provided in the robot working area; Among them, the robot working area is the spatial range where the robot is located when performing tasks, such as an automated production line, etc. The multi-class sensors are multiple different types of data acquisition devices installed in the robot working area for obtaining environmental information. The preset period refers to the time interval for the sensor to collect data, which is a fixed and pre-set time length. The multi-modal perception data refers to the data obtained through multiple different types of sensors, and these data reflect the state of the environment from multiple perspectives. LiDAR (Light Detection and Ranging) is a sensor that uses lasers for distance measurement and environmental modeling, and LiDAR data refers to the distance information and three-dimensional point cloud data collected by LiDAR. The infrared sensor is a sensor that obtains information by detecting the infrared radiation emitted by an object, and the infrared sensor data refers to information such as temperature and reflectivity collected by the infrared sensor. The visual sensor usually refers to a camera for obtaining image or video data, and the visual sensor data refers to the two-dimensional image or video information collected by the camera. The conveying device refers to a mechanical device for transporting goods, usually including a conveyor belt, a conveyor, etc.

[0021] In some embodiments of the present application, a variety of sensors are installed in the working area of the robot, including LiDAR, infrared sensors, and visual sensors. These sensors are distributed around the conveying device and can cover the conveyor belt and the goods on it. The LiDAR is used to sense the spatial position of the conveyor belt and the goods. The infrared sensor is used to judge the wear condition of the conveyor belt and the thermal characteristics of the goods. The visual sensor is used to capture the image information of the conveyor belt and the goods to identify the appearance characteristics of the goods, such as color, shape, and size. All sensors collect data synchronously at a preset time period to obtain multi-modal perception data. Among them, the preset time period can be 3 seconds.

[0022] For example Figure 2 As shown in the figure, Figure 2 is a schematic structural diagram of a robot multi-modal perception and motion coordination control system provided by the present application, including a robot, a conveying device, a robotic arm, and multi-class sensors. The multi-class sensors include LiDAR, visual sensors, and infrared sensors. After the system is started, multi-class sensor information is obtained at a preset period through multi-class sensors arranged in the robot working area as multi-modal perception data.

[0023] S102, analyze the conveying parameters of the conveyor belt of the conveying device and the goods parameters of the goods on the conveyor belt based on the multi-modal perception data; Among them, the transfer parameters are parameters describing the operating state of the transfer device, including the running speed, inclination angle, and wear data of the conveyor belt; the cargo parameters are characteristic parameters describing the cargo on the conveyor belt, including the central coordinates, size, and relative position to the edge of the conveyor belt of the cargo.

[0024] In some embodiments of the present application, a high-definition camera installed in front of the conveyor belt captures images of the cargo on the conveyor belt at a frequency of 30 frames per second to record the appearance characteristics of the cargo. A lidar installed on the side of the conveyor belt scans the conveyor belt and the cargo in real time to generate high-precision three-dimensional point cloud data for measuring the central coordinates, size, and relative position to the edge of the conveyor belt of the cargo. An infrared sensor arranged on the side of the conveyor belt detects the temperature change and wear condition of the conveyor belt, and simultaneously monitors the thermal characteristics of the cargo. The visual sensor captures the marking points on the conveyor belt, and combines with the scanning data of the lidar to calculate the real-time running speed of the conveyor belt. For example, the displacement of the marking point in two consecutive frames of images divided by the time interval gives the conveyor belt speed. The lidar is used to measure the height difference between both sides of the conveyor belt, and combines with the length of the conveyor belt to calculate the inclination angle of the conveyor belt.

[0025] S103. Dynamically adjust the historical action parameters of the robotic arm of the robot according to the transfer parameters and the cargo parameters to obtain the optimal motion coordination parameters applicable to the transfer device; Among them, the historical action parameters of the robotic arm refer to the action parameters used by the robotic arm of the robot in previous operations.

[0026] In some embodiments of the present application, the specific process of dynamically adjusting the historical action parameters of the robotic arm of the robot according to the transfer parameters and the cargo parameters to obtain the optimal motion coordination parameters applicable to the transfer device includes: converting the running speed and inclination angle into frequency domain features through dynamic Fourier transform; constructing a wear thermal map of the conveyor belt according to the wear data; fusing the frequency domain features with the wear thermal map to obtain an environmental encoding vector; using an attention mechanism to perform attention encoding on the central coordinates, size, and relative position to generate a cargo encoding vector for each cargo; inputting the environmental encoding vector and the cargo encoding vector into a pre-trained action parameter analysis model to output the current action parameters of the robotic arm of the robot; in the case where the current action parameters are inconsistent with the historical action parameters of the robotic arm of the robot, replacing the historical action parameters with the current action parameters to obtain the optimal motion coordination parameters applicable to the transfer device.

[0027] Among them, the dynamic Fourier transform is a mathematical tool that converts a time-domain signal into a frequency-domain signal for analyzing the frequency components of the signal. Frequency-domain features refer to the representation of a signal in the frequency domain, including information such as frequency components, amplitude, and phase. Wear data refers to the relevant information describing the wear degree of the conveyor belt, including wear depth, temperature change, reflectivity change, etc. The wear heat map is a visualization tool that represents the wear degree of different areas of the conveyor belt by the shade of color. The environmental coding vector is a vector that comprehensively represents the operating state and wear condition of the conveyor belt, obtained by fusing the frequency-domain features and the spatial information of the wear heat map. The attention mechanism is a neural network technology that simulates human attention. The attention mechanism can be used to identify the key features of the goods (such as the central coordinates, size, and relative position) and generate a goods coding vector. Through the attention mechanism, the robot can more accurately identify and operate the goods. The goods coding vector can integrate the key features of the goods and provide goods-related information for adjusting the motion parameters of the robot. The motion parameter analysis model is a pre-trained neural network model that can learn the relationship between the transfer parameters, goods parameters, and manipulator motion parameters, and dynamically adjust the motion parameters of the manipulator to adapt to different operating environments.

[0028] In the embodiments of the present application, through the fusion of the frequency-domain features and the wear heat map, the robot can more comprehensively understand the operating state and wear condition of the conveyor belt. This multi-dimensional perception enables the robot to accurately identify the position and state of the goods, reducing sorting errors caused by misjudgment. At the same time, the motion parameter analysis model can quickly output the optimal motion parameters, enabling the robot to immediately adjust its motion to adapt to the change in the conveyor belt speed and the dynamic position of the goods, thereby improving the sorting efficiency.

[0029] In some embodiments of the present application, the specific process of converting the running speed and tilt angle into frequency-domain features through the dynamic Fourier transform is as follows: filtering, denoising, and normalizing the running speed and tilt angle that change with time to obtain the preprocessed time-domain signal of the conveying device; using the dynamic Fourier transform expression to convert the time-domain signal of the conveying device into a frequency-domain signal to obtain the frequency spectrum data of the conveying device; identifying the main frequency components, the amplitude of each frequency component, and the phase of each frequency component in the frequency spectrum data of the conveying device as the frequency-domain information of the running speed and tilt angle; fusing the frequency-domain information of the running speed and tilt angle to obtain the frequency-domain features.

[0030] Among them, the dynamic Fourier transform expression is:

[0031] Among them, is the complex value of the th frequency component of the frequency-domain signal, is the The value of a sampling point, is the total number of sampling points, the sampling index of the time-domain signal, is the frequency index of the frequency-domain signal, is a complex exponential function used to transform the time-domain signal to the frequency domain.

[0032] For example, a low-pass filter is used to remove high-frequency noise and normalize the signal to the range of [0, 1]. The preprocessed time-domain signal is transformed into a frequency-domain signal using the Fast Fourier Transform (FFT). The main frequency components, the amplitude and phase of each frequency component are extracted from the frequency-domain signal. An example of the signal diagram between the time-domain signal and the frequency-domain signal is Figure 3 shown as follows.

[0033] In some embodiments of the present application, the specific process of constructing a wear heat map of the conveyor belt based on wear data is as follows: from the wear data, the wear depth of the conveyor belt, the temperature change in the wear area, and the reflectivity change in the wear area are extracted; the surface of the conveyor belt is divided into multiple small grids based on preset grid division parameters, and each small grid corresponds to a wear data point; the wear depth of the conveyor belt, the temperature change in the wear area, and the reflectivity change in the wear area are normalized and then mapped onto each small grid to form a wear data matrix; by performing spatial smoothing processing on the wear data matrix, a raster layer representing the wear degree is generated; according to the magnitude of the wear data in the wear data matrix, different shades of colors are assigned to the raster layer to obtain the wear heat map of the conveyor belt.

[0034] For example, the generated wear heat map can intuitively display the wear condition of the conveyor belt, helping maintenance personnel to detect potential problems in advance and reduce the occurrence frequency of sudden failures. At the same time, combined with the robot control system, the action parameters can be dynamically adjusted to optimize the production process and improve production efficiency.

[0035] In some embodiments of the present application, the specific process of fusing the frequency-domain features with the wear heat map to obtain an environmental coding vector includes: extracting the spatial information of the wear heat map; mapping the spatial information of the wear heat map onto the frequency-domain features to splice the frequency-domain features with the spatial information of the wear heat map to obtain a spliced feature; encoding the spliced feature to obtain an environmental coding vector.

[0036] Among them, the pre-trained action parameter analysis model includes a weight matrix loading layer, a quantization layer, an attention layer, a dot product operation layer, a feature mapping layer, and a clipping layer, for example Figure 4 shown as follows.

[0037] In some embodiments of the present application, the specific process of inputting the environmental encoding vector and the cargo encoding vector into a pre-trained motion parameter analysis model to output the current motion parameters of the robotic arm includes: the weight matrix loading layer loads the pre-trained weight matrix; the quantization layer uses the pre-trained weight matrix to quantize the environmental encoding vector and the cargo encoding vector respectively to obtain a query vector, a key vector, and a value vector; the attention layer calculates an attention score according to the query vector, the key vector, and the value vector; the attention score is used to measure the importance of the environmental encoding vector to the cargo encoding vector; the dot product operation layer performs dot product operations on the environmental encoding vector and the cargo encoding vector with the attention score respectively and sums them to obtain a joint feature containing environmental and cargo interaction information; the feature mapping layer maps the joint feature to a preset robotic arm motion space to obtain a quantity matching the dimension of the robotic arm degrees of freedom; the clipping layer uses the pre-stored physical constraints of the robotic arm to clip the quantity matching the dimension of the robotic arm degrees of freedom to obtain the current motion parameters of the robotic arm of the robot.

[0038] The quantization formula is: ; Where is the query vector, is the key vector, is the value vector, ([[]] ) is the pre-trained weight matrix for the query vector, key vector, and value vector of the model, is the environmental encoding vector, is the cargo encoding vector; Where the attention score calculation formula is:

[0039] Where is the attention score, is the dimension of the key vector, used to scale the dot product result; where the mapping formula is: ; Where is the quantity matching the dimension of the robotic arm degrees of freedom, is the output layer weight matrix, is the output layer bias term, is the joint feature; where the clipping expression is: ; Where is the maximum allowable angular velocity of each joint of the robotic arm, is the clipping function of the model.

[0040] S104. Construct and execute an action control instruction corresponding to the optimal motion coordination parameter to operate on the cargo on the conveyor belt.

[0041] Among them, the motion control instruction is a specific control instruction generated according to the optimal motion coordination parameters. The motion control instruction converts the optimal motion coordination parameters into specific commands that the robot can execute, ensuring that the robotic arm completes tasks according to the preset motion trajectory and parameters.

[0042] In some embodiments of the present application, specific motion control instructions are generated according to the optimal motion coordination parameters. These instructions need to convert the abstract parameters into specific commands that the robot can execute. For example, the speed parameter is converted into the rotation speed instruction of the motor, and the joint angle is converted into the target position instruction of the joint motor. According to the information such as the position and size of the goods on the conveyor belt, combined with the optimal motion coordination parameters, the specific motion sequence for the robotic arm to complete the operation task is determined. For example, the motion sequence may include moving above the goods, descending to the grasping height, grasping the goods, ascending, moving to the target position, placing the goods, etc. The constructed motion control instructions are sent to the robot control system, usually through a network interface or a directly connected controller. After receiving the motion control instructions, the robot control system analyzes the instruction content and extracts the specific motion parameters. For example, the control system will analyze the target position coordinates, speed, acceleration, etc. information that the robotic arm needs to move to. The robot control system controls the various joint motors and actuators of the robotic arm according to the analyzed instructions and executes the actions according to the preset motion parameters. For example, the robotic arm moves above the goods on the conveyor belt according to the instructions, adjusts the posture of the grasping tool, and then grasps the goods with appropriate force.

[0043] In the embodiments of the present application, on the one hand, through multi-modal perception data, the robot can obtain the running state of the conveyor belt and the detailed information of the goods in real time, thereby dynamically adjusting the historical motion parameters of the robotic arm to achieve precise and efficient sorting operations. Precise sorting operations can significantly improve the sorting efficiency, reduce the damage or sorting errors of goods caused by misoperations, and thus reduce the costs generated by frequent equipment maintenance and repair. On the other hand, by dynamically adjusting the historical motion parameters of the robotic arm to adapt to the real-time running state of the conveyor belt and the characteristics of the goods, the robot can maintain a stable operation performance under different working conditions. This adaptive ability can effectively avoid operation interruptions or incorrect operations caused by problems such as changes in the conveyor belt speed, offset of the goods position, or wear of the conveyor belt, thereby ensuring the stability of the entire production process.

[0044] Please refer to Figure 5 , which is a schematic flowchart of the model training method of a pre-trained reinforcement learning model provided by the embodiments of the present application. As Figure 5 shown, the method of the embodiments of the present application may include the following steps: S201, Simulate different conveyor belt working conditions in the robot operation area in the simulation environment to obtain sample conveyor parameters; Among them, the simulation environment refers to one or more computer system models created using software tools. In the simulation environment of the robot working area, different conveyor belt conditions can be simulated, including the speed, inclination angle, load condition, etc. of the conveyor belt, providing a safe and controllable platform for the development and testing of the robot control strategy. The sample transfer parameters refer to the specific values of the conveyor belt operation state simulated in the simulation environment, such as the speed, inclination angle, wear degree, etc. of the conveyor belt. These parameters are used to train and verify the action parameter analysis model of the robot, helping the robot learn how to adjust the action parameters under different conveyor belt conditions.

[0045] S202, during the simulation process, randomly generate sample cargo parameters; S203, record the sequence of joint angular velocities of the robotic arm input by the user during the simulation process as the supervision signal; S204, construct a model training sample according to the supervision signal, sample transfer parameters, and sample cargo parameters; In some embodiments of the present application, the specific process of constructing a model training sample according to the supervision signal, sample transfer parameters, and sample cargo parameters includes: constructing a sample environment encoding vector according to the sample transfer parameters; constructing a sample cargo encoding vector according to the sample cargo parameters; quantifying the sample environment encoding vector and the sample cargo encoding vector into a sample query vector, a sample key vector, and a sample value vector respectively; using the supervision signal to perform label annotation on the sample query vector, the sample key vector, and the sample value vector to obtain the model training sample.

[0046] For example, import the robot model into the simulation software (such as Delmia, MATLAB, etc.), configure the joint type and motion range to ensure that the model is consistent with the actual physical robot. Add elements such as conveyor belts and goods to the simulation environment, set the initial parameters of the conveyor belt (such as speed, inclination angle, etc.), and construct a complete robot working area. In the simulation environment, simulate different working conditions by modifying parameters such as the speed, inclination angle, and load of the conveyor belt. For example, set the conveyor belt speeds to 0.5 m / s, 1 m / s, and 1.5 m / s respectively, and the inclination angles to 0°, 5°, and 10° respectively. Under each working condition, record the operating parameters of the conveyor belt, including speed, inclination angle, wear degree, etc., to form sample transfer parameters. During the simulation process, randomly generate parameters such as the position, size, and weight of the goods to simulate the diversity of goods in actual production. In the simulation environment, record the sequence of joint angular velocities of the robotic arm input by the user according to the current conveyor belt condition and cargo parameters as the supervision signal. Integrate the sample transfer parameters, sample cargo parameters, and supervision signal together to form the sample data for training the robot action parameter analysis model.

[0047] S205. Create and initialize an action parameter analysis model using a neural network. The model includes a weight matrix loading layer, a quantization layer, an attention layer, a dot product operation layer, a feature mapping layer, and a clipping layer. S206. Perform machine learning on the initialized action parameter analysis model using model training samples so that the action parameter analysis model learns the weight association relationship between the sample transfer parameters, sample cargo parameters, and supervision signals, and obtain a pre-trained weight matrix. S207. Use the model loss function of the action parameter analysis model to quantify the loss value of the pre-trained weight matrix. Among them, the loss function of the model is:

[0048] Among them, is the loss value, is the total number of samples of the model training samples, is the th predicted action parameter of the sample, is the th true action parameter represented by the supervision signal of the sample; among them, is the dimension of the attention score, is the th predicted value of the attention score, is the th true value of the attention score, is the regularization coefficient, used to control the strength of regularization, is the pre-trained weight matrix, where represents any layer in the model.

[0049] S208. Obtain the pre-trained action parameter analysis model when the loss value reaches the minimum.

[0050] In the embodiments of the present application, on the one hand, through multi-modal perception data, the robot can obtain the running state of the conveyor belt and the detailed information of the goods in real time, so as to dynamically adjust the historical action parameters of the robotic arm, realize precise and efficient sorting operations. Precise sorting operations can significantly improve sorting efficiency, reduce goods damage or sorting errors caused by misoperations, and thus reduce the costs generated by frequent equipment maintenance and repair. On the other hand, by dynamically adjusting the historical action parameters of the robotic arm to adapt to the real-time running state of the conveyor belt and the characteristics of the goods, the robot can maintain stable operation performance under different working conditions. This adaptive ability can effectively avoid operation interruptions or misoperations caused by problems such as changes in conveyor belt speed, offset of goods position, or wear of the conveyor belt, thus ensuring the stability of the entire production process.

[0051] The following are device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0052] Please refer to Figure 6 , which shows a schematic structural diagram of a robot multi-modal perception and motion coordination control device provided by an exemplary embodiment of the present application. The robot multi-modal perception and motion coordination control device can be implemented as all or part of an electronic device through software, hardware, or a combination of both. The device 1 includes a sensor information acquisition module 10, a parameter analysis module 20, an action parameter adjustment module 30, and an instruction execution module 40.

[0053] The sensor information acquisition module 10 is configured to obtain multi-class sensor information as multi-modal perception data at a preset period through multi-class sensors arranged in the robot operation area. The multi-modal perception data includes lidar data, infrared sensor data, and visual sensor data. A conveying device is provided in the robot operation area; The parameter analysis module 20 is configured to analyze the conveying parameters of the conveyor belt of the conveying device and the cargo parameters of the cargo on the conveyor belt based on the multi-modal perception data; The action parameter adjustment module 30 is configured to dynamically adjust the historical action parameters of the robot's robotic arm according to the conveying parameters and the cargo parameters to obtain optimal motion coordination parameters applicable to the conveying device; The instruction execution module 40 is configured to construct and execute an action control instruction corresponding to the optimal motion coordination parameters to operate on the cargo on the conveyor belt.

[0054] It should be noted that when the robot multi-modal perception and motion coordination control device provided in the above embodiment executes the robot multi-modal perception and motion coordination control method, only the above division of each functional module is used for illustration. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the robot multi-modal perception and motion coordination control device provided in the above embodiment and the method embodiment of the robot multi-modal perception and motion coordination control belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0055] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0056] In the embodiments of the present application, on the one hand, through multi-modal perception data, the robot can obtain the operating state of the conveyor belt and the detailed information of the goods in real time, thereby dynamically adjusting the historical action parameters of the robotic arm to achieve precise and efficient sorting operations. Precise sorting operations can significantly improve the sorting efficiency, reduce the damage or sorting errors of goods caused by misoperations, and thus reduce the costs incurred due to frequent maintenance and repair of equipment. On the other hand, by dynamically adjusting the historical action parameters of the robotic arm to adapt to the real-time operating state of the conveyor belt and the characteristics of the goods, the robot can maintain a stable operation performance under different working conditions. This adaptive ability can effectively avoid operation interruptions or incorrect operations caused by problems such as changes in the conveyor belt speed, offset of the goods position, or wear of the conveyor belt, thereby ensuring the stability of the entire production process.

[0057] The present application also provides a computer-readable medium, on which program instructions are stored. When the program instructions are executed by a processor, the robotic arm multi-modal perception and motion cooperative control method provided by each of the above method embodiments is implemented.

[0058] The present application also provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the robotic arm multi-modal perception and motion cooperative control method of each of the above method embodiments.

[0059] Please refer to Figure 7 , which is a schematic structural diagram of an electronic device provided by the embodiments of the present application. As Figure 7 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.

[0060] Among them, the communication bus 1002 is used to realize the connection and communication between these components.

[0061] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.

[0062] Among them, the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0063] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire electronic device 1000 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling the data stored in the memory 1005, it performs various functions of the electronic device 1000 and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and may be implemented separately by a single chip.

[0064] Among them, the memory 1005 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1005 may also be at least one storage system located far from the aforementioned processor 1001. As Figure 7 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a robot multi-modal perception and motion coordination control application program.

[0065] In Figure 7In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an interface for the user to input and obtain the data input by the user. The processor 1001 can be used to call the robot multi-modal perception and motion coordination control application program stored in the memory 1005 and specifically perform the following operations: Obtain multi-class sensor information from multiple types of sensors deployed in the robot working area at a preset period as multi-modal perception data. The multi-modal perception data includes lidar data, infrared sensor data, and visual sensor data. A conveying device is provided in the robot working area; Analyze the conveying parameters of the conveyor belt of the conveying device and the cargo parameters of the goods on the conveyor belt based on the multi-modal perception data; Dynamically adjust the historical motion parameters of the robot's robotic arm according to the conveying parameters and the cargo parameters to obtain the optimal motion coordination parameters applicable to the conveying device; Construct and execute the motion control instruction corresponding to the optimal motion coordination parameters to operate on the goods on the conveyor belt.

[0066] In one embodiment, when the processor 1001 executes dynamically adjusting the historical motion parameters of the robot's robotic arm according to the conveying parameters and the cargo parameters to obtain the optimal motion coordination parameters applicable to the conveying device, it specifically performs the following operations: Convert the running speed and tilt angle into frequency domain features through dynamic Fourier transform; Construct a wear thermal map of the conveyor belt according to the wear data; Fuse the frequency domain features with the wear thermal map to obtain an environmental encoding vector; Use the attention mechanism to perform attention encoding on the center coordinates, size, and relative position to generate a cargo encoding vector for each piece of cargo; Input the environmental encoding vector and the cargo encoding vector into a pre-trained motion parameter analysis model to output the current motion parameters of the robot's robotic arm; In the case where the current motion parameters are inconsistent with the historical motion parameters of the robot's robotic arm, replace the historical motion parameters with the current motion parameters to obtain the optimal motion coordination parameters applicable to the conveying device.

[0067] In one embodiment, when the processor 1001 executes converting the running speed and tilt angle into frequency domain features through dynamic Fourier transform, it specifically performs the following operations: Filter, denoise, and normalize the running speed and tilt angle that change with time to obtain a preprocessed time domain signal of the conveying device; Use the dynamic Fourier transform expression to convert the time domain signal of the conveying device into a frequency domain signal to obtain the frequency spectrum data of the conveying device; Identify the main frequency components, the amplitude of each frequency component, and the phase of each frequency component in the spectrum data of the conveyor device as the frequency-domain information of the running speed and the tilt angle; Fuse the frequency-domain information of the running speed and the tilt angle to obtain frequency-domain features.

[0068] In one embodiment, when the processor 1001 executes to construct a wear heat map of the conveyor belt according to the wear data, the following operations are specifically performed: Extract the wear depth of the conveyor belt, the temperature change in the wear area, and the reflectivity change in the wear area from the wear data; Divide the surface of the conveyor belt into multiple small grids based on preset grid division parameters, and each small grid corresponds to a wear data point; Normalize the wear depth of the conveyor belt, the temperature change in the wear area, and the reflectivity change in the wear area and map them onto each small grid to form a wear data matrix; Generate a grid layer representing the wear degree by performing spatial smoothing processing on the wear data matrix; Assign different shades of color to the grid layer according to the magnitude of the wear data in the wear data matrix to obtain the wear heat map of the conveyor belt.

[0069] In one embodiment, when the processor 1001 executes to fuse the frequency-domain features with the wear heat map to obtain an environment encoding vector, the following operations are specifically performed: Extract the spatial information of the wear heat map; Map the spatial information of the wear heat map onto the frequency-domain features to splice the frequency-domain features with the spatial information of the wear heat map to obtain spliced features; Encode the spliced features to obtain an environment encoding vector.

[0070] In one embodiment, when the processor 1001 executes to input the environment encoding vector and the cargo encoding vector into a pre-trained action parameter analysis model and output the current action parameters of the robot's robotic arm, the following operations are specifically performed: The weight matrix loading layer loads the pre-trained weight matrix; The quantization layer uses the pre-trained weight matrix to quantize the environment encoding vector and the cargo encoding vector respectively to obtain a query vector, a key vector, and a value vector; The attention layer calculates attention scores according to the query vector, the key vector, and the value vector; the attention scores are used to measure the importance of the environment encoding vector to the cargo encoding vector; The dot product operation layer performs dot product operations on the environment encoding vector and the cargo encoding vector with the attention scores respectively and sums them to obtain a joint feature containing environment-cargo interaction information; The feature mapping layer maps the joint features to a preset robotic arm action space, obtaining a quantity that matches the degrees of freedom of the robotic arm in dimension. The clipping layer uses pre-stored physical constraints of the robotic arm to clip the quantity that matches the degrees of freedom of the robotic arm in dimension, obtaining the current action parameters of the robotic arm of the robot.

[0071] In one embodiment, when the processor 1001 executes to generate a pre-trained action parameter analysis model, it specifically performs the following operations: Simulate different conveyor belt working conditions in the robotic operation area in a simulation environment to obtain sample conveyor parameters; During the simulation process, randomly generate sample cargo parameters; Record the sequence of joint angular velocities of the robotic arm input by the user during the simulation process as a supervision signal; Construct a model training sample according to the supervision signal, sample conveyor parameters, and sample cargo parameters; Create and initialize an action parameter analysis model using a neural network. This model includes a weight matrix loading layer, a quantization layer, an attention layer, a dot product operation layer, a feature mapping layer, and a clipping layer; Perform machine learning on the initialized action parameter analysis model using the model training sample, so that the action parameter analysis model learns the weight association relationship between the sample conveyor parameters, sample cargo parameters, and the supervision signal, obtaining a pre-trained weight matrix; Quantify the loss value of the pre-trained weight matrix using the model loss function of the action parameter analysis model; When the loss value reaches the minimum, obtain the pre-trained action parameter analysis model.

[0072] In one embodiment, when the processor 1001 executes the process of constructing a model training sample according to the supervision signal, sample conveyor parameters, and sample cargo parameters, it specifically performs the following operations: Construct a sample environment encoding vector according to the sample conveyor parameters; Construct a sample cargo encoding vector according to the sample cargo parameters; Quantize the sample environment encoding vector and the sample cargo encoding vector into a sample query vector, a sample key vector, and a sample value vector respectively; Perform label annotation on the sample query vector, the sample key vector, and the sample value vector using the supervision signal to obtain a model training sample; where the loss function of the model is:

[0073] Among them, is the loss value, is the total number of samples of the model training sample, is the The predicted action parameters of the is the true action parameter characterized by the supervision signal of the th sample; where is the dimension of the attention score, is the predicted value of the th attention score, is the true value of the th attention score, is the regularization coefficient, used to control the strength of regularization, is the pre-trained weight matrix, where represents any layer in the model.

[0074] In the embodiments of the present application, on the one hand, through multi-modal perception data, the robot can obtain the operating state of the conveyor belt and the detailed information of the goods in real time, so as to dynamically adjust the historical action parameters of the robotic arm, realizing precise and efficient sorting operations. Precise sorting operations can significantly improve the sorting efficiency, reduce the damage or sorting errors of goods caused by misoperations, and thus reduce the costs generated by frequent maintenance and repair of equipment. On the other hand, by dynamically adjusting the historical action parameters of the robotic arm to adapt to the real-time operating state of the conveyor belt and the characteristics of the goods, the robot can maintain stable operation performance under different working conditions. This adaptive ability can effectively avoid operation interruptions or incorrect operations caused by problems such as changes in conveyor belt speed, offset of goods position, or wear of the conveyor belt, thus ensuring the stability of the entire production process.

[0075] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program for robot multi-modal perception and motion collaborative control can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium of the program for robot multi-modal perception and motion collaborative control can be a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0076] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited by this. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A multi-modal perception and motion collaborative control method for a robot, characterized in that, Applied to a robot, the method includes: Obtaining multi-class sensor information at a preset period through multi-class sensors deployed in the robot's working area as multi-modal perception data, where the multi-modal perception data includes lidar data, infrared sensor data, and visual sensor data, and a conveying device is provided in the robot's working area; Analyzing the conveying parameters of the conveyor belt of the conveying device and the cargo parameters of the cargo on the conveyor belt based on the multi-modal perception data; Dynamically adjusting the historical action parameters of the robot's robotic arm according to the conveying parameters and the cargo parameters to obtain optimal motion coordination parameters applicable to the conveying device; Constructing and executing an action control instruction corresponding to the optimal motion coordination parameters to operate on the cargo on the conveyor belt.

2. The method according to claim 1, wherein The conveying parameters include the running speed, inclination angle, and wear data of the conveyor belt; the cargo parameters include the central coordinates, size, and relative position with respect to the conveyor belt edge of the cargo on the conveyor belt; The dynamically adjusting the historical action parameters of the robot's robotic arm according to the conveying parameters and the cargo parameters to obtain optimal motion coordination parameters applicable to the conveying device includes: Converting the running speed and inclination angle into frequency domain features through dynamic Fourier transform; Constructing a wear heat map of the conveyor belt according to the wear data; Fusing the frequency domain features with the wear heat map to obtain an environment encoding vector; Using an attention mechanism to perform attention encoding on the central coordinates, size, and relative position to generate a cargo encoding vector for each cargo; Inputting the environment encoding vector and the cargo encoding vector into a pre-trained action parameter analysis model to output the current action parameters of the robot's robotic arm; In the case where the current action parameters are inconsistent with the historical action parameters of the robot's robotic arm, replacing the historical action parameters with the current action parameters to obtain optimal motion coordination parameters applicable to the conveying device.

3. The method according to claim 2, wherein The converting the running speed and inclination angle into frequency domain features through dynamic Fourier transform includes: Filtering, denoising, and normalizing the running speed and inclination angle that change over time to obtain a pre-processed time domain signal of the conveying device; Using a dynamic Fourier transform expression to convert the time domain signal of the conveying device into a frequency domain signal to obtain conveying device spectrum data; Identifying the main frequency components, the amplitude of each frequency component, and the phase of each frequency component in the conveying device spectrum data as the frequency domain information of the running speed and inclination angle; Fusing the frequency domain information of the running speed and inclination angle to obtain frequency domain features.

4. The method according to claim 2, wherein The constructing the wear heat map of the conveyor belt according to the wear data includes: Extracting the wear depth of the conveyor belt, the temperature change of the wear area, and the reflectivity change of the wear area from the wear data; Dividing the surface of the conveyor belt into multiple small grids based on preset grid division parameters, and each small grid corresponds to a wear data point; Normalize the wear depth of the conveyor belt, the temperature change in the worn area, and the reflectivity change in the worn area, and then map them onto each small grid to form a wear data matrix; Generate a raster layer representing the wear degree by performing spatial smoothing on the wear data matrix; According to the magnitude of the wear data in the wear data matrix, assign different shades of color to the raster layer to obtain the wear heat map of the conveyor belt.

5. The method according to claim 2, characterized in that, The fusion of the frequency domain features and the wear heat map to obtain an environmental coding vector includes: Extract the spatial information of the wear heat map; Map the spatial information of the wear heat map onto the frequency domain features to splice the frequency domain features and the spatial information of the wear heat map to obtain a spliced feature; Encode the spliced feature to obtain an environmental coding vector.

6. The method according to claim 2, wherein The pre-trained action parameter analysis model includes a weight matrix loading layer, a quantization layer, an attention layer, a dot product operation layer, a feature mapping layer, and a clipping layer; Inputting the environmental coding vector and the cargo coding vector into the pre-trained action parameter analysis model to output the current action parameters of the robot's robotic arm includes: The weight matrix loading layer loads the pre-trained weight matrix; The quantization layer uses the pre-trained weight matrix to quantize the environmental coding vector and the cargo coding vector respectively to obtain a query vector, a key vector, and a value vector; The attention layer calculates an attention score according to the query vector, the key vector, and the value vector; the attention score is used to measure the importance of the environmental coding vector to the cargo coding vector; The dot product operation layer performs dot product operations on the environmental coding vector and the cargo coding vector with the attention score respectively and sums them to obtain a joint feature containing environmental and cargo interaction information; The feature mapping layer maps the joint feature to a preset robotic arm action space to obtain a dimension matching the degree of freedom of the robotic arm; The clipping layer uses the pre-stored physical constraints of the robotic arm to clip the dimension matching the degree of freedom of the robotic arm to obtain the current action parameters of the robot's robotic arm.

7. The method according to claim 6, wherein The quantization formula is: ; Among them, is the query vector, is the key vector, is the value vector, ( ) is the pre-trained weight matrix for the query vector, key vector, and value vector of the model, is the environment encoding vector, is the goods encoding vector; Among them, the calculation formula for the attention score is: Among them, is the attention score, is the dimension of the key vector, which is used to scale the dot product result; among them, the mapping formula is: ; Among them, is the matching quantity of the dimension and the degrees of freedom of the robotic arm, is the weight matrix of the output layer, is the bias term of the output layer, is the joint feature; among them, the expression of clipping is: ; Among them, is the maximum allowable angular velocity of each joint of the robotic arm, is the clipping function of the model.

8. The method according to any one of claims 2-7, characterized in that Generate a pre-trained action parameter analysis model according to the following steps, including: Simulate different conveyor belt conditions in the robot operation area in a simulation environment to obtain sample conveyor parameters; During the simulation process, randomly generate sample cargo parameters; Record the sequence of joint angular velocities of the robotic arm input by the user during the simulation process as a supervision signal; Construct a model training sample according to the supervision signal, the sample conveyor parameters, and the sample cargo parameters; Create and initialize an action parameter analysis model using a neural network. This model includes a weight matrix loading layer, a quantization layer, an attention layer, a dot product operation layer, a feature mapping layer, and a clipping layer; Use the model training sample to perform machine learning on the initialized action parameter analysis model so that the action parameter analysis model learns the weight correlation relationship between the sample conveyor parameters, the sample cargo parameters, and the supervision signal to obtain a pre-trained weight matrix; Quantify the loss value of the pre-trained weight matrix by using the model loss function of the action parameter analysis model; Obtain the pre-trained action parameter analysis model when the loss value reaches the minimum.

9. The method according to claim 8, characterized in that Constructing a model training sample according to the supervision signal, the sample transfer parameter, and the sample cargo parameter includes: Construct a sample environment encoding vector according to the sample transfer parameter; Construct a sample cargo encoding vector according to the sample cargo parameter; Quantize the sample environment encoding vector and the sample cargo encoding vector into a sample query vector, a sample key vector, and a sample value vector respectively; Use the supervision signal to label the sample query vector, the sample key vector, and the sample value vector to obtain a model training sample; wherein, the loss function of the model is: Among them, is the loss value, is the total number of samples of the model training samples, is the predicted action parameter of the th sample, is the true action parameter of the supervised signal representation of the th sample; among them, is the dimension of the attention score, is the predicted value of the th attention score, is the true value of the th attention score, is the regularization coefficient, used to control the strength of regularization, is the pre-trained weight matrix, where represents any layer in the model.

10. A robot multi-modal perception and motion cooperative control device, characterized in that, The device includes: A sensor information acquisition module, configured to acquire multi-class sensor information at a preset period through multi-class sensors arranged in the robot operation area as multi-modal perception data, the multi-modal perception data including lidar data, infrared sensor data, and visual sensor data, and a conveying device is provided in the robot operation area; A parameter analysis module, configured to analyze the conveying parameter of the conveyor belt of the conveying device and the cargo parameter of the cargo on the conveyor belt based on the multi-modal perception data; An action parameter adjustment module, configured to dynamically adjust the historical action parameters of the robot's robotic arm according to the conveying parameter and the cargo parameter to obtain optimal motion coordination parameters applicable to the conveying device; An instruction execution module, configured to construct and execute an action control instruction corresponding to the optimal motion coordination parameter to operate on the cargo on the conveyor belt.

Citation Information

Patent Citations

  • Robot motion control method, motion control device and robot system

    CN110799911A

  • Energy-saving control method and system for belt conveyor based on transportation volume

    CN118850677A

  • Robot control method and device based on multi-modal data fusion

    CN119260752A

  • Rolling bearing fault diagnosis method based on time-frequency image and transfer learning

    CN120086689A

  • EA202490400A1

Cited By

  • Conveying system and method based on AI commodity identification

    CN120853110A

  • Robot autonomous task planning and execution control method based on reinforcement learning

    CN121004603A

  • A Reinforcement Learning-Based Method for Autonomous Task Planning and Execution Control of Robots

    CN121004603B

  • Multi-robot cooperative control method and system for industrial control

    CN121028705A