Self-adaptive gravity center adjusting method and system for transfer robot
Feature extraction through multimodal sensor data acquisition and deep learning algorithm, combined with Q-learning optimization algorithm, the problem that the transporting robot is difficult to perceive the state of the object is solved, the robot is high stability and adaptability in the dynamic environment, and the success rate and efficiency of the transporting task are improved.
Patent Information
- Application Number
- CN202510475183.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing transport robots rely on a single sensor or limited sensor data, making it difficult to sense the state of an object comprehensively and accurately. Especially when the center of gravity of the object or the contact conditions suddenly change, it is difficult for the robot to quickly adjust its operating strategy, resulting in tilting, falling or damage to the object.
Multimodal sensor data acquisition is adopted, including three-dimensional information, contact force distribution information and contact force information between the robot and the object. Visual and tactile feature vectors are extracted through deep learning algorithms, and weighted average fusion is carried out to build a motion state model, real-time feedback is combined with Q-learning optimization algorithm, and control strategies are adjusted.
It realizes the robot's comprehensive and accurate perception of the object state, improves the stability and adaptability of the robot in a dynamic environment, reduces the risk of tilting, falling or damage of items, and improves the success rate and efficiency of handling tasks.
Smart Images

Figure CN119974028A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a method and system for adaptively adjusting the center of gravity of a handling robot. Background Art
[0002] In recent years, with the rapid development of intelligent manufacturing, logistics automation and service robot technology, handling robots, as an important part of the automation system, have been widely used in many industries. The main task of handling robots is to grasp, carry and place objects accurately and quickly. However, problems such as unstable center of gravity of objects, diversity of grasping and irregularity of object surface often make the stability and efficiency of the handling process face challenges. In order to improve the reliability and flexibility of handling robots, related technologies have gradually developed towards adaptive control and multimodal sensing technology. Adaptive control technology enables robots to adjust their operating strategies in real time according to changes in the environment and objects, while multimodal sensing technology provides robots with rich perceptual inputs by combining visual, tactile and force information, which can help robots perform precise operations in more complex environments.
[0003] Although the existing technology has made significant progress, there are still some shortcomings. Most of the existing handling robots rely on a single sensor or limited sensor data, and it is difficult to fully and accurately perceive the state of objects. For example, although traditional visual sensors can provide information such as the shape and size of objects, they have limited perception capabilities for dynamic characteristics such as the weight and center of gravity of objects. Force sensors can measure contact force, but it is difficult to provide full information about the object. In addition, the existing center of gravity adjustment methods mainly rely on preset control rules or empirical models, lacking sufficient flexibility and adaptability. In particular, when the center of gravity position or contact conditions of the object suddenly change, it is difficult for the robot to quickly adjust its operating strategy, which may cause the object to tilt, fall or be damaged. Summary of the invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for adaptively adjusting the center of gravity of a transport robot to solve the problem that most existing transport robots rely on a single sensor or limited sensor data and are difficult to fully and accurately perceive the state of an object.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for adaptively adjusting the center of gravity of a handling robot, which comprises: The sensors are used to collect the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object, and the collected information is integrated to obtain a multimodal data set and preprocessed; Using deep learning algorithms, visual feature vectors are extracted from 3D image information, and tactile feature vectors are extracted from contact force distribution information and contact force information; The extracted visual feature vector and tactile feature vector are fused to obtain a multimodal feature set; Construct a motion state model based on a deep learning algorithm, use historical multimodal feature sets to input the motion state model, and output a control strategy; Utilize real-time feedback information to optimize control algorithms and control strategies.
[0007] As a preferred solution of the method for adaptively adjusting the center of gravity of the handling robot of the present invention, the sensors are used to respectively collect the three-dimensional information of the object, the contact force distribution information and the contact force information between the robot and the object, and the collected information is integrated to obtain a multimodal data set and preprocessed, and the specific steps are as follows: The RGB-D camera is used to obtain the three-dimensional image information of the object. At the same time, a pressure sensor array is arranged on the robot end effector to collect the contact force distribution information. In addition, a torque sensor is installed on the robot's joint actuator to measure the contact force information between the robot and the object. Combining three different types of data into a multimodal dataset , the expression is: ; in, Indicates at time 3D image information, Indicates at time The contact force distribution information, Indicates at time contact force information.
[0008] As a preferred solution of the method for adaptively adjusting the center of gravity of a handling robot according to the present invention, wherein: Select a dimension of time series data from the original signal ; Apply continuous wavelet transform (CWT) to the original signal To decompose; Through continuous wavelet transform CWT, a series of different scales are obtained The detail coefficient and approximation coefficient under ; Set the noise threshold. When the detail coefficient at a certain scale is less than the noise threshold, this part is considered to be noise and is set to zero. The signal is reconstructed using the denoised detail coefficients and approximate coefficients. The expression is: ; in, represents the reconstructed signal, represents the inverse continuous wavelet transform, represents the detail coefficient after noise threshold processing, represents the approximate coefficient, An index variable representing the scale, Indicates the maximum scale index.
[0009] As a preferred solution of the method for adaptively adjusting the center of gravity of the handling robot described in the present invention, a convolutional neural network CNN is used to obtain the three-dimensional image information. Extract visual feature vectors from ; Using long short-term memory network LSTM to obtain contact force distribution information and contact force information Extract tactile feature vectors from .
[0010] As a preferred solution of the method for adaptively adjusting the center of gravity of a handling robot according to the present invention, the extracted visual feature vector and tactile feature vector are fused to obtain a multimodal feature set, and the specific steps are as follows: The weighted average method is used to fuse the visual feature vector and the tactile feature vector to obtain a multimodal feature set. , the expression is: ; in, and Represent visual feature vectors and the tactile feature vector The weight coefficient of Indicates time, Indicates that the item is at time The visual feature vector of Indicates that the item is at time The tactile feature vector of .
[0011] As a preferred solution of the method for adaptive center of gravity adjustment of the handling robot of the present invention, wherein: the motion state model is constructed based on the deep learning algorithm, the motion state model is input using the historical multimodal feature set, and the control strategy is output, and the specific steps are: The motion state model is constructed based on the combination of long short-term memory network LSTM and multi-layer perceptron MLP. The long short-term memory network LSTM is used to process the dynamic characteristics of time series data and can capture the dependency of multimodal features on the time axis. The multi-layer perceptron MLP is used to process the fused multimodal features and learn the complex mapping relationship between the features. Collecting historical multimodal feature sets , and input it into the motion state model, and output the control strategy, the expression is: ; in, Indicates the robot at time The control strategies adopted represents the activation function Sigmoid, represents the weight matrix of the output layer of the multilayer perceptron MLP, Represents the bias term of the multi-layer perceptron MLP output layer; The control strategy to be output It is transmitted to the robot actuator to guide the robot's movement and posture adjustment.
[0012] As a preferred solution of the method for adaptively adjusting the center of gravity of the handling robot of the present invention, the specific steps of utilizing real-time feedback information to optimize the control algorithm and control strategy are as follows: When the robot performs tasks, it collects real-time feedback information based on sensors; The original signal collected by the sensor , the original signal After a series of preprocessing steps, the current state of the robot is finally obtained ; Based on the current state of the robot and control strategies , the Q-value function is used to evaluate the quality of the control strategy. The goal is to maximize long-term returns. The Q-learning update formula is: ; in, Indicates the current status and control strategies The Q value under represents the learning rate, Indicates at time Instant rewards, represents the discount factor, represents the updated control strategy, Indicates that the selection is in state The control strategy that can obtain the maximum future reward is: Indicates the robot at time status, Indicates the robot at time The control strategies adopted; according to The control strategy that produces the maximum Q value in a given state is selected to guide the robot's actions.
[0013] In a second aspect, the present invention provides a handling robot adaptive center of gravity adjustment system, including: a collection module, a feature extraction module, a strategy generation module, a feedback module and an optimization module: The acquisition module is responsible for acquiring and preprocessing the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object; The feature extraction module is responsible for extracting visual feature vectors from the three-dimensional image data and tactile feature vectors from the contact force distribution and contact force information. Then, these features are fused by weighted averaging to generate a final multimodal feature set. The strategy generation module is used to construct a robot motion state model based on the extracted multimodal feature set and predict the robot's motion state; The feedback module is used to collect and analyze feedback information during the execution of the robot in real time to optimize the control strategy; The optimization module is responsible for continuously collecting new sensor data, updating historical data sets, and adjusting the motion state model through optimization algorithms.
[0014] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for adaptive center of gravity adjustment of a handling robot as described in the first aspect of the present invention is implemented.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the method for adaptive center of gravity adjustment of a transport robot as described in the first aspect of the present invention is implemented.
[0016] The beneficial effects of the present invention are as follows: the present invention collects multimodal sensor data and combines it with a deep learning algorithm for feature extraction and fusion, and proposes an adaptive center of gravity adjustment method. By using an RGB-D camera, a pressure sensor array and a torque sensor, the interaction information between the object and the robot can be fully and accurately obtained, thereby providing a reliable basis for subsequent control. Wavelet transform is used to denoise and enhance signals to ensure high data quality. The convolutional neural network CNN and the long short-term memory network LSTM are combined to effectively extract visual and tactile features. Then, through weighted average fusion, a multimodal feature set is obtained to accurately reflect the state of the object. Based on these features, the constructed motion state model can predict and adaptively adjust the robot control strategy. Real-time feedback combined with the Q-learning optimization algorithm further improves the robot's decision-making ability and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0018] Figure 1 This is a flow chart of the method for adaptively adjusting the center of gravity of the transport robot in Example 1.
[0019] Figure 2 Schematic diagram of the adaptive center of gravity adjustment system of the transport robot in Example 1. DETAILED DESCRIPTION
[0020] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0022] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0023] Example 1, reference Figure 1 and Figure 2, which is the first embodiment of the present invention, provides a method for adaptively adjusting the center of gravity of a handling robot, comprising the following steps: S1. Using sensors to collect the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object, respectively, and integrating the collected information to obtain a multimodal data set and preprocessing it; The RGB-D camera is used to obtain three-dimensional image information of the object, which provides key visual features about the shape and size of the object. At the same time, a pressure sensor array is arranged on the robot's end effector to collect contact force distribution information, which reflects the pressure situation at the contact point between the object surface and the robot. In addition, a torque sensor is installed on the robot's joint actuator to measure the contact force information between the robot and the object, which is crucial for understanding the object's weight distribution and grasping stability. Combining three different types of data into a multimodal dataset , the expression is: ; in, Indicates at time 3D image information, Indicates at time The contact force distribution information, Indicates at time Contact force information; The three-dimensional image information includes the distance value and color value of each pixel point, the contact force distribution information includes the pressure value of each sensor unit, and the contact force information between the robot and the object includes the force components along the X-axis, Y-axis and Z-axis; From multimodal datasets The time series data of a dimension selected in is used to represent the change of the physical quantity under this dimension over time, and the original signal is obtained ; Considering that the data provided by different sensors contain noise or missing values, an effective algorithm is needed to perform denoising and enhancement processing. For this step, a method based on wavelet transform is selected as the optimal solution for signal processing. Wavelet transform can effectively separate the high-frequency noise components in the signal while retaining the effective low-frequency information, thereby improving the accuracy of subsequent analysis; Specifically, the continuous wavelet transform (CWT) is applied to the original signal Decomposed, the expression is: ; in, Represents the original signal after wavelet transform At different scales The transformation result is: represents the index of the scale, represents the integral variable, Represents the mother wavelet function In scale The complex conjugate of represents differential elements; Through continuous wavelet transform CWT, a series of different scales are obtained The detail coefficient and approximation coefficient under ; Set a noise threshold. When the detail coefficient at a certain scale is less than the noise threshold, this part is considered to be noise and is set to zero. This hard threshold method can effectively remove high-frequency noise while retaining effective low-frequency information. The signal is reconstructed using the denoised detail coefficients and approximate coefficients to achieve the final denoising effect. The expression is: ; in, represents the reconstructed signal, represents the inverse continuous wavelet transform, represents the detail coefficient after noise threshold processing, represents the approximate coefficient, An index variable representing the scale, Indicates the maximum scale index.
[0024] S2. Using deep learning algorithms, extract visual feature vectors from three-dimensional image information; extract tactile feature vectors from contact force distribution information and contact force information; Using convolutional neural network CNN from 3D image information Extract visual feature vectors from ; By extracting visual features, the geometric model of the object is established to provide accurate spatial information for center of gravity adjustment. This solves the problem that image feature processing in traditional methods is not sufficient to reflect the shape of complex objects. CNN can analyze layer by layer from local details to overall structure, avoiding the limitations of manually set features, improving the recognition accuracy and adaptability of complex objects, and providing more reliable input data for subsequent posture adjustment and force control strategies. Using long short-term memory network LSTM to obtain contact force distribution information and contact force information Extract tactile feature vectors from ; By dynamically analyzing tactile data, the force model of the object is established to support the robot to adjust the grasping strategy and posture control, solving the problem that the traditional force control system can only process static data and is difficult to capture dynamic force changes. LSTM analyzes the connection between historical data and the current state through the memory and forgetting mechanism, provides more flexible force perception capabilities, enhances the robot's perception of dynamic force changes, improves its adaptability and stability to complex handling tasks, and helps maintain balance and grasping stability in irregular objects or dynamic environments; Through the step-by-step extraction of visual and tactile features and multi-dimensional data processing, the present invention has significant advantages in adjusting the center of gravity of the robot. The visual features provide geometric information support, while the tactile features capture dynamic force changes. The combination of the two can form a more comprehensive description of the object state. This multimodal perception solution overcomes the defect that a single sensor data is not sufficient to accurately reflect complex scenes, and significantly improves the robot's environmental adaptability and task completion rate. Especially in scenarios that require high-precision posture adjustment (such as precision assembly or flexible handling), the present invention provides a more reliable and intelligent solution.
[0025] S3, fusing the extracted visual feature vector and tactile feature vector to obtain a multimodal feature set; The weighted average method is used to fuse the visual feature vector and the tactile feature vector to obtain a multimodal feature set. , the expression is: ; in, and Represent visual feature vectors and the tactile feature vector The weight coefficient of Indicates time, Indicates that the item is at time The visual feature vector of Indicates that the item is at time The tactile feature vector of The multimodal feature fusion method proposed in the present invention overcomes the inadequacy of the expressive ability of single-mode features by weighted combination of visual and tactile features. In traditional technologies, visual or tactile features are usually processed separately and cannot provide a description of the global state of the object, resulting in unstable or inaccurate control strategies. By using multimodal feature fusion and integrating visual and tactile information, the robot is provided with a more complete object perception capability, which solves the limitation of single sensor data expression, especially the problem of delayed response of visual features to dynamic environment and lack of global information of tactile features. Combining the spatial distribution advantage of visual information with the dynamic feedback advantage of tactile information improves the robot's environmental adaptability and control accuracy; Step S3 fuses visual and tactile features through weighted averaging to construct a multimodal feature set , plays a core role in the robot control algorithm. This design not only improves the data expression ability and adaptability, but also introduces dynamic time parameters and weight coefficient and , which solves the problem of ignoring the importance of different features in traditional methods. Its innovation lies in making full use of the complementarity of visual and tactile information, enabling the robot to maintain high-precision control and adaptive adjustment in dynamic and complex environments to meet practical application needs.
[0026] S4, build a motion state model based on deep learning algorithm, use historical multimodal feature set to input the motion state model, and output control strategy; The motion state model is constructed based on the combination of long short-term memory network LSTM and multi-layer perceptron MLP. The long short-term memory network LSTM is used to process the dynamic characteristics of time series data and can capture the dependency of multimodal features on the time axis. The multi-layer perceptron MLP is used to process the fused multimodal features and learn the complex mapping relationship between the features. In this way, the motion state model can predict the current motion state of the robot based on historical data, thereby achieving adaptive adjustment. LSTM is suitable for processing time series data and can capture the impact of past information on current predictions. It is particularly suitable for robot kinematics and dynamic control problems. MLP can learn multiple input features in parallel and is good at extracting high-dimensional information from different features. Especially when we need to process fused visual and tactile features, the advantages of MLP are more obvious. Collecting historical multimodal feature sets , and input it into the motion state model, and output the control strategy, the expression is: ; in, , represents the activation function Sigmoid, represents the weight matrix of the output layer of the multilayer perceptron MLP, Represents the bias term of the multi-layer perceptron MLP output layer; The control strategy to be output The robot actuator guides the robot's movement and posture adjustment. During the execution process, the robot corrects its motion state based on real-time feedback (such as the output of pressure sensors and torque sensors) and updates the historical feature set. , continuously optimize the motion state model; Control strategy The range of is [0,1], which indicates the probability of the control strategy. For example, the probability of the control strategy pointing to a certain posture adjustment, and the value in the range will indicate the probability of performing the action; Weight Matrix It is learned through the back propagation algorithm. Specifically, the motion state model is based on the training data (i.e., the historical multimodal feature set and corresponding control strategies Keep adjusting And other network parameters, during the training process, the optimization algorithm Adam is used to minimize the loss function, thereby updating and other parameters.
[0027] S5. Use real-time feedback information to optimize control algorithms and control strategies; When the robot performs a task, it collects real-time feedback information based on sensors, including real-time contact force information, real-time contact force distribution information, and real-time motion state information; Based on the collected real-time feedback information, Q-learning in reinforcement learning is used to update the motion state model. Q-learning can evaluate the effect of the current control strategy and optimize future decisions through real-time feedback information; The original signal collected by the sensor , the original signal After a series of preprocessing steps (such as denoising, wavelet transform, CNN and LSTM), the current state of the robot is finally obtained , including visual feature vectors, tactile feature vectors and multimodal feature sets; Based on the current state of the robot and control strategies , the Q-value function is used to evaluate the quality of the control strategy. The goal is to maximize the long-term return (i.e., the cumulative reward). The Q-learning update formula is: ; in, Indicates the current status and control strategies The Q value under represents the learning rate, Indicates at time The immediate reward is the feedback value obtained by the robot after executing the control strategy in this state. represents the discount factor, represents the updated control strategy, Indicates that the selection is in state The control strategy that can obtain the maximum future reward is: Indicates the robot at time status, Indicates the robot at time The control strategies adopted; according to The control strategy that produces the maximum Q value in a given state is selected to guide the robot's actions.
[0028] This embodiment also provides a handling robot adaptive center of gravity adjustment system, including: a collection module, a feature extraction module, a strategy generation module, a feedback module and an optimization module: The acquisition module is responsible for collecting and preprocessing the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object; The feature extraction module is responsible for extracting visual feature vectors from 3D image data and tactile feature vectors from contact force distribution and contact force information. These features are then fused through weighted averaging to generate the final multimodal feature set. A strategy generation module is used to build a robot motion state model based on the extracted multimodal feature set and predict the robot's motion state; Feedback module, used to collect and analyze feedback information during the robot execution process in real time and optimize the control strategy; The optimization module is responsible for continuously collecting new sensor data, updating historical data sets, and adjusting the motion state model through optimization algorithms.
[0029] This embodiment also provides a computer device, which is suitable for the case of an adaptive center of gravity adjustment method for a handling robot, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the adaptive center of gravity adjustment method for a handling robot as proposed in the above embodiment.
[0030] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0031] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for realizing adaptive center of gravity adjustment of a handling robot as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0032] In summary, the present invention collects multimodal sensor data and combines it with a deep learning algorithm to extract and fuse features, and proposes an adaptive center of gravity adjustment method. By using an RGB-D camera, a pressure sensor array and a torque sensor, the interaction information between the object and the robot can be fully and accurately acquired, thereby providing a reliable basis for subsequent control. Wavelet transform is used to denoise and enhance signals to ensure high data quality. Convolutional neural network CNN and long short-term memory network LSTM are combined to effectively extract visual and tactile features. Multimodal feature sets are obtained through weighted average fusion to accurately reflect the state of the object. Based on these features, the constructed motion state model can predict and adaptively adjust the robot control strategy. Real-time feedback combined with Q-learning optimization algorithm further improves the robot's decision-making ability and stability.
[0033] Example 2, referring to Table 1, is the second example of the present invention. To further verify the technical solution of the present invention, experimental simulation data of the adaptive center of gravity adjustment method of the handling robot are provided.
[0034] This experiment aims to verify the effectiveness and superiority of the robot control method based on multimodal feature fusion in complex object grasping and posture adjustment. The experimental platform includes: Robotic system: A six-degree-of-freedom robotic arm equipped with an end effector and force sensor.
[0035] Visual sensor: Depth camera, used to capture three-dimensional image information of objects.
[0036] Tactile sensor: An array of pressure sensors mounted on the end of a robotic arm to measure contact force distribution.
[0037] Data processing: Use high-performance processors.
[0038] The experimental objects are three objects of different shapes and surfaces, including regular rectangular objects (object A), cylindrical objects (object B), and objects with irregular surfaces (object C). The experimental process is as follows: Data collection phase: Use RGB-D cameras to collect 3D image information of objects , with a resolution of 1280x720 and a frame rate of 30fps, recording object size and shape characteristics.
[0039] Use pressure sensors to record contact force distribution information , with a sampling rate of 200 Hz, to measure the pressure distribution between the actuator and the object.
[0040] Torque sensor collects contact force information between the robot and the object , with a resolution of 0.1N.
[0041] All collected data are synchronized and a multimodal dataset is constructed .
[0042] Data preprocessing stage: The collected data was denoised and the continuous wavelet transform (CWT) was used to decompose and reconstruct the signal. The threshold was set to 1.5σ (standard deviation).
[0043] The denoised signals are resynthesized to ensure high signal-to-noise ratio for both visual and tactile data.
[0044] Feature extraction stage: Visual features are extracted through convolutional neural network (CNN) and output as visual feature vectors .
[0045] The tactile features are extracted through the long short-term memory network LSTM, and the tactile feature vector is output. .
[0046] Multimodal feature fusion stage: According to the expression , the visual and tactile features are integrated by weighted averaging, and the weights are set =0.6, =0.4 to enhance the responsiveness of dynamic information during the crawling process.
[0047] Control strategy generation and execution phase: Input fusion feature set To LSTM-MLP model, generate control strategy .
[0048] The control strategy is optimized in real time according to Q-learning, and the grasping posture is corrected by the immediate reward r(t).
[0049] The details are shown in Table 1 below: Table 1 Experimental data and comparative analysis table Experimental data show that the proposed method is superior to existing methods in key indicators such as grasping success rate, number of posture adjustments, average force error and object stability score. The following is a specific analysis of each indicator: Grasping success rate: The success rate of the present invention is generally about 10-23 percentage points higher than that of existing methods, especially showing significant advantages on objects C with irregular surfaces, which verifies the adaptability of multimodal feature fusion to complex objects; Number of posture adjustments: The existing method requires multiple adjustments to the grasping posture to achieve a stable state, while the present invention reduces the number of adjustments to 3-4 times based on real-time feedback and optimization control strategy of multimodal data, thereby improving execution efficiency; Average force error: The average error of the present invention is controlled between 0.12-0.20N, which is significantly lower than 0.45-0.72N of the existing method. This shows that the fusion of visual and tactile features effectively improves the perception accuracy of the force state of the object; Data processing and execution time: The present invention reduces the data processing and execution phases by 20-35ms and 30-45ms respectively, which is mainly due to the efficient denoising of wavelet transform and the fast feature extraction of deep learning model; Object stability score: The score of the method of the present invention is above 9.0, far exceeding the average score of 6.5 of existing methods, which further proves that multimodal feature fusion can improve the robot's grasping stability.
[0050] It can be seen from the experimental results that the present invention has significant innovation in the field of complex object grasping and posture control through multimodal data fusion, feature extraction and dynamic control strategy optimization. Compared with the traditional single sensor or independent processing mode, the present invention makes full use of the complementarity of visual and tactile features, so that the robot shows higher robustness and adaptability in dynamic environments. Especially in the grasping of irregular or complex-surface objects, traditional methods are prone to frequent or failed posture adjustments, while the present invention significantly improves the control accuracy and task success rate through real-time feedback and optimization mechanisms. Therefore, the invention provides a more efficient and reliable solution for application scenarios such as automated production lines, intelligent handling and precision assembly.
[0051] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for adaptively adjusting the center of gravity of a handling robot, characterized in that: include: The sensors are used to collect the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object, and the collected information is integrated to obtain a multimodal data set and preprocessed; Using deep learning algorithms, visual feature vectors are extracted from 3D image information, and tactile feature vectors are extracted from contact force distribution information and contact force information; The extracted visual feature vector and tactile feature vector are fused to obtain a multimodal feature set; Construct a motion state model based on a deep learning algorithm, use historical multimodal feature sets to input the motion state model, and output a control strategy; Utilize real-time feedback information to optimize control algorithms and control strategies.
2. The method for adaptively adjusting the center of gravity of a handling robot according to claim 1, characterized in that: The sensors are used to respectively collect the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object, and the collected information is integrated to obtain a multimodal data set and preprocessed. The specific steps are: The RGB-D camera is used to obtain the three-dimensional image information of the object. At the same time, a pressure sensor array is arranged on the robot end effector to collect the contact force distribution information. In addition, a torque sensor is installed on the robot's joint actuator to measure the contact force information between the robot and the object. Combining three different types of data into a multimodal dataset , the expression is: ; in, Indicates at time Three-dimensional image information, Indicates at time The contact force distribution information, Indicates at time contact force information.
3. The method for adaptively adjusting the center of gravity of a handling robot according to claim 2, characterized in that: The pre-processing comprises the following specific steps: From multimodal datasets Select a dimension of time series data from the original signal ; Apply continuous wavelet transform (CWT) to the original signal To decompose; Through continuous wavelet transform CWT, a series of different scales are obtained The detail coefficient and approximation coefficient under ; Set the noise threshold. When the detail coefficient at a certain scale is less than the noise threshold, this part is considered to be noise and is set to zero. The signal is reconstructed using the denoised detail coefficients and approximate coefficients. The expression is: ; in, represents the reconstructed signal, represents the inverse continuous wavelet transform, represents the detail coefficient after noise threshold processing, represents the approximate coefficient, An index variable representing the scale, Indicates the maximum scale index.
4. The method for adaptively adjusting the center of gravity of a handling robot according to claim 3, characterized in that: The method uses a deep learning algorithm to extract visual feature vectors from three-dimensional image information; and extracts tactile feature vectors from contact force distribution information and contact force information. The specific steps are as follows: Using convolutional neural network CNN from 3D image information Extract visual feature vectors from ; Using long short-term memory network LSTM to obtain contact force distribution information and contact force information Extract tactile feature vectors from .
5. The method for adaptively adjusting the center of gravity of a handling robot according to claim 4, characterized in that: The extracted visual feature vector and tactile feature vector are fused to obtain a multimodal feature set, and the specific steps are as follows: The weighted average method is used to fuse the visual feature vector and the tactile feature vector to obtain a multimodal feature set. , the expression is: ; in, and Represent visual feature vectors and the tactile feature vector The weight coefficient of Indicates time, Indicates that the item is at time The visual feature vector of Indicates that the item is at time The tactile feature vector of .
6. The method for adaptively adjusting the center of gravity of a handling robot according to claim 5, characterized in that: The motion state model is constructed based on the deep learning algorithm, the motion state model is input using the historical multimodal feature set, and the control strategy is output. The specific steps are: The motion state model is constructed based on the combination of long short-term memory network LSTM and multi-layer perceptron MLP; Collecting historical multimodal feature sets , and input it into the motion state model, and output the control strategy, the expression is: ; in, Indicates the robot at time The control strategies adopted represents the activation function Sigmoid, represents the weight matrix of the output layer of the multilayer perceptron MLP, Represents the bias term of the multi-layer perceptron MLP output layer; The control strategy to be output It is transmitted to the robot actuator to guide the robot's movement and posture adjustment.
7. The method for adaptively adjusting the center of gravity of a handling robot according to claim 6, characterized in that: The specific steps of utilizing real-time feedback information to optimize the control algorithm and control strategy are as follows: When the robot performs tasks, it collects real-time feedback information based on sensors; The original signal collected by the sensor , the original signal After a series of preprocessing steps, the current state of the robot is finally obtained ; Based on the current state of the robot and control strategies , the Q-value function is used to evaluate the quality of the control strategy. The goal is to maximize long-term returns. The Q-learning update formula is: ; in, Indicates the current status and control strategies The Q value under represents the learning rate, Indicates at time Instant rewards, represents the discount factor, represents the updated control strategy, Indicates that the selection is in state The control strategy that can obtain the maximum future reward is: Indicates the robot at time status, Indicates the robot at time The control strategies adopted; according to The control strategy that produces the maximum Q value in a given state is selected to guide the robot's actions.
8. A handling robot adaptive center of gravity adjustment system, based on the handling robot adaptive center of gravity adjustment method according to any one of claims 1 to 7, characterized in that: include: Acquisition module, feature extraction module, strategy generation module, feedback module and optimization module: The acquisition module is responsible for acquiring and preprocessing the three-dimensional information of the object, the contact force distribution information, and the contact force information between the robot and the object; The feature extraction module is responsible for extracting visual feature vectors from the three-dimensional image data and tactile feature vectors from the contact force distribution and contact force information. Then, these features are fused by weighted averaging to generate a final multimodal feature set. The strategy generation module is used to construct a robot motion state model based on the extracted multimodal feature set and predict the robot's motion state; The feedback module is used to collect and analyze feedback information during the execution of the robot in real time to optimize the control strategy; The optimization module is responsible for continuously collecting new sensor data, updating historical data sets, and adjusting the motion state model through optimization algorithms.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for adaptively adjusting the center of gravity of a handling robot according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for adaptively adjusting the center of gravity of a handling robot according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-mode object grabbing method and system based on combination of touch and vision
CN111055279A
Multi-degree-of-freedom auxiliary outer limb grabbing robot system fused with visual touch active perception
CN114131635A
Mechanical arm control method and system based on multi-mode driving and storage medium
CN118752495A
Efficient robot vision system based on deep learning and multi-modal fusion
CN118865042A
Screw flaw detection and feeding control method and system based on machine vision
CN119016362A
Cited By
Microscopic positioning control method and system for self-adaptive vision and force sense fusion
CN120259435A
Dexterous hand multi-object stable grabbing method and system based on vision and touch perception
CN120395966A
Robot joint positioning method and device and storage medium
CN121061841A
Logistics robot sensor system and data processing method thereof
CN121083660A
Sensor system of a logistics robot and data processing method thereof
CN121083660B