Robot control method, control system and computer readable storage medium
By adding image encoding of the image coded stream in the robot control system, aligning it with the force signal stream, and using the prediction model to obtain deviation correction parameters, the problem of low robot control accuracy in the prior art is solved, and higher control accuracy and system stability are achieved.
Patent Information
- Application Number
- CN202510348004.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing robot control technology has the problem of low control accuracy in the combination of vision and force sense, especially when the image may be blocked during the contact of the target, resulting in the missing feedback of the target position and state, and the force sense signal will not generate an effective signal before contacting the target.
By adding image encoding in the image encoding stream, aligning it with the force signal stream, forming a fused data stream, and inputting the target data stream into the prediction model to obtain bias correction parameters, adjusting the operation of the robot's future moment.
It ensures the precise coordination of image and force-sensing information at the same time point, eliminates errors caused by signal misalignment, improves robot control accuracy, enhances system stability, and reduces errors and interrupts during operation.
Smart Images

Figure CN120190816A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of robot control, and particularly to a control method, a control system, and a computer-readable storage medium for a robot. Background Art
[0002] Vision and force sensing are two of the most commonly used perception methods by humans. Vision contains rich information, especially in terms of target localization, including space, geometric shape, color, texture, etc., and is the most intuitive perception method. In contrast, force sensing can obtain the most direct contact information when contacting the target and the environment.
[0003] Traditional vision technologies usually process images through binarization or gradient to find target features. However, during the contact process with the target, the image may be blocked, resulting in the lack of feedback on the target position and state at the key steps of the operation. On the other hand, force sensing signals do not generate effective signals before contacting the target, so there is a lack of target search and localization before contact.
[0004] The disadvantage is that although vision and force sensing technologies are used in robot control, the current accuracy of robot control is relatively low. Summary of the Invention
[0005] The control method, control system, and computer-readable storage medium for a robot provided by this application can improve the control accuracy of the robot.
[0006] In a first aspect, this application provides a control method for a robot. The method includes: encoding a target image stream using a vision encoding module to obtain an image encoding stream; and receiving a force signal stream sent by a force sensing component; wherein, the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream; increasing the image encodings in the image encoding stream, and after increasing the image encodings in the image encoding stream, aligning the image encoding stream with the force signal stream to obtain a fused data stream; inputting the target data stream into a prediction model to obtain a correction parameter output by the prediction model; the correction parameter is used to adjust the robot at a future moment; wherein, the target data stream includes the current force signal and the historical data in the fused data stream; controlling the robot according to the correction parameter.
[0007] Wherein, increasing the image encodings in the image encoding stream includes: performing an interpolation operation on the current image encoding stream in the image encoding stream according to the acquisition frequency corresponding to the force signal stream to increase the image encodings in the image encoding stream; and / or, using the prediction model to predict the current image encoding stream in the image encoding stream to obtain interpolation image encodings between adjacent image encodings in the current image encoding stream, and increasing the image encodings in the image encoding stream.
[0008] Among them, the prediction model includes an encoding layer, a Transformer layer, and a fully connected layer; inputting the target data stream into the prediction model to obtain the correction parameters output by the prediction model, including: globally encoding and position encoding the target data stream using the encoding layer to obtain the target encoding; extracting features from the target encoding using the Transformer layer to obtain the target features; predicting the correction parameters from the target features using the fully connected layer to obtain the correction parameters.
[0009] Among them, the Transformer layer includes a layer normalization layer, a multi-head self-attention mechanism layer, and a fully connected feed-forward network. Extracting features from the target encoding using the Transformer layer to obtain the target features includes: performing layer normalization on the target encoding using the layer normalization layer to obtain the normalized features; performing self-attention learning on the normalized features using the multi-head self-attention mechanism layer to obtain the self-attention features; processing the self-attention features using the fully connected feed-forward network to obtain the target features.
[0010] Among them, the method further includes: inputting the target data stream into the prediction model to obtain the predicted image at a future moment output by the prediction model.
[0011] Among them, encoding the target image stream using the visual encoding module to obtain the image encoding stream includes: performing binary segmentation on each image in the target image stream using the visual encoding module to obtain binary images; sampling the targets in the binary images to obtain the corresponding image encodings, and further obtaining the image encoding stream.
[0012] Among them, the model in the visual encoding module uses a weighted cross-entropy loss function for parameter optimization during training.
[0013] Among them, the prediction model uses a mean square error loss function for parameter optimization during training.
[0014] In a second aspect, the present application provides a control system for a robot. The control system for the robot includes: an image acquisition component for acquiring a target image stream; a force sensing component for acquiring a force signal stream; a processing component connected to the image acquisition component, the force sensing component, and the robot, and configured to implement the control method provided in the first aspect.
[0015] In a third aspect, the present application provides a computer-readable storage medium for storing a computer program, which when executed by a processor, is configured to implement the control method provided in the first aspect.
[0016] The beneficial effects of the present application are as follows: Different from the prior art, the control method, control system, and computer-readable storage medium of the robot provided in the present application consider the problem that the acquisition frequency of the force signal stream is greater than that of the target image stream. By increasing the image encoding in the image encoding stream, after increasing the image encoding in the image encoding stream, aligning the image encoding stream with the force signal stream, and obtaining the fused data stream, it ensures the precise coordination of the two perception information (image and force sense) at the same time point, eliminates the error caused by signal misalignment, and uses the high-frequency force signal to drive the fusion signal output of the image encoding and the force signal, enabling the system to respond quickly in a high-dynamic scene; then inputting the target data stream into the prediction model to obtain the correction parameter output by the prediction model; the correction parameter is used to adjust the robot at a future moment; where the target data stream includes the current force signal and the historical data in the fused data stream; controlling the robot according to the correction parameter, thereby improving the control accuracy of the robot, enhancing the system stability, and effectively reducing the mistakes and interruptions during the operation of the robot.
[0017] Further, the control method, control system, and computer-readable storage medium of the robot in the present application can realize high-precision and high-efficiency operations in complex dynamic scenes by combining the rich information of vision with the real-time feedback of force sense, and also significantly improve the stability and success rate in static operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0019] Figure 1 is a schematic flowchart of an embodiment of the control method of the robot provided by the present application;
[0020] Figure 2 is Figure 1 a schematic flowchart of an embodiment of step 11 in
[0021] Figure 3 is Figure 1 a schematic flowchart of an embodiment of step 14 in
[0022] Figure 4 is Figure 3 a schematic flowchart of an embodiment of step 142 in
[0023] Figure 5 is a schematic structural diagram of the control system of the robot provided by the present application;
[0024] Figure 6 It is a schematic diagram of the data stream and memory pool distribution at the current moment of current provided by this application;
[0025] Figure 7 It is a schematic structural diagram of an embodiment of the control system of the robot provided by this application;
[0026] Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. It can be understood that the specific embodiments described herein are only used to explain this application, rather than limiting this application. In addition, it should be noted that for the convenience of description, only parts related to this application rather than all structures are shown in the drawings. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0028] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in conjunction with the embodiments can be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0029] Vision and force sense are two of the most commonly used perception methods by humans. Vision contains rich information, especially in terms of target positioning, including space, geometric morphology, color, texture, etc., and is the most intuitive perception method. In contrast, force sense can obtain the most direct contact information when contacting the target and the environment.
[0030] Traditional vision technologies usually process images through binarization or gradient to find target features. However, during the contact process with the target, the image may be blocked, resulting in the lack of feedback on the target position and state in the key steps of the operation. On the other hand, force sense signals do not generate effective signals before contacting the target, so there is a lack of target search and positioning before contact.
[0031] The disadvantage is that although vision and force sense technologies are used in robot control, the current accuracy of robot control is relatively low.
[0032] Based on this, in view of the problem that the acquisition frequency corresponding to the force signal stream is greater than that of the target image stream, by increasing the image encoding in the image encoding stream, after increasing the image encoding in the image encoding stream, aligning the image encoding stream with the force signal stream to obtain a fused data stream, it ensures the precise coordination of the two perception information (image and force sense) at the same time point, eliminates the error caused by signal misalignment, and uses the high-frequency force signal to drive the fusion signal output of the image encoding and the force signal, enabling the system to respond quickly in a high-dynamic scenario; then inputting the target data stream into the prediction model to obtain the correction parameters output by the prediction model; the correction parameters are used to adjust the robot at a future moment; where the target data stream includes the current force signal and the historical data in the fused data stream; controlling the robot according to the correction parameters to improve the control accuracy of the robot, enhance the system stability, and effectively reduce the mistakes and interruptions during the robot operation. Specifically, refer to the technical solutions of any one of the following embodiments.
[0033] Refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the control method for a robot provided by this application. The method includes:
[0034] Step 11: Encode the target image stream using a visual encoding module to obtain an image encoding stream.
[0035] In some embodiments, an image acquisition component can be used to acquire the target image stream of the working area.
[0036] In some embodiments, the image acquisition component can be disposed on the execution mechanism of the robot for closely acquiring the target image stream of the working area. In some embodiments, the robot may include a corresponding robotic arm, and the corresponding work is completed by controlling the robotic arm. At this time, the image acquisition component can be disposed on the robotic arm.
[0037] In some embodiments, the image acquisition component can be disposed at other positions (non-robot positions) for acquiring the target image stream of the entire working area.
[0038] In some embodiments, the working area can be determined according to the actual situation. For example, if the robot is used for a welding task, the working area is the welding area, and there are workpieces to be welded in the welding area.
[0039] In some embodiments, traditional vision algorithms can be used to encode the image; it can also be extended to encode other custom parameters, such as the number of trajectories, the trajectory deflection angle, etc.
[0040] In some embodiments, refer to Figure 2 , step 11 may be the following process:
[0041] Step 111: Use the visual coding module to perform binary segmentation on each image in the target image stream to obtain a binary image.
[0042] In some embodiments, the target area in the image can be used as the foreground, and other areas as the background for binary segmentation to obtain a binary image. For example, the target area is represented by white, and other areas are represented by black.
[0043] Step 112: Sample the targets in the binary image to obtain corresponding image encodings, and then obtain an image encoding stream.
[0044] In some embodiments, taking the weld seam as an example of the target for illustration:
[0045] After denoising the binary image, sample the trajectory points from it. For example, along the y-axis direction, every 5 pixel rows, take the average value of the pixel x coordinates in the pixel row as the sampling point for that row, and obtain d / 5 sampling points. Use spline interpolation to fit the sampling points to obtain a trajectory line, and let the lowest point of the trajectory line be (x target , y target ), then obtain the image encoding x encode .
[0046] Step 12: Receive the force signal stream sent by the force sensing component.
[0047] In some embodiments, the force sensing component can be set on the execution mechanism of the robot. When the execution mechanism of the robot contacts the target in the image, the force sensing component will generate a corresponding force signal to represent the magnitude of the contact force when the execution mechanism contacts the target in the image.
[0048] In some embodiments, the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream. That is, within the same time length, the number of force signals in the force signal stream is greater than the number of images in the target image stream.
[0049] In some embodiments, the force sensing components can be distributed on different axes of the robot, that is, force signals in different dimensions can be acquired. For example, if the robot is a three-axis robot, three-dimensional force signals can be acquired. For example, if the robot is a four-axis robot, four-dimensional force signals can be acquired. For example, if the robot is a five-axis robot, five-dimensional force signals can be acquired. For example, if the robot is a six-axis robot, six-dimensional force signals can be acquired.
[0050] Step 13: Increase the image encodings in the image encoding stream. After increasing the image encodings in the image encoding stream, align the image encoding stream with the force signal stream to obtain a fused data stream.
[0051] In some embodiments, the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream. Therefore, the image encoding in the image encoding stream and the force signal in the force signal stream cannot be aligned in time. Thus, the image encoding in the image encoding stream is increased. After increasing the image encoding in the image encoding stream, the image encoding stream and the force signal stream are aligned to obtain a fused data stream.
[0052] In some embodiments, interpolation operations can be performed on the current image encoding stream in the image encoding stream according to the acquisition frequency corresponding to the force signal stream to increase the image encoding in the image encoding stream.
[0053] In some embodiments, a prediction model can be used to predict the current image encoding stream in the image encoding stream to obtain interpolated image encodings between adjacent image encodings in the current image encoding stream, and the image encoding in the image encoding stream is increased.
[0054] In some embodiments, interpolation operations can be performed on the current image encoding stream in the image encoding stream according to the acquisition frequency corresponding to the force signal stream to increase the image encoding in the image encoding stream, and a prediction model can be used to predict the current image encoding stream in the image encoding stream to obtain interpolated image encodings between adjacent image encodings in the current image encoding stream, and the image encoding in the image encoding stream is increased. That is, two methods are used to increase the image encoding. Among them, when there are two image encodings at the same time, the one with better quality is selected and retained according to the quality of the image encoding.
[0055] Step 14: Input the target data stream into the prediction model to obtain the correction parameters output by the prediction model; the correction parameters are used to adjust the robot at a future moment; wherein, the target data stream includes the current force signal and the historical data in the fused data stream.
[0056] In some embodiments, the prediction model can be constructed based on at least one network model such as Transformer, LSTM, RNN, iTransformer, etc.
[0057] In some embodiments, during the movement and operation of the robot, an offset phenomenon will occur. That is, there is a difference between the expected position and the actual position. Therefore, through the above method, the correction parameter prediction can be combined with the force signal and the image, and the correction parameters are used to adjust the robot at a future moment to improve the control accuracy of the robot.
[0058] In some embodiments, since the frequency in the target data stream is consistent with the acquisition frequency of the force signal, the frequency of the correction parameters output by the prediction model is also consistent with the acquisition frequency of the force signal. Furthermore, the correction parameters with a higher frequency can be used to adjust the robot at a future moment, increasing the number of correction adjustments per unit time for the robot, and thus improving the control accuracy of the robot.
[0059] In some embodiments, the prediction model includes an encoding layer, a Transformer layer, and a fully connected layer. Refer to Figure 3 , step 14 may be the following process:
[0060] Step 141: Use the encoding layer to perform global encoding and positional encoding on the target data stream to obtain the target encoding.
[0061] In some embodiments, the encoding layer is divided into two encoding parts. The first encoding part enables the Transformer layer to perceive the order of the data by encoding the global number (global encoding), and the second encoding part explicitly affects the model prediction result by encoding the distance from the current time to the next visual signal (positional encoding).
[0062] Among them, the global number encoding formula is:
[0063]
[0064] Among them, pos represents the label of the data, i represents the dimension index, and d model represents the dimension of the prediction model.
[0065] Among them, the distance encoding encodes the distance through a multi-layer perceptron, such as the following formula:
[0066] PE distance = MLP(n%D).
[0067] Among them, n is the current moment, and D represents the ratio between the image encoding frequency and the force signal acquisition frequency.
[0068] Generally speaking, the positional encoding can be expressed as:
[0069] PE(x) = x + PE number (x) + PE distance (x).
[0070] Step 142: Use the Transformer layer to extract features from the target encoding to obtain the target features.
[0071] In some embodiments, the number of Transformer layers can be multiple. Multiple Transformer layers are hierarchically connected. That is, the first Transformer layer is connected to the encoding layer, and then the remaining Transformer layers are sequentially connected after the first Transformer layer, and the fully connected layer is connected after the last Transformer layer. The structures of the Transformer layers are the same.
[0072] In some embodiments, the Transformer layer includes a layer normalization layer, a multi-head self-attention mechanism layer, and a fully connected feed-forward network. Refer to Figure 4 , step 142 may be the following process:
[0073] Step 31: Use the layer normalization layer to perform layer normalization on the target encoding to obtain normalized features.
[0074] Among them, the formula for layer normalization is:
[0075]
[0076] Among them, μ is the mean, σ is the standard deviation, is a very small constant used to prevent division by zero, and γ and β are learnable parameters.
[0077] Step 32: Use the multi-head self-attention mechanism layer to perform self-attention learning on the normalized features to obtain self-attention features.
[0078] Specifically, the multi-head self-attention mechanism can be expressed as:
[0079] MultiHead(Q, K, V) = Concat(head1, head2, …, head h )W O .
[0080] Among them, the calculation method of each head is:
[0081] head i = Attention(QW Qi , KW Ki , VW Vi ).
[0082] The self-attention mechanism is defined as follows:
[0083]
[0084] Here, the query (Query), key (Key), and value (Value) vectors are calculated by projecting the same input matrix X:
[0085] Q = XW Q , K = XW K , V = XW V .
[0086] Among them, is a learnable weight matrix, d k = d v = d x . In the multi-head self-attention mechanism, W OIt is a learnable weight matrix for linearly transforming the outputs of multiple heads.
[0087] Step 33: Process the self-attention features using a fully-connected feed-forward network to obtain the target features.
[0088] In addition, the Transformer layer also contains a fully-connected feed-forward network. The feed-forward network at each position is the same and is executed independently. The feed-forward network contains two linear transformations and a GeLU activation function:
[0089] FFN(x) = GeLU(xW1 + b1)W2 + b2.
[0090] Among them, W1 and W2 are weight matrices, and b1 and b2 are bias terms. Through layer normalization and residual connections, the Transformer can train deep networks more stably and efficiently. The formula for layer normalization is:
[0091]
[0092] Among them, μ is the mean, σ is the standard deviation, is a very small constant used to prevent division by zero, and γ and β are learnable parameters.
[0093] The formula for the residual connection is:
[0094] Output = x + SubLayer(x).
[0095] Here, SubLayer(x) can be the output of the self-attention layer or the feed-forward network layer. By introducing the residual connection, the problem of gradient disappearance in the training of deep networks can be effectively alleviated, and the training efficiency and stability of the model can be improved.
[0096] Step 143: Use a fully-connected layer to predict the correction parameters for the target features to obtain the correction parameters.
[0097] In some embodiments, the target data stream is input into the prediction model to obtain the predicted image at a future time output by the prediction model. The predicted image at a future time can be used for feed-forward control, that is, the robot can perceive the information in front of the target in advance through the predicted image at a future time. In an application scenario, such perception includes the front path and the position to be assembled soon, etc., providing pre-feedback for the robot in trajectory following and making the operation smoother and more accurate.
[0098] Step 15: Control the robot according to the correction parameters.
[0099] In some embodiments, the model in the above visual encoding module uses a weighted cross-entropy loss function to optimize the parameters during training.
[0100] In some embodiments, the above prediction model uses the mean squared error loss function to optimize parameters during training.
[0101] In this embodiment, considering the problem that the acquisition frequency of the force signal stream is greater than that of the target image stream, by increasing the image encoding in the image encoding stream, after increasing the image encoding in the image encoding stream, aligning the image encoding stream with the force signal stream, and obtaining the fused data stream, it ensures the precise coordination of the two perception information (image and force sense) at the same time point, eliminates the error caused by signal misalignment, and uses the high-frequency force signal to drive the fusion signal output of the image encoding and the force signal, enabling the system to respond quickly in a high-dynamic scenario; then inputting the target data stream into the prediction model to obtain the correction parameters output by the prediction model; the correction parameters are used to adjust the robot at a future moment; where the target data stream includes the current force signal and the historical data in the fused data stream; controlling the robot according to the correction parameters, thereby improving the control accuracy of the robot, enhancing the system stability, and effectively reducing the mistakes and interruptions during the robot operation.
[0102] In an application scenario, combined with Figure 5 for illustration:
[0103] The control system of the entire robot includes a visual coding module, a temporal memory pool, and a temporal multi-modal fusion module. During the actual operation process, the control system of the robot receives image and force data in real time and processes them separately. When the image data stream (target image stream) arrives, the robot encodes the image through the visual coding module and stores it in the temporal data memory pool according to the time stamp. When the six-dimensional force data stream (force signal stream) arrives, the robot retrieves some historical data from the temporal data memory pool and sends them together to the temporal multi-modal fusion module to calculate and output high-frequency correction parameters.
[0104] Among them, the role of the visual coding module is to preprocess the image, perform image encoding through a deep learning model, and finally store the encoding result in the temporal data memory pool. The specific process of the visual coding module includes image cropping and image encoding.
[0105] The process of image cropping is as follows: Crop the working area at a fixed position from the image. The size of the image is [h, w], and the pixel coordinates of the robot tool end in the image are (x p , y p ). Given the cropping size of d, then the cropped image region [max(0, x - d / 2):min(h - 1, x + d / 2), y:min(w - 1, y + d)] obtains a target region with a size of [d, d], and the coordinates of the machine arm tool end pixel are updated to (x' p , y' p(Fixed at the midpoint of the upper boundary of the image (d / 2, 0)).
[0106] The image encoding process is as follows: The cropped image is input into a deep learning model for image encoding. In trajectory tracking, the image is encoded as the shortest distance from the pixel at the tool tip to the trajectory, i.e., x encode = x' p - x target , in other tasks, the image will be encoded as an interpolable custom parameter. Taking the workpiece welding image as an example, first, the image is input into the UNet deep learning network, and the image is binary segmented with the weld seam as the foreground point. The output result is a binary image with the weld seam area in white and other areas in black. After simple denoising processing of the model output image, trajectory points are sampled from it. Specifically, along the y-axis direction, every 5 pixel rows, the average value of the pixel x coordinates in the pixel row is taken as the sampling point of that row, and d / 5 sampling points are obtained. Spline interpolation is used to fit the sampling points to obtain a trajectory line. Let the lowest point of the trajectory line be (x target , y target ), then the image encoding x encode is obtained.
[0107] The loss function of the deep learning model can be in the following way:
[0108] The deep learning model is a segmentation network that adopts the UNet architecture and selects ResNet as the backbone network. The initial parameters of the model are selected as the model parameters pre-trained on ImageNet, and the model is fine-tuned on the local dataset. Specifically, the goal of the model is to predict the weld seam area as the foreground point and other areas as the background point, and a weighted cross-entropy loss function is used for parameter optimization during training. The weighted cross-entropy loss function can be defined as:
[0109] where N is the number of image samples, y e is the true label of the e-th pixel, taking values of 0 or 1, p e is the probability that the i-th pixel is predicted as 1, and α and β are the weight coefficients of the foreground and background respectively, used to balance the ratio of foreground pixels and background pixels. For example, α = 3, β = 1, d×d represents the area size of the region where the pixel is located.
[0110] The specific process of the time-series data memory pool is as follows:
[0111] The time-series data memory pool is responsible for integrating the image encoding data output at different frequencies, the six-axis force signal, and the correction parameters output by the time-series model (prediction model) in the time-series multi-modal fusion module. Among them, the image encoding is V, the six-axis force signal is F, and the correction parameter output by the time-series model is S. Assuming that the current time is t = n and the magnification of the image encoding output frequency to the time-series model output frequency is D, the data distribution in the time-series memory pool is: V = {v0, v′1, …, v′ n}, where v′ indicates that the data is obtained by time-series model prediction and interpolation, that is, in the time-series data memory pool, the image encoding will be interpolated into high-frequency parameters; F = {f0, f1, …, f n}, where each force signal is actual sensor data; S = {0, s1, s2, …, s′ n-1 , s′ n}, where s′ n-1 , s′ n indicates that the data is the output of the time-series model at the previous moment, and other data are historical data. The formula shows that the model predicted two correction parameters at the previous moment.
[0112] As Figure 6 shown, a schematic diagram of the data flow and memory pool distribution at the current time of current is given. The data in the memory pool is divided into three categories (three time axes) according to the format, namely the six-axis force signal flow, the image encoding flow, and the time-series model output flow; it is divided into three categories (three colors) according to the form, namely the actual acquisition (sensor or visual model output) data, the pseudo data (obtained by time-series model output or interpolation), and the time-series model output.
[0113] The specific process of the time-series multi-modal fusion module is as follows:
[0114] The time-series multi-modal fusion module extracts all the data within the given time window W from the time-series data memory pool as input, outputs high-frequency correction parameters for the control module to use; outputs the pseudo representation of the camera signal that has not arrived to mitigate the impact of image transmission and model response time on model inference. This module is constructed using Transformer, and the specific process is as follows.
[0115] First is the model input. The time-series multi-modal fusion module extracts the data within the time window W at the current time t = n from the time-series data memory pool as input, including [win(V), win(F), win(S)], then the continuous data stream within the W time window can be obtained, denoted as C n = {c n-W+1 , c n-W+2 , …, c n}, that is, the input of the model at the moment t = n is C n .
[0116] Then there is the temporal encoding, which is used to inject temporal information into the data so that the Transformer can perceive the order of the data and the prediction relationship. The temporal encoding in this paper is divided into two parts. The first part enables the Transformer to perceive the order of the data by encoding the global number, and the second part explicitly affects the model prediction result by encoding the distance from the current time to the next visual signal. The global number encoding formula is:
[0117] where pos represents the label of the data, i represents the dimension index, and d model represents the dimension of the model. The distance encoding encodes the distance through a multi-layer perceptron:
[0118] PE distance = MLP(n%D). Where n is the current moment, and D represents the ratio of the image encoding frequency to the six-axis force acquisition frequency. Generally speaking, the position encoding can be expressed as:
[0119] PE(x) = x + PE number (x) + PE distance (x).
[0120] Then, in the Transformer layer, the relationship features and temporal features in the data are learned through the self-attention mechanism and the feed-forward network, and finally high-frequency correction parameters are output. The self-attention encoder structure of the Transformer is stacked by multiple Transformer layers, and each Transformer layer contains layer normalization, the multi-head self-attention mechanism layer, and the fully connected feed-forward network.
[0121] In addition, the Transformer layer also contains a fully connected feed-forward network. The feed-forward network at each position is the same and is executed independently. The feed-forward network contains two linear transformations and a GeLU activation function:
[0122] FFN(x) = GeLU(xW1 + b1)W2 + b2. Where W1 and W2 are weight matrices, and b1 and b2 are bias terms. Through layer normalization and residual connection, the Transformer can train deep networks more stably and efficiently. The formula for layer normalization is:
[0123] where μ is the mean, σ is the standard deviation, is a very small constant used to prevent division by zero, and γ and β are learnable parameters.
[0124] The formula for the residual connection is:
[0125] Output = x + SubLayer(x). Here, SubLayer(x) can be the output of a self-attention layer or a feed-forward network layer. By introducing residual connections, the problem of gradient vanishing in the training of deep networks can be effectively alleviated, and the training efficiency and stability of the model can be improved.
[0126] The overall process of the temporal multi-modal fusion module is as follows:
[0127] First, the input data C at the current time point is obtained n , and then data dimensionality increase and temporal encoding are performed through PE(C N ). Feature extraction is performed by stacking multiple Transformer layers to obtain O n , which is the feature output of the Transformer model. Then, the last feature vector of O n is taken and input into the fully connected layer for deviation correction parameter prediction and visual signal pseudo-data prediction. These pseudo-data can be added as image encoding in the image encoding stream.
[0128] The loss function of the temporal multi-modal fusion module is as follows:
[0129] The goal of the temporal model of the temporal multi-modal fusion module is to predict the next time step and the image signal pseudo-data at all times when the image signals have not arrived. To explore the value of future trajectories, the prediction task of the temporal model is further extended: predicting the unarrived visual signals (images) within the future W2 window, and using the data in the training set for supervised learning to improve the prediction ability of the model. For all predictions of the model, the mean squared error loss (MSE) is used to optimize the model:
[0130] where is the predicted value of the model, y i is the value in the image label, N is the number of samples, and u is the number of all parameters that the model needs to predict, such as u = 2 + W2.
[0131] Refer to Figure 7 , Figure 7 is a schematic structural diagram of an embodiment of the control system of the robot provided by this application. The control system 100 of the robot includes: an image acquisition component 10, a force sensing component 20, and a processing component 30.
[0132] The image acquisition component 10 is used to acquire the target image stream.
[0133] The force sensing component 20 is used to acquire the force signal stream.
[0134] The processing component 30 is connected to the image acquisition component 10, the force sensing component 20 and the robot, and is used to encode the target image stream by using the visual coding module to obtain an image coding stream; and receive the force signal stream sent by the force sensing component 20; wherein, the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream; increase the image coding in the image coding stream, and after increasing the image coding in the image coding stream, align the image coding stream with the force signal stream to obtain a fused data stream; input the target data stream into the prediction model to obtain the correction parameter output by the prediction model; the correction parameter is used to adjust the robot at a future moment; wherein, the target data stream includes the current force signal and the historical data in the fused data stream; control the robot according to the correction parameter.
[0135] In some embodiments, the processing component 30 is further configured to perform an interpolation operation on the current image coding stream in the image coding stream according to the acquisition frequency corresponding to the force signal stream to increase the image coding in the image coding stream; and / or, use the prediction model to predict the current image coding stream in the image coding stream to obtain the interpolation image coding between adjacent image codings in the current image coding stream, and increase the image coding in the image coding stream.
[0136] In some embodiments, the prediction model includes an encoding layer, a Transformer layer and a fully connected layer; the processing component 30 is further configured to perform global encoding and position encoding on the target data stream by using the encoding layer to obtain a target encoding; perform feature extraction on the target encoding by using the Transformer layer to obtain a target feature; perform correction parameter prediction on the target feature by using the fully connected layer to obtain a correction parameter.
[0137] In some embodiments, the Transformer layer includes a layer normalization layer, a multi-head self-attention mechanism layer, and a fully connected feed-forward network. The processing component 30 is further configured to perform layer normalization on the target encoding by using the layer normalization layer to obtain a normalized feature; perform self-attention learning on the normalized feature by using the multi-head self-attention mechanism layer to obtain a self-attention feature; perform processing on the self-attention feature by using the fully connected feed-forward network to obtain a target feature.
[0138] In some embodiments, the processing component 30 is further configured to input the target data stream into the prediction model to obtain the predicted image at a future moment output by the prediction model.
[0139] In some embodiments, the processing component 30 is further configured to perform binary segmentation on each image in the target image stream by using the visual coding module to obtain a binary image; sample the target in the binary image to obtain the corresponding image coding, and further obtain the image coding stream.
[0140] In some embodiments, the processing component 30 is further configured to optimize parameters using a weighted cross-entropy loss function during the model training process in the visual encoding module.
[0141] In some embodiments, the processing component 30 is further configured to optimize parameters using a mean squared error loss function during the prediction model training process.
[0142] In some embodiments, the processing component 30 is further configured to implement the method of any of the above embodiments.
[0143] Refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 80 is used to store a computer program 81, and when the computer program 81 is executed by a processor, it is used to implement the following method:
[0144] Encoding a target image stream using a visual encoding module to obtain an image encoding stream; and receiving a force signal stream sent by a force sensing component; wherein, the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream; increasing the image encoding in the image encoding stream, and after increasing the image encoding in the image encoding stream, aligning the image encoding stream with the force signal stream to obtain a fused data stream; inputting the target data stream into a prediction model to obtain a correction parameter output by the prediction model; the correction parameter is used to adjust the robot at a future moment; wherein, the target data stream includes the current force signal and historical data in the fused data stream; controlling the robot according to the correction parameter.
[0145] In some embodiments, when the computer program 81 is executed by a processor, it is further configured to implement the method of any of the above embodiments.
[0146] In summary, for the control method, control system, and computer-readable storage medium of the robot provided by the present application, considering the problem that the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream, by increasing the image encoding in the image encoding stream, and after increasing the image encoding in the image encoding stream, aligning the image encoding stream with the force signal stream to obtain a fused data stream, it ensures the precise coordination of the two perception information (image and force sense) at the same time point, eliminates the error caused by signal misalignment, and uses the high-frequency force signal to drive the fusion signal output of the image encoding and the force signal, enabling the system to respond quickly in a high-dynamic scenario; then inputting the target data stream into the prediction model to obtain a correction parameter output by the prediction model; the correction parameter is used to adjust the robot at a future moment; wherein, the target data stream includes the current force signal and historical data in the fused data stream; controlling the robot according to the correction parameter, thereby improving the control accuracy of the robot, enhancing the system stability, and effectively reducing the mistakes and interruptions during the operation of the robot.
[0147] Furthermore, this application aims to achieve a deep integration of vision (images) and force sensing (force signals) to ensure stable, high-precision, and high-speed output. To achieve this goal, the key technical points of this application are as follows:
[0148] The timestamps of the vision and force sensing signals are highly aligned. To avoid signal misalignment affecting the output accuracy, the present invention designs a mechanism to achieve a high degree of alignment between the timestamps of the vision signal and the force sensing signal. This ensures the precise coordination of the two signals at the same time point, improving the accuracy of the overall operation.
[0149] Output of high-frame-rate fusion signals: The present invention drives the output of the vision and force sensing fusion signals through high-frame-rate force sensing signals. Since vision processing consumes a large amount of computing resources, using high-frame-rate force sensing signals as the driving force can effectively improve the output speed of the fusion signals and meet the operation requirements in high-dynamic scenarios.
[0150] Pre-feedback mechanism: To achieve high stability and high precision, this application makes full use of the wide field of view of the vision signal to pre-perceive the information in front of the target. This perception includes the path ahead and the position to be assembled soon, etc., providing pre-feedback for the robot during trajectory following and making the operation smoother and more accurate.
[0151] Countermeasure strategy for the delay of vision processing signals: This application designs a strategy to deal with the delay of vision processing signals to solve the problem of insufficient data input caused by the delay. By combining the current force sensing signal and the previous vision and force fusion data for prediction, the system can still make high-precision real-time decisions even when the vision signal is delayed.
[0152] These key technical points work together to overcome the shortcomings of the prior art, achieve a deep integration of vision and force sensing, improve the accuracy, stability, and speed of the robot in industrial operations, and adapt to complex and changing dynamic and static operating environments.
[0153] In several embodiments provided by this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0154] If the integrated units in the above-mentioned other embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processing circuit component (processor) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0155] The above are only the embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A robot control method, characterized in that: The method comprises: Encoding the target image stream using a visual encoding module to obtain an image encoding stream; and receiving a force signal stream sent by a force sensing component; wherein the acquisition frequency corresponding to the force signal stream is greater than the acquisition frequency of the target image stream; adding image codes to the image coding stream, and after adding the image codes to the image coding stream, aligning the image coding stream with the force signal stream to obtain a fused data stream; Input the target data stream into the prediction model to obtain the correction parameters output by the prediction model; the correction parameters are used to adjust the robot at a future time; wherein the target data stream includes the current force signal and the historical data in the fused data stream; The robot is controlled according to the deviation correction parameters.
2. The control method according to claim 1, characterized in that: The step of adding the image code in the image coding stream comprises: According to the acquisition frequency corresponding to the force signal stream, an interpolation operation is performed on the current image coding stream in the image coding stream to increase the image code in the image coding stream; And / or, using the prediction model to predict the current image coding stream in the image coding stream, obtain interpolated image codes between adjacent image codes in the current image coding stream, and increase the image codes in the image coding stream.
3. The control method according to claim 1, characterized in that: The prediction model includes a coding layer, a Transformer layer and a fully connected layer; the target data stream is input into the prediction model to obtain the correction parameters output by the prediction model, including: Performing global encoding and position encoding on the target data stream using the encoding layer to obtain a target encoding; Using the Transformer layer to extract features from the target code to obtain target features; The fully connected layer is used to predict the correction parameters of the target features to obtain the correction parameters.
4. The control method according to claim 3, characterized in that: The Transformer layer includes a layer normalization layer, a multi-head self-attention mechanism layer, and a fully connected feedforward network. The Transformer layer is used to extract features of the target code to obtain target features, including: Using the layer normalization layer to perform layer normalization on the target code to obtain normalized features; Using the multi-head self-attention mechanism layer to perform self-attention learning on the normalized features to obtain self-attention features; The self-attention feature is processed using a fully connected feed-forward network to obtain the target feature.
5. The control method according to claim 1, characterized in that: The method further includes: inputting the target data stream into the prediction model to obtain a prediction image at a future moment output by the prediction model.
6. The control method according to claim 1, characterized in that: The method of encoding the target image stream by using the visual encoding module to obtain the image encoding stream includes: Using the visual encoding module to perform binary segmentation on each image in the target image stream to obtain a binary image; The target in the binary image is sampled to obtain the corresponding image code, and then the image code stream is obtained.
7. The control method according to claim 1, characterized in that: The model in the visual encoding module uses a weighted cross entropy loss function for parameter optimization during training.
8. The control method according to claim 1, characterized in that: The prediction model uses a mean square error loss function to optimize parameters during the training process.
9. A robot control system, characterized in that: The control system of the robot comprises: An image acquisition component, used for acquiring a target image stream; A force sensing component for collecting force signal streams; A processing component is connected to the image acquisition component, the force sensing component and the robot, and is used to implement the control method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, it is used to implement the control method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Grabbed object recognition method based on tactile vibration signal and visual image fusion
CN112388655A
Multi-modal data processing method applied to robot interaction
CN113894779A
Robot precision assembly control method and system based on visual and tactile fusion
CN113927602A
Underwater SLAM (Simultaneous Localization and Mapping) system fused with visual inertial pressure sensor and method thereof
CN117029809A
Data fusion method and system based on reservoir multi-source data
CN119475254A