Biped humanoid robot attitude detection method based on time sequence multi-scale feature fusion

Through the temporal multi-scale feature fusion and attention mechanism processing, combined with the jump connection module, the problem of low accuracy and recall of existing bipedal humanoid robot posture detection methods is solved, and higher detection accuracy and robustness are achieved.

CN120047929AActive Publication Date: 2025-05-27BEIJING SCI & TECH PATENT OFFICE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411858947.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-27
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The accuracy and recall rate of existing bipedal humanoid robot posture detection methods are low, mainly due to the simple structure of the model network, and the information in the picture cannot be fully mined.

Method used

The detection method based on time-series multi-scale feature fusion is adopted, and the original image of multiple frames is continuously acquired, feature maps of different scales are extracted, and feature fusion is carried out in time-series multi-scale feature fusion. Combining the attention mechanism and jump connection module, the robustness of feature expression and the feature decoder's feature decoding ability is enhanced.

Benefits of technology

The accuracy and recall of posture detection of bipedal humanoid robots is improved, the robustness of robot recognition for different positions and sizes is enhanced, and the calculation cost and inference speed are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047929A_ABST
    Figure CN120047929A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a biped humanoid robot attitude detection method based on time sequence multi-scale feature fusion, which comprises the following steps: acquiring an original image of a target biped humanoid robot, and extracting a plurality of feature maps with different scales; iteratively carrying out time sequence multi-scale feature fusion according to the feature map of the current scale of the current frame, the fusion feature map of the previous scale of the current frame and the feature map of the current scale of the previous frame to obtain a plurality of fusion feature maps of the current frame; processing the fusion feature map of the minimum scale and the fusion feature maps of other scales by using a preset attention module and a preset jump connection module according to the plurality of fusion feature maps of the current frame to obtain a coding feature and a plurality of jump connection features of the current frame; and according to the coding feature and the plurality of jump connection features, decoding to obtain an output feature map of the current frame. Through time sequence multi-scale feature fusion, the accuracy and recall rate of posture detection of the biped humanoid robot are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a biped humanoid robot pose detection method based on temporal multi-scale feature fusion. Background Art

[0002] Biped humanoid robot pose detection refers to identifying key points in images captured by a camera using various detection methods. Currently, there are generally two types of methods: The first type is to directly apply mature human pose key point detection methods to biped humanoid robot pose detection. The problem with this method is that there are significant differences in the morphology and size between humans and biped humanoid robots. Directly applying methods designed for humans to biped humanoid robot pose detection will result in low accuracy and low recall rate. The second type of method is to design a dedicated network structure based on existing publicly available biped humanoid robot key point datasets to achieve key point detection of biped humanoid robots. However, due to the relatively simple network structure used in this method, the accuracy is low.

[0003] The problem with current methods for biped humanoid robot pose detection is that the model network structure is relatively simple, and the information in biped humanoid robot images is not fully exploited, resulting in low detection accuracy and recall rate. Summary of the Invention

[0004] In view of this, the present invention provides a biped humanoid robot pose detection method based on temporal multi-scale feature fusion to solve the problem of low accuracy and recall rate of existing detection methods.

[0005] In a first aspect, the present invention provides a biped humanoid robot pose detection method based on temporal multi-scale feature fusion. The method includes:

[0006] Continuously obtain multiple frames of original images of a target biped humanoid robot, and respectively preprocess each original image and extract multiple feature maps of different scales;

[0007] Iteratively perform temporal multi-scale feature fusion based on the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale to obtain multiple fused feature maps of the current frame;

[0008] Sort the multiple fused feature maps of the current frame according to the scale size to obtain the fused feature map of the smallest scale and multiple fused feature maps of other scales;

[0009] Process the fused feature map of the smallest scale using a preset attention module to obtain the encoded feature of the current frame;

[0010] Process the fused feature maps of other scales using a preset skip connection module to obtain multiple skip connection features of the current frame;

[0011] Decode the output feature map of the current frame based on the coding features and multiple skip connection features.

[0012] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided by the present invention increases the robustness of identifying robots at different positions and different sizes through temporal multi-scale feature fusion, adds an attention mechanism between encoding and decoding to enhance the feature decoding ability of the decoder, and utilizes the decoder to enhance the ability to process spatial features, thereby improving the accuracy and recall rate of bipedal humanoid robot pose detection.

[0013] In an optional implementation manner, preprocess each original image and then extract multiple feature maps of different scales, including:

[0014] Normalize the original image to obtain a standard image suitable for input to the deep network;

[0015] Extract features of different scales from the standard image to obtain multiple feature maps of different scales.

[0016] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided by the present invention is convenient for subsequent processing by the deep network through normalization processing, and is beneficial to improving the accuracy of subsequent bipedal humanoid robot pose detection by extracting features of different scales and understanding multiple features of the original image at different levels.

[0017] In an optional implementation manner, iteratively perform temporal multi-scale feature fusion based on the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame to obtain multiple fused feature maps of the current frame, including:

[0018] Fuse the feature map of the smallest scale of the current frame and the feature map of the smallest scale of the previous frame to obtain the fused feature map of the smallest scale of the current frame;

[0019] Take the fused feature map of the smallest scale of the current frame as the fused feature map of the previous scale of the current frame in the first temporal multi-scale feature fusion process, and repeat the temporal multi-scale feature fusion according to the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame to obtain multiple fused feature maps of different scales of the current frame.

[0020] In an optional implementation manner, perform temporal multi-scale feature fusion according to the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame, including:

[0021] Perform a first preset convolution process on the feature map of the current frame at the current scale to obtain a first feature map, perform a preset upsampling process on the fused feature map of the previous scale of the current frame to obtain a second feature map, and perform a second preset convolution process on the feature map of the previous frame at the current scale to obtain a third feature map;

[0022] Add the first feature map and the second feature map and then splice them with the third feature map to obtain a spliced feature map;

[0023] Perform a depthwise separable convolution process on the spliced feature map to obtain the fused feature map of the current frame at the current scale.

[0024] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided by the present invention enhances the robustness of feature expression by fusing the same-scale features of adjacent temporal frames and different-scale features of the same temporal frame. The depthwise separable convolution process can significantly reduce the model parameters and computational volume after feature splicing in the channel dimension and improve the model processing speed.

[0025] In an alternative embodiment, use a preset skip connection module to process the fused feature maps of other scales respectively to obtain multiple skip connection features of the current frame, including:

[0026] Input the fused feature maps of other scales into the preset skip connection module respectively and perform a first preset convolution process to obtain the skip connection features of the corresponding scale of the current frame.

[0027] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided by the present invention enables the direct transmission of low-level features (such as edges, textures, etc.) in the network to higher-level feature maps through skip connections of the fused features of other scales, which is beneficial to maintaining spatial information and improving the accuracy of target detection.

[0028] In an alternative embodiment, perform decoding based on the encoded feature and multiple skip connection features to obtain the output feature map of the current frame, including:

[0029] Sort the multiple skip connection features from smallest to largest in scale to obtain a skip connection feature sequence;

[0030] Perform iterative preset asymmetric convolution processing based on the encoded feature and the skip connection feature sequence to obtain the output encoded feature of the current frame;

[0031] Perform a second preset convolution process on the output encoded feature of the current frame to obtain the output feature map of the current frame.

[0032] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided by the present invention can achieve more powerful learning ability without significantly increasing the number of parameters by using skip connections and feature addition. Since the amount of data to be processed is reduced, the computational cost can be reduced and the inference speed can be accelerated.

[0033] In an alternative embodiment, the method further includes:

[0034] Extracting multiple key points and skeletons of the target bipedal humanoid robot from the output feature map of the current frame;

[0035] Visualizing the output features according to the key points and skeletons to obtain the heat map of the target bipedal humanoid robot.

[0036] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided by the present invention facilitates the intuitive observation of the specific features captured by the network model at different levels by visualizing the output feature map, better understanding the process of the model extracting information from the input data and making predictions, so as to better optimize the model and improve the accuracy of the model in detecting the pose of the bipedal humanoid robot.

[0037] In a second aspect, the present invention provides a bipedal humanoid robot pose detection device, the device includes:

[0038] A multi-scale feature extraction module, configured to continuously obtain multiple frames of original images of the target bipedal humanoid robot, and respectively preprocess each original image and extract multiple feature maps of different scales;

[0039] A temporal multi-scale feature fusion module, configured to iteratively perform temporal multi-scale feature fusion according to the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame to obtain multiple fused feature maps of the current frame;

[0040] A scale sorting module, configured to sort the multiple fused feature maps of the current frame according to the scale size to obtain the fused feature map of the smallest scale and multiple other-scale fused feature maps;

[0041] An attention mechanism processing module, configured to process the fused feature map of the smallest scale by using a preset attention module to obtain the encoded feature of the current frame;

[0042] A skip connection processing module, configured to process the other-scale fused feature maps by using a preset skip connection module to obtain multiple skip connection features of the current frame;

[0043] A feature decoding module, configured to perform decoding according to the encoded feature and multiple skip connection features to obtain the output feature map of the current frame.

[0044] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any corresponding embodiment thereof.

[0045] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the method according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 FIG. is a flowchart of a method for biped humanoid robot pose detection based on temporal multi-scale feature fusion according to an embodiment of the present invention;

[0048] Figure 2 FIG. is a schematic diagram of the positions of key points and skeletons of a biped humanoid robot in a method for biped humanoid robot pose detection based on temporal multi-scale feature fusion according to an embodiment of the present invention;

[0049] Figure 3 FIG. is a schematic diagram of a key point heat map and a skeleton heat map of a biped humanoid robot in a method for biped humanoid robot pose detection based on temporal multi-scale feature fusion according to an embodiment of the present invention;

[0050] Figure 4 FIG. is a flowchart of another method for biped humanoid robot pose detection based on temporal multi-scale feature fusion according to an embodiment of the present invention;

[0051] Figure 5 FIG. is a schematic diagram of a neural network structure executed by a method for biped humanoid robot pose detection based on temporal multi-scale feature fusion according to an embodiment of the present invention;

[0052] Figure 6 FIG. is a structural block diagram of a device for biped humanoid robot pose detection based on temporal multi-scale feature fusion according to an embodiment of the present invention;

[0053] Figure 7 FIG. is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] The embodiments of the present invention provide a bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion, which improves the accuracy and recall rate of bipedal humanoid robot pose detection by performing temporal multi-scale feature fusion on the collected images.

[0056] According to the embodiments of the present invention, there is provided an embodiment of a bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0057] In this embodiment, a bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion is provided, which can be used in the above computer system. Figure 1 is a flowchart of a bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion according to the embodiments of the present invention, as Figure 1 shown, the process includes the following steps:

[0058] Step S101, continuously obtain multiple frames of original images of the target bipedal humanoid robot, and respectively preprocess each original image and then extract multiple feature maps of different scales.

[0059] Specifically, multiple color images of the target bipedal humanoid robot can be continuously collected through a camera as the original images. Since the original images are consecutive image frames, for each frame of the original image, it may contain one or more complete bipedal humanoid robots, one or more parts of the bipedal humanoid robots, or may not contain a bipedal humanoid robot. When the original image does not contain a bipedal humanoid robot, the responses of the key points and the skeleton feature map output by the bipedal humanoid robot pose detection method provided in this embodiment will be relatively small, lower than the threshold, so ultimately no robot is detected. For each original image, after preprocessing according to the preset processing method, feature extraction of different scales is performed to obtain multiple feature maps of different scales.

[0060] Step S102: Iteratively perform temporal multi-scale feature fusion based on the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale to obtain multiple fused feature maps of the current frame.

[0061] Specifically, the temporal multi-scale feature fusion is achieved by inputting the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale into a preset feature fusion module. If the current scale of the current frame is the smallest scale and there is no fused feature map at the previous scale, then directly perform temporal fusion on the feature map of the current frame at the current scale and the feature map of the previous frame at the current scale to obtain the fused feature map of the current frame at the current scale. If the current scale of the current frame is not the smallest scale, there is a fused feature map of the current frame at the previous scale for the feature map of the previous frame at the current scale, then use the preset feature fusion module to fuse the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale according to the preset rules. Finally, multiple fused feature maps of the current frame are obtained.

[0062] Step S103: Sort the multiple fused feature maps of the current frame according to the scale size to obtain the fused feature map of the smallest scale and multiple fused feature maps of other scales.

[0063] Specifically, the scale sizes of the multiple fused feature maps of the current frame can be 1 / 32, 1 / 16, 1 / 8, 1 / 4, 1 / 2. Sorting according to the scale size, the 1 / 32 fused feature map, 1 / 16 fused feature map, 1 / 8 fused feature map, 1 / 4 fused feature map, and 1 / 2 fused feature map of the current frame are obtained. Among them, the 1 / 32 fused feature map is the fused feature map of the smallest scale, only as an example, but not limited thereto.

[0064] Step S104: Process the fused feature map of the smallest scale using a preset attention module to obtain the encoded feature of the current frame.

[0065] Specifically, input the fused feature map of the smallest scale into the preset attention module for attention mechanism processing to obtain the first encoder feature of the current frame. The preset attention module is an existing channel-based attention module, and the attention mechanism processing process is a mature existing technology, which will not be elaborated here.

[0066] Step S105: Process the fused feature maps of other scales using a preset skip connection module to obtain multiple skip connection features of the current frame.

[0067] Specifically, while inputting the fused feature map of the smallest scale into the preset attention module, input the fused feature maps of other scales into the preset skip connection module to obtain the corresponding skip connection features respectively.

[0068] Step S106: Decode according to the encoded features and multiple skip connection features to obtain the output feature map of the current frame.

[0069] Specifically, for the linkage relationship between the key points and the skeleton of the bipedal humanoid robot, a decoder with an asymmetric convolution structure is designed to reduce the number of parameters, mitigate overfitting, enhance the ability of the decoder module to process spatial features, and effectively improve the accuracy of the final pose detection.

[0070] In the bipedal humanoid robot pose detection, an attention mechanism is added between the encoder module and the decoder module, so that the importance of each channel of the features finally output by the encoder for the decoding work of the decoder varies according to different images, improving the feature decoding ability of the decoder and facilitating the improvement of the accuracy of the final detection result.

[0071] The bipedal humanoid robot pose detection method provided in this embodiment, through temporal multi-scale feature fusion, increases the robustness of identifying robots at different positions and of different sizes. An attention mechanism is added between the encoding and decoding to improve the feature decoding ability of the decoder. The decoder is used to enhance the ability to process spatial features, and the accuracy and recall rate of the bipedal humanoid robot pose detection are improved.

[0072] In some alternative embodiments, the method further includes:

[0073] Extract multiple key points and skeletons of the target bipedal humanoid robot according to the output feature map of the current frame.

[0074] Specifically, if the size of the output feature map of the current frame is where H represents the height of the original image after preprocessing, W represents the length of the original image after preprocessing, p represents the number of key points of the bipedal humanoid robot, and s represents the number of skeletons of the bipedal humanoid robot.

[0075] Each channel of the output feature map corresponds to each key point or each segment of the skeleton of the robot. In each output feature map, the feature value at the image position corresponding to the key point or skeleton is relatively large, and the response at the remaining positions is relatively small. Therefore, the key points can be extracted by the method of non-maximum suppression, and different key points of the bipedal humanoid robot can be combined by post-processing the feature map of the skeleton. As Figure 2 shown in the schematic diagram of the positions of the key points and skeletons of the bipedal humanoid robot, where the red dots represent the key points, and the orange lines between two red dots represent the skeletons.

[0076] Visualize the output features according to the key points and skeletons to obtain the heat map of the target bipedal humanoid robot.

[0077] Specifically, each channel of the output feature map is visualized and displayed in the form of a heat map, as Figure 3 shown are the key point heat map and the skeleton heat map of the bipedal humanoid robot. It can be seen that the responses of the feature maps at the corresponding positions of the key points and the skeleton are relatively large, while the responses of the feature maps at other positions are relatively small.

[0078] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided in this embodiment facilitates intuitively observing the specific features captured by the network model at different levels by visualizing the output feature map, better understanding the process of the model extracting information from the input data and making predictions, so as to better optimize the model and improve the accuracy of the model in detecting the pose of the bipedal humanoid robot.

[0079] In this embodiment, a bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion is provided, which can be used in a computer system, Figure 4 is a flowchart of the bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion according to an embodiment of the present invention, as Figure 4 shown, this process includes the following steps:

[0080] Step S201, continuously obtain multiple frames of original images of the target bipedal humanoid robot, and respectively preprocess each original image and extract multiple feature maps of different scales.

[0081] Specifically, the above step S201 includes:

[0082] Step S2011, normalize the original image to obtain a standard image suitable for input to the deep network.

[0083] Specifically, the t-th frame of the original color image is I t , after normalizing I t , adjust the size to H×W, where H is the height and W is the length, to obtain a standard image I t ', and the height and length of the standard image should be greater than 200 pixels. It can be an image containing one or more complete bipedal humanoid robots, one or more parts of bipedal humanoid robots, or an image that does not contain bipedal humanoid robots.

[0084] Step S2012, perform feature extraction on the standard image at different scales to obtain multiple feature maps of different scales.

[0085] Specifically, any feature extraction module can be used to perform feature extraction on the standard image at different scales. The feature maps at smaller scales represent deeper semantic features, while the feature maps at larger scales represent simpler features (such as texture, shape, etc.). Since the positions and distances of the bipedal humanoid robot in the image vary, the main task of feature extraction is to extract a feature with rotational invariance, scale invariance, stability, and generality for pose estimation.

[0086] As Figure 5 shown in the schematic diagram of the neural network structure relied on during the execution of this embodiment, where the feature extraction module may include multiple feature extractors at different scales. In this embodiment, five feature extractors are taken as an example for illustration, only as an example, but not limited thereto. The standard input image I t ′ obtained after preprocessing the original image through normalization / size adjustment is input into the first feature extractor to obtain the first feature map t The scale (size) of the first feature map is where C1 is the number of feature channels of the first feature map; the first feature map is input into the second feature extractor to obtain the second feature map The scale of the second feature map is where C2 is the number of feature channels of the second feature map; the second feature map is input into the third feature extractor to obtain the third feature map The scale of the third feature map is where C3 is the number of feature channels of the third feature map; the third feature map is input into the fourth feature extractor to obtain the fourth feature map The scale of the fourth feature map is where C4 is the number of feature channels of the fourth feature map; the fourth feature map is input into the fifth feature extractor to obtain the fifth feature map The scale of the fifth feature map is where C5 is the number of feature channels of the fifth feature map. In summary, the feature extraction module finally outputs the first feature map corresponding to the t-th frame, the second feature map the third feature map the fourth feature map the fifth feature map These five feature maps at different scales.

[0087] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided in this embodiment is convenient for subsequent deep network processing through normalization processing, and by extracting features at different scales, it can understand multiple features of the original image at different levels, which is beneficial to improving the accuracy of subsequent bipedal humanoid robot pose detection.

[0088] Step S202: Iteratively perform temporal multi-scale feature fusion based on the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale to obtain multiple fused feature maps of the current frame.

[0089] Specifically, the above-mentioned step S202 includes:

[0090] Step S2021: Perform feature fusion on the feature map of the smallest scale of the current frame and the feature map of the smallest scale of the previous frame to obtain the fused feature map of the smallest scale of the current frame.

[0091] Specifically, the main component structure of the preset feature fusion module is FLNet, whose function is to fuse the k-th feature map of the t-th frame, the (k - 1)-th feature map of the t-th frame (if it exists), and the k-th feature map of the (t - 1)-th frame, and fuse the same-scale features of adjacent temporal frames and different-scale features of the same temporal frame to enhance the robustness of feature expression. For the feature map of the smallest scale of the current frame, since there is no feature map of the previous scale of the current frame, it is only necessary to perform feature fusion on the feature map of the smallest scale of the current frame and the feature map of the smallest scale of the previous frame to obtain the fused feature map of the smallest scale of the current frame. The smallest scale of the current frame is equal to the smallest scale of the previous frame.

[0092] Step S2022: Use the fused feature map of the smallest scale of the current frame as the fused feature map of the previous scale of the current frame in the first temporal multi-scale feature fusion process, and repeat the temporal multi-scale feature fusion according to the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale to obtain multiple fused feature maps of different scales of the current frame.

[0093] Specifically, taking feature maps of five different scales as an example, the fused feature map of the smallest scale of the t-th frame is used as the fifth fused feature map Fuse the fourth feature map of the t-th frame, the fifth fused feature map, and the fourth feature map of the (t - 1)-th frame to obtain the fourth fused feature map Fuse the third feature map of the t-th frame, the fourth fused feature map, and the third feature map of the (t - 1)-th frame to obtain the third fused feature map Fuse the second feature map of the t-th frame, the third fused feature map, and the second feature map of the (t - 1)-th frame to obtain the second fused feature map Fuse the first feature map of the t-th frame, the second fused feature map, and the first feature map of the (t - 1)-th frame to obtain the first fused feature map Finally, five fused feature maps of different scales of the t-th frame are obtained, which is only for example and not limited thereto. It should be noted that the k-th feature map of the t-th frame has the same scale as the k-th feature map of the (t - 1)-th frame.

[0094] In some alternative embodiments, in step S2022, temporal multi-scale feature fusion is performed based on the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale, including:

[0095] Step a1, perform a first preset convolution process on the feature map of the current frame at the current scale to obtain a first feature map, perform a preset upsampling process on the fused feature map of the current frame at the previous scale to obtain a second feature map, and perform a second preset convolution process on the feature map of the previous frame at the current scale to obtain a third feature map.

[0096] Specifically, the structure for feature fusion in the preset feature fusion module is FLNet, and FLNet has three inputs: Input1 is the k-th feature map of the t-th frame, Input2 is the k-1 fused feature map of the t-th frame (if it exists), and Input3 is the k-th feature map of the t-1-th frame. The processing process of FLNet for the three inputs includes: Input1 first undergoes a convolution process (Convolution, Conv) with a convolution kernel of 1×1 to obtain a first feature map, Input2 undergoes bilinear interpolation upsampling by a factor of 2 to obtain a second feature map, and Input3 undergoes a convolution process with a convolution kernel of 3x3 to obtain a third feature map.

[0097] Step a2, add the first feature map and the second feature map and then concatenate them with the third feature map to obtain a concatenated feature map.

[0098] Specifically, add the first feature map and the second feature map and then concatenate them with the third feature map in the channel dimension to obtain a concatenated feature map.

[0099] Step a3, perform a depthwise separable convolution process on the concatenated feature map to obtain the fused feature map of the current frame at the current scale.

[0100] Specifically, after passing the concatenated feature map through two depthwise separable convolution processes, an output feature map is obtained, where the depthwise separable convolution can significantly reduce the model parameters and computational amount after the features are concatenated in the channel dimension.

[0101] For feature maps with five different scales, the overall processing process of the preset feature fusion module includes: the fifth feature map of the t-th frame the fifth feature map of the t-1-th frame pass through the fifth FLNet to obtain the fifth fused feature of the t-th frame the fourth feature map of the t-th frame the fourth feature map of the t-1-th frame the fifth fused feature of the t-th frame pass through the fourth FLNet to obtain the fourth fused feature of the t-th frame the third feature map of the t-th frame The third feature map of the (t - 1)-th frame The fourth fused feature of the t-th frame The third fused feature of the t-th frame is obtained through the processing of the third FLNet The second feature map of the t-th frame The second feature map of the (t - 1)-th frame The third fused feature of the t-th frame The second fused feature of the t-th frame is obtained through the processing of the second FLNet The first feature map of the t-th frame The first feature map of the (t - 1)-th frame The second fused feature The first fused feature of the t-th frame is obtained through the processing of the first FLNet

[0102] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided in this embodiment enhances the robustness of feature expression by fusing the same-scale features of adjacent temporal frames and different-scale features of the same temporal frame. The depthwise separable convolution processing can significantly reduce the model parameters and computational amount after the features are concatenated in the channel dimension, and improve the model processing speed.

[0103] Step S203: Sort the multiple fused feature maps of the current frame according to the scale size to obtain the fused feature map with the smallest scale and multiple other-scale fused feature maps. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.

[0104] Step S204: Process the fused feature map with the smallest scale using a preset attention module to obtain the encoded feature of the current frame. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.

[0105] Step S205: Process the other-scale fused feature maps using a preset skip connection module respectively to obtain multiple skip connection features of the current frame.

[0106] Specifically, the above step S205 includes:

[0107] Input the other-scale fused feature maps into the preset skip connection module respectively to perform the first preset convolution processing to obtain the skip connection features corresponding to the current frame at the corresponding scale.

[0108] Specifically, the preset skip connection module includes multiple convolutional layers with a convolution kernel of 1×1. The first fused feature of the t-th frame After being processed by the convolution with a convolution kernel of 1×1, the first skip connection feature of the t-th frame is obtained The second fused feature of the t-th frame After convolution processing with a 1×1 convolution kernel, the second skip connection feature of the t-th frame is obtained The third fusion feature of the t-th frame After convolution processing with a 1×1 convolution kernel, the third skip connection feature of the t-th frame is obtained The fourth fusion feature of the t-th frame After convolution processing with a 1×1 convolution kernel, the fourth skip connection feature of the t-th frame is obtained

[0109] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided in this embodiment enables low-level features (such as edges, textures, etc.) in the network to be directly transmitted to higher-level feature maps by performing skip connections on fusion features of other scales, which is beneficial to maintaining spatial information and improving the accuracy of target detection.

[0110] Step S206: According to the encoded feature and multiple skip connection features, perform decoding to obtain the output feature map of the current frame.

[0111] Specifically, the above step S206 includes:

[0112] Step S2061: Sort the multiple skip connection features from smallest to largest in scale to obtain a skip connection feature sequence.

[0113] Specifically, as Figure 5 shown, the scale of the fourth skip connection feature is 1 / 16, the scale of the third skip connection feature is 1 / 8, the scale of the second skip connection feature is 1 / 4, and the scale of the first skip connection feature is 1 / 2. Then the skip connection feature sequence is: the fourth skip connection feature, the third skip connection feature, the second skip connection feature, the first skip connection feature, which is only for example and not limited thereto.

[0114] The decoding process is implemented through a preset decoding module. The preset decoding module includes multiple DCNet asymmetric convolution structures. The processing process of DCNet includes:

[0115] (1) Receive the input feature map input, and divide input into a first branch and a second branch for separate processing.

[0116] (2) The first branch is sequentially processed through convolutions with 1×3 and 3×1 convolution kernels to obtain the first branch feature, and the second branch is sequentially processed through convolutions with 3×1 and 1×3 convolution kernels to obtain the second branch feature.

[0117] (3) After adding the first branch feature and the second branch feature, an intermediate feature is obtained.

[0118] (4) Process the intermediate features through the third and fourth branches to obtain the third branch features and the fourth branch features: The third branch consists of convolutions with kernel sizes of 1×3 and 3×1, and the fourth branch is processed by convolutions with kernel sizes of 3×1 and 1×3.

[0119] (5) After adding the third branch features and the fourth branch features, perform batch normalization and ReLu activation function processing to finally obtain the output feature map output corresponding to the input feature map input.

[0120] Step S2062: According to the encoded features and the skip connection feature sequence, iteratively perform preset asymmetric convolution processing to obtain the output encoded features of the current frame.

[0121] Specifically, as Figure 5 shown, the processing process of the preset decoding module includes: The first encoder features of the t-th frame First, after being processed by the first DCNet, are upsampled by bilinear interpolation by a factor of two and then added to the fourth skip connection feature The result of the addition is processed by the second DCNet, then upsampled by bilinear interpolation by a factor of two and added to the third skip connection feature The result of the addition is processed by the third DCNet, then upsampled by bilinear interpolation by a factor of two and added to the second skip connection feature The result of the addition is processed by the fourth DCNet, then upsampled by bilinear interpolation by a factor of two and added to the first skip connection feature to obtain the output encoded features of the t-th frame.

[0122] Step S2063: Perform a second preset convolution processing on the output encoded features of the current frame to obtain the output feature map of the current frame.

[0123] Specifically, the output encoded features of the t-th frame are processed by a convolution with a kernel size of 3×3 to obtain the output feature map of the current frame.

[0124] The bipedal humanoid robot pose detection method based on temporal multi-scale feature fusion provided in this embodiment can achieve more powerful learning ability without significantly increasing the number of parameters by using skip connections and feature addition. Since the amount of data to be processed is reduced, the computational cost can be reduced and the inference speed can be accelerated.

[0125] In a specific embodiment, the camera captures a color image I t with a size of 1920×1080. After subtracting the mean values [123.68, 116.779, 103.939] from the three RGB channels of the image respectively, and then dividing by the variances [58.393, 57.12, 57.375] for normalization, and then performing size adjustment to obtain the standard input image I of the deep networkt ‘’, with dimensions H = 384 and W = 384. The feature extraction module is composed of ResNet34. The dimensions of the first feature map, second feature map, third feature map, fourth feature map, and fifth feature map output by its first feature extractor, second feature extractor, third feature extractor, fourth feature extractor, and fifth feature extractor are 192×192×64, 96×96×64, 48×48×128, 24×24×256, and 12×12×512 respectively.

[0126] In the preset feature fusion module, the parameters of the two groups of depthwise separable convolutions in FLNet are as follows: the kernel size of the depthwise convolution is 3×3, and the number of output channels of the pointwise convolution is 128. The process of temporal multi-scale feature fusion includes:

[0127] In the fifth FLNet, the size of Input1 is 12×12×512, Input2 does not exist, and the size of Input3 is 12×12×512; the size of the feature map obtained after Input1 is processed by 1×1 convolution is 12×12×512, and the size of the feature map obtained after Input3 is processed by 3×3 convolution is 12×12×512; the size of the output fifth fusion feature is 12×12×128.

[0128] In the fourth FLNet, the size of Input1 is 24×24×256, the size of Input2 is 12×12×512, and the size of Input3 is 24×24×256; the size of the feature map obtained after Input1 is processed by 1×1 convolution is 24×24×512, and the size of the feature map obtained after Input3 is processed by 3×3 convolution is 24×24×512; the size of the output fourth fusion feature is 24×24×128.

[0129] In the third FLNet, the size of Input1 is 48×48×128, the size of Input2 is 24×24×256, and the size of Input3 is 48×48×128; the size of the feature map obtained after Input1 is processed by 1×1 convolution is 48×48×256, and the size of the feature map obtained after Input3 is processed by 3×3 convolution is 48×48×256; the size of the output third fusion feature is 48×48×128.

[0130] In the second FLNet, the size of Input1 is 96×96×64, the size of Input2 is 48×48×128, and the size of Input3 is 96×96×64; the size of the feature map obtained after Input1 is processed by 1×1 convolution is 96×96×128, and the size of the feature map obtained after Input3 is processed by 3×3 convolution is 96×96×128; the size of the output second fusion feature is 96×96×128.

[0131] In the first FLNet, the size of Input1 is 192×192×64, the size of Input2 is 96×96×64, and the size of Input3 is 192×192×64; the size of the feature map obtained after Input1 is processed by 1×1 convolution is 192×192×64, and the size of the feature map obtained after Input3 is processed by 3×3 convolution is 192×192×64; the size of the output first fusion feature is 192×192×128.

[0132] In the preset skip connection module, the sizes of the first skip connection feature, the second skip connection feature, the third skip connection feature, the fourth skip connection feature, and the fifth skip connection feature are 192×192×64, 96×96×64, 48×48×64, 24×24×64, and 12×12×64 respectively.

[0133] In the preset decoding module, the sizes of the output encoded features of the first DCNet, the second DCNet, the third DCNet, and the fourth DCNet are 12×12×64, 24×24×64, 48×48×64, and 96×96×64 respectively.

[0134] Finally, the size of the output feature map is 192×192×(13 + 14).

[0135] To obtain the parameters of each module in the above deep network, a mean square error loss function is constructed using the output of the network and the labeled ground truth, and training is performed with training samples to obtain the parameters of each module of the above neural network.

[0136] In this embodiment, a bipedal humanoid robot posture detection device based on temporal multi-scale feature fusion is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0137] This embodiment provides a bipedal humanoid robot posture detection device based on temporal multi-scale feature fusion, as Figure 6 shown, including:

[0138] The multi-scale feature extraction module 601 is configured to continuously obtain multiple frames of original images of the target bipedal humanoid robot, and respectively preprocess each original image and then extract multiple feature maps of different scales.

[0139] The temporal multi-scale feature fusion module 602 is configured to iteratively perform temporal multi-scale feature fusion based on the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale, to obtain multiple fused feature maps of the current frame.

[0140] The scale sorting module 603 is configured to sort the multiple fused feature maps of the current frame according to the scale size, to obtain the fused feature map of the smallest scale and multiple fused feature maps of other scales.

[0141] The attention mechanism processing module 604 is configured to process the fused feature map of the smallest scale by using a preset attention module to obtain the encoded feature of the current frame.

[0142] The skip connection processing module 605 is configured to process the fused feature maps of other scales respectively by using a preset skip connection module to obtain multiple skip connection features of the current frame.

[0143] The feature decoding module 606 is configured to perform decoding based on the encoded feature and multiple skip connection features to obtain the output feature map of the current frame.

[0144] In some alternative embodiments, the multi-scale feature extraction module 601 includes:

[0145] An image normalization unit, configured to normalize the original image to obtain a standard image suitable for input to the deep network.

[0146] A feature extraction unit, configured to perform feature extraction of different scales on the standard image to obtain multiple feature maps of different scales.

[0147] In some alternative embodiments, the temporal multi-scale feature fusion module 602 includes:

[0148] A minimum fusion unit, configured to perform feature fusion on the feature map of the smallest scale of the current frame and the feature map of the smallest scale of the previous frame to obtain the fused feature map of the smallest scale of the current frame.

[0149] An iterative fusion unit, configured to use the fused feature map of the smallest scale of the current frame as the fused feature map of the current frame at the previous scale during the first temporal multi-scale feature fusion process, and repeatedly perform temporal multi-scale feature fusion based on the feature map of the current frame at the current scale, the fused feature map of the current frame at the previous scale, and the feature map of the previous frame at the current scale, to obtain multiple fused feature maps of different scales of the current frame.

[0150] In some alternative embodiments, the iterative fusion unit includes:

[0151] A feature convolution processing subunit, configured to perform a first preset convolution processing on the feature map of the current frame at the current scale to obtain a first feature map, perform a preset upsampling processing on the fused feature map of the current frame at the previous scale to obtain a second feature map, and perform a second preset convolution processing on the feature map of the previous frame at the current scale to obtain a third feature map.

[0152] A feature splicing subunit, configured to splice the first feature map and the second feature map after addition with the third feature map to obtain a spliced feature map.

[0153] A depthwise separable convolution processing subunit, configured to perform depthwise separable convolution processing on the spliced feature map to obtain the fused feature map of the current frame at the current scale.

[0154] In some alternative embodiments, the feature decoding module 606 includes:

[0155] A skip connection feature sorting unit, configured to sort a plurality of skip connection features from smallest to largest in scale to obtain a skip connection feature sequence.

[0156] An iterative decoding unit, configured to iteratively perform a preset asymmetric convolution processing according to the encoded feature and the skip connection feature sequence to obtain the output encoded feature of the current frame.

[0157] An output feature determination unit, configured to perform a second preset convolution processing on the output encoded feature of the current frame to obtain the output feature map of the current frame.

[0158] The further function descriptions of the above-mentioned respective modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.

[0159] The bipedal humanoid robot pose detection device based on temporal multi-scale feature fusion in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0160] The embodiment of the present invention further provides a computer device having the above-mentioned Figure 6 shown bipedal humanoid robot pose detection device based on temporal multi-scale feature fusion.

[0161] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 7As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 7 In [the figure], one processor 10 is taken as an example.

[0162] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0163] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0164] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0165] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0166] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0167] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0168] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A bipedal humanoid robot posture detection method based on temporal multi-scale feature fusion, characterized in that: The method comprises: Continuously acquire multiple frames of original images of the target bipedal humanoid robot, and extract multiple feature maps of different scales after preprocessing each original image; Iteratively perform temporal multi-scale feature fusion according to the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame to obtain multiple fused feature maps of the current frame; Sort multiple fused feature maps of the current frame according to their scales to obtain a fused feature map of the smallest scale and multiple fused feature maps of other scales; Processing the minimum-scale fusion feature map using a preset attention module to obtain a coding feature of the current frame; Use the preset skip connection module to process the fusion feature maps of other scales respectively to obtain multiple skip connection features of the current frame; According to the encoding features and multiple jump connection features, decoding is performed to obtain an output feature map of the current frame.

2. The method according to claim 1, characterized in that After preprocessing each original image, multiple feature maps of different scales are extracted, including: Normalize the original image to obtain a standard image suitable for deep network input; Feature extraction of different scales is performed on the standard image to obtain a plurality of feature maps of different scales.

3. The method according to claim 1, characterized in that According to the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame, the temporal multi-scale feature fusion is iteratively performed to obtain multiple fused feature maps of the current frame, including: The feature map of the minimum scale of the current frame is fused with the feature map of the minimum scale of the previous frame to obtain the fused feature map of the minimum scale of the current frame; The fused feature map of the minimum scale of the current frame is used as the fused feature map of the previous scale of the current frame in the first temporal multi-scale feature fusion process. The temporal multi-scale feature fusion is repeatedly performed based on the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame to obtain multiple fused feature maps of different scales in the current frame.

4. The method according to claim 3, characterized in that The step of performing temporal multi-scale feature fusion according to the feature map of the current scale of the current frame, the fused feature map of the previous scale of the current frame, and the feature map of the current scale of the previous frame includes: Performing a first preset convolution process on a feature map of a current frame at a current scale to obtain a first feature map, performing a preset upsampling process on a fused feature map of a previous scale of the current frame to obtain a second feature map, and performing a second preset convolution process on a feature map of a previous frame at a current scale to obtain a third feature map; Adding the first feature map and the second feature map together and concatenating the sum with the third feature map to obtain a concatenated feature map; The concatenated feature map is subjected to a depthwise separable convolution process to obtain a fused feature map of a current frame and a current scale.

5. The method according to claim 1, characterized in that: The preset skip connection module is used to process the fusion feature maps of other scales respectively to obtain multiple skip connection features of the current frame, including: The fused feature maps of other scales are respectively input into the preset skip connection module, and the first preset convolution process is performed to obtain the skip connection features of the scale corresponding to the current frame.

6. The method according to claim 1, characterized in that According to the coding feature and the multiple skip connection features, decoding is performed to obtain an output feature map of the current frame, including: Sort multiple skip connection features from small to large scales to obtain a skip connection feature sequence; Iteratively performing a preset asymmetric convolution process according to the coding feature and the skip connection feature sequence to obtain an output coding feature of the current frame; A second preset convolution process is performed on the output coding features of the current frame to obtain an output feature map of the current frame.

7. The method according to claim 1 or 6, characterized in that: The method further comprises: Extract multiple key points and skeleton of the target bipedal humanoid robot according to the output feature map of the current frame; The output features are visualized according to the key points and the skeleton to obtain a thermal map of the target bipedal humanoid robot.

8. A bipedal humanoid robot posture detection device based on temporal multi-scale feature fusion, characterized in that: The device comprises: A multi-scale feature extraction module is used to continuously acquire multiple frames of original images of the target bipedal humanoid robot, and extract multiple feature maps of different scales after preprocessing each original image; A temporal multi-scale feature fusion module is used to iteratively fuse temporal multi-scale features according to a feature map of the current scale of the current frame, a fused feature map of the previous scale of the current frame, and a feature map of the current scale of the previous frame, to obtain multiple fused feature maps of the current frame; The scale sorting module is used to sort multiple fused feature maps of the current frame according to the scale size to obtain the fused feature map of the smallest scale and multiple fused feature maps of other scales; An attention mechanism processing module, used to process the minimum-scale fusion feature map using a preset attention module to obtain a coding feature of the current frame; A skip connection processing module, used to use a preset skip connection module to process fusion feature maps of other scales respectively, to obtain multiple skip connection features of the current frame; The feature decoding module is used to decode the encoded features and multiple jump connection features to obtain an output feature map of the current frame.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method fusing multi-scale spatial-temporal characteristics

    CN116229304A

  • Gait recognition method and system based on pedestrian time sequence contour reconstruction and restoration

    CN118762397A