A method and system for motion recognition and imitation of a humanoid robot
By using grayscale differential background segmentation method and tracking algorithm in humanoid robots, combined with static and dynamic recognition technology to extract and imitate motion characteristics, the problems of motion capture and imitation accuracy and cost in the existing technology are solved, and more efficient and accurate motion imitation is achieved.
Patent Information
- Application Number
- CN202411203627.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-08-29
AI Technical Summary
The prior art has high precision but expensive optical or inertial sensor-limited application scenarios in the motion capture and imitation of humanoid robots, and the calculation of action behavior mapping technology is large and complex, affecting the similarity and real-time nature of action imitation.
The static scene view is extracted from the original motion video by using the grayscale differential background segmentation method. The motion trajectory of the moving target is tracked through the tracking algorithm, key action frames are captured and fused action feature extraction is performed. Combined with static recognition and dynamic recognition, action descriptions are generated and sent to the humanoid robot for action imitation.
It realizes more accurate motion description and more efficient motion imitation, reduces dependence on equipment, improves the similarity and accuracy of motion imitation, and lays the foundation for the widespread application of humanoid robots.
Smart Images

Figure CN119169502B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of humanoid robots, and in particular to a method and system for motion recognition and imitation of a humanoid robot. Background Art
[0002] As an intelligent device with human appearance and behavioral capabilities, humanoid robots have demonstrated their unique application value in many fields, especially in medical, industrial, entertainment, academic research and daily life. In the medical field, especially for the treatment of children with autism, they are usually afraid of social contact and have communication barriers, but they tend to show a higher acceptance of interacting with robots. By imitating simple body movements, humanoid robots can help these children improve their imitation and learning abilities and promote the development of social skills; in industrial production, as the industrial environment becomes increasingly complex and harsh, human operations in certain high-risk environments become more and more dangerous. Humanoid robots can replace humans in these environments by imitating human movements, thereby reducing workers' exposure to danger, liberating labor and improving production efficiency; today, as the human-computer interaction experience is increasingly valued, the entertainment field has a great demand for humanoid robots. The demand for robots is also growing. Humanoid robots can imitate various human movements, such as walking, dancing, playing table tennis, playing football, etc., thus bringing novel entertainment experiences to users. In academic research, humanoid robots, as a practical and convenient academic platform, can be used to study multidisciplinary theories such as image recognition, video retrieval, and motor control. Through robot motion imitation, scholars can verify and develop these theories in reality and promote technological progress. In daily life, humanoid robots are gradually replacing humans to complete some repetitive and physical labor, such as doing housework, carrying heavy objects, etc., which not only improves the convenience of life, but also provides a development direction for future smart homes.
[0003] Although humanoid robots have shown broad application prospects in many fields, existing technologies still have some shortcomings in motion capture and imitation. Traditional motion capture devices, such as optical or inertial sensors, are expensive and have strict environmental requirements despite their high accuracy, which limits their application scenarios. In terms of humanoid robot motion imitation, existing motion behavior mapping technology mainly relies on inverse kinematics solution, which is computationally intensive and complex, and may affect the similarity and real-time performance of motion imitation. Although the method of solving using spatial vectors is computationally simple, it lacks accurate analysis of the relationship between the skeletal structure and degrees of freedom of the human body and the robot, resulting in the need to improve the similarity and accuracy of motion imitation. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a method and system for motion recognition and imitation of a humanoid robot, which achieves more accurate motion description and more efficient motion imitation by integrating static recognition and dynamic recognition, which not only reduces the dependence on equipment, but also improves the similarity and accuracy of motion imitation, laying the foundation for the widespread application of humanoid robots.
[0005] A method for motion recognition and imitation of a humanoid robot, comprising:
[0006] The grayscale difference background segmentation method is used to extract the static scene view containing the moving target to be identified from the original motion video.
[0007] A moving object is separated from the static scene view.
[0008] A tracking algorithm is used to track the motion trajectory of the moving target.
[0009] Capture the key action frames of the moving target, extract fused action features of the moving target, and obtain a fused action feature vector of the moving target.
[0010] The key action frames are annotated with action labels, and the motion target fusion action feature vectors and the corresponding action labels are sorted to obtain a training data set.
[0011] The training data set is trained using the target fused action feature vector, and static recognition and dynamic recognition are combined to obtain an action description of the moving target.
[0012] The action description is sent to the humanoid robot for action imitation.
[0013] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, the grayscale difference background segmentation method is used to extract a static scene view containing a moving target to be identified from an original motion video, including:
[0014] Extract several initial image frames from the original motion video and calculate the average grayscale value of these frames as the initial background model Among them, G t (x, y) is the grayscale value of the tth frame, (x, y) is the pixel coordinate in each of the image frames, and N is the number of the extracted image frames.
[0015] The initial background model is updated using dynamic change smoothing to obtain an updated background model B t+1 (x,y)=(1-α)×B t (x,y)+α×G t+1 (x,y), where α is the smoothing factor, 0<α<1.
[0016] For each frame, the grayscale difference between the current image frame and the initial background model is calculated to obtain the grayscale difference model ΔG t (x,y)=|G t (x,y)-B t (x,y)∣.
[0017] For the grayscale difference model ΔG t (x, y), distinguish the foreground area from the background area by setting the threshold P, the scene foreground mask A value of 1 in the scene foreground mask indicates a foreground area, and a value of 0 in the scene foreground mask indicates a background area.
[0018] Using the scene foreground mask M t (x,y) extracts the foreground region F from the current frame t (x,y)=G t (x,y)×M t (x,y).
[0019] The average value of the cumulatively extracted foreground areas in all the initial image frames is calculated to obtain a cumulative foreground area for representing a static scene view containing a moving target to be identified. Wherein, T is the total number of frames used to extract static views.
[0020] The technical effect is as follows: by calculating the average gray value of the image frame and using dynamic changes to smoothly update the background model, it can effectively capture and update the background information of the scene, improve the accuracy of the background model, adapt to slow changes in the environment, reduce the interference of instantaneous dynamic changes, and ensure the stability and accuracy of the background model over a long period of time; through the setting of the gray difference model and the threshold P, the foreground area is clearly distinguished from the background area, the foreground area of the dynamic target is accurately extracted, and the influence of background interference on the extraction of the moving target is reduced.
[0021] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, a smoothing filter is applied to the accumulated foreground area for smoothing.
[0022] The technical effect is that by applying a smoothing filter to smooth the accumulated foreground area, the image can be effectively smoothed, the noise introduced by the camera equipment, environmental changes or segmentation algorithms can be reduced, the smoothness of the target edge can be enhanced, the continuity of the target area can be improved, the target area can be made more coherent in space, the stability of feature extraction can be enhanced, the foreground area after smoothing has fewer details and a more uniform distribution, and the computational complexity of the system can be reduced.
[0023] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, the step of separating the moving target from the static scene view comprises:
[0024] From the static scene view, a grayscale difference between a current image frame and the accumulated foreground region is calculated.
[0025] The threshold is dynamically adjusted according to the statistical characteristics of the grayscale difference map, and a moving object foreground mask is generated according to the threshold.
[0026] The moving object foreground mask is used to extract the moving object from the current image frame.
[0027] Its technical effects are: by calculating the grayscale difference between the current image frame and the accumulated foreground area, the difference between the moving target and the background can be more accurately identified; the threshold is dynamically adjusted by using the statistical characteristics of the grayscale difference map to adapt to different scenes and changes in lighting conditions, and the accuracy of segmentation is improved by dynamically adjusting the threshold.
[0028] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, the step of tracking the motion trajectory of the moving target using a tracking algorithm comprises:
[0029] According to the initial area of the moving target, set the tracking window with a width of Height is Wherein, (x0, y0) is the initial center coordinate of the moving target, W0 is the initial width of the tracking window, and H0 is the initial height of the tracking window.
[0030] Perform feature extraction on each pixel (x, y) in the current tracking window and calculate the gray value I t (x, y), calculate the relative feature weight of each pixel in the current tracking window Generate a weight distribution graph for representing the weighted feature distribution of each pixel in the current tracking window, using the weight distribution model W dist (x,y)=w(x,y)×I t (x,y) represents.
[0031] The weighted center position (x c ,y c ),in,
[0032] According to the concentration of weight distribution, adjust the current tracking window, the width is W t+1 =α·W t+(1-α)·Var(W dist ), height is H t+1 =α·H t +(1-α)·Var(H dist ), where α is the smoothing factor, Var(W dist ) and Var(H dist ) is the variance of the weight distribution graph in the x and y directions.
[0033] According to the position of the weighted center, the tracking window in each frame is updated, and the position is (x t+1 =x c ,y t+1 =y c ), with a width of W t+1 , height is H t+1 .
[0034] The tracking window is repeatedly tracked for each frame of the original motion video until the original motion video ends or the moving target is lost.
[0035] Its technical effects are: by setting the tracking window and calculating the weighted feature distribution of pixels in each frame, the position and shape changes of the moving target can be accurately tracked, the continuity of the target between consecutive frames can be maintained, and the possibility of target loss can be reduced; the size of the tracking window can be dynamically adjusted according to the concentration of the weight distribution to adapt to changes in the target size. When the target moves, it may be enlarged or reduced. By adjusting the width and height of the tracking window, the actual size of the target can be better captured, ensuring that the tracking window always covers the target, and effectively preventing the target from leaving the tracking range.
[0036] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the capturing of the key motion frame of the moving target, extracting the fused motion feature of the moving target, and obtaining the fused motion feature vector of the moving target comprises:
[0037] In the video of the motion trajectory of the moving target, the difference between each frame and the previous frame is calculated. When the difference exceeds a set threshold, the frame is identified as a key action frame.
[0038] For each of the key action frames, extract the geometric feature vector G j =[R j ,A r ,P r ,SM t ], the geometric morphological feature vector, where R j is the aspect ratio of the target area, A r is the ratio of the area to the area of the circumscribed rectangle, P r is the ratio of the perimeter to the perimeter of the circumscribed rectangle, SMt is the shape moment parameter, and the calculation formula is
[0039] For each of the key action frames, the global invariant feature vector H is calculated by global invariant moment. j .
[0040] The geometric feature vector and the global invariant feature vector are combined to obtain a fused motion feature vector V j =[G j ,H j ].
[0041] Its technical effects are: by calculating the difference between each frame and the previous frame and setting a threshold, important action change moments in the video can be effectively identified, and the captured key action frames can represent the significant action features of the target; by fusing geometric morphological features and global invariant features, the action features of the target are comprehensively described from different angles. The geometric morphological features reflect the changes in the target's morphology and geometric structure, while the global invariant features provide a description of shape invariance, rotation and scaling invariance. By merging these two types of features, the target's action state and change trend can be more accurately portrayed, forming a comprehensive description of the action.
[0042] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, the key motion frames are annotated with motion labels, and the motion target fusion motion feature vector and the corresponding motion label are sorted to obtain a training data set, which includes:
[0043] Define the type of action tag, including at least one of walking, jumping, and waving.
[0044] The fused action feature vector of the moving target in each of the key action frames is combined with the corresponding action label to obtain a training data set.
[0045] Its technical effect is: by annotating key action frames with action labels and organizing them with fused action feature vectors, a high-quality training data set is obtained, which helps to improve the classification accuracy of the model, enhance the representativeness of the training data set, and optimize the learning effect of the model.
[0046] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the training data set is trained using the target fusion motion feature vector, and the motion description of the moving target is obtained by combining static recognition and dynamic recognition, which includes:
[0047] The geometric morphology feature vector is processed by a linear classifier to obtain a geometric morphology linear classification output f G (Gj )=W G ·G j +b G , where W G is the weight matrix of the geometric morphology feature vector, with a dimension of C×d G , d G is the geometric morphology feature vector G j The dimension of C is the number of categories, b is G is the bias term of geometric features.
[0048] The globally invariant feature vector is processed by a linear classifier to obtain a globally invariant linear classification output f H (H j )=W H ·H j +b H , where W H is the weight matrix of the global invariant feature vector, with dimension C×d H , d H is the global invariant eigenvector H j The dimension of C is the number of categories, b is H is the bias term of geometric features.
[0049] The geometric linear classification output and the global invariant linear classification output are fused by weighted averaging to obtain a classification probability distribution model for static recognition. Among them, α is the weight parameter of the geometric morphological feature vector, β is the weight parameter of the global invariant feature vector, and α+β=1.
[0050] A 3D convolutional spatiotemporal network model is established, wherein the 3D convolutional spatiotemporal network model includes an input layer, a downsampling layer, a residual layer and an output layer.
[0051] The input layer uses a 3D convolution kernel with a size of 3×3×4 and a step size of 1×2×2. The input layer outputs a feature map with a size of N×T×C1×H / 4×W / 4 as the input of the residual layer, where N is the batch size, T is the number of frames, C1 is the number of channels, H is the height, and W is the width.
[0052] The downsampling layer uses a 3D convolution kernel with a size of 1×2×2 and a step size of 1×2×2.
[0053] Each convolution block of the residual layer includes a 3×3×3 convolution kernel layer, a 3×7×7 convolution kernel layer, an MLP layer, a normalization layer and a GELU activation function layer, respectively, to obtain an output feature map of the residual layer y=x+MLP(GELU(MLP(LN(conv3(x)+conv7(x)+conv3x3x3(x))))), where x is the input of each convolution block.
[0054] The output layer uses a global average pooling layer and a fully connected layer for final classification. The feature map size of the global average pooling layer output is N×512, and the output of the fully connected layer is N×C, which serves as the final output of the 3D convolutional spatiotemporal network model.
[0055] The 3D convolutional spatiotemporal network model is trained to obtain a dynamic recognition classification probability distribution model Among them, W d is the weight matrix of the fully connected layer, b d is the bias vector of the fully connected layer.
[0056] The classification probability distribution model of the static recognition and the classification probability distribution model of the dynamic recognition are weightedly fused to obtain a fusion model Among them, γ s is the weight parameter of the static recognition result, γ d is the weight parameter of the dynamic recognition result, γ s +γ d =1.
[0057] Select the category with the highest probability The final action classification result for the action description of the moving target is obtained.
[0058] The technical effects are as follows: the geometric morphology feature vector and the global invariant feature vector are processed by linear classifiers respectively, so that the static recognition model can accurately extract key information from each feature vector and classify it. By weighted fusion of the geometric morphology linear classification output and the global invariant linear classification output, static recognition and dynamic recognition are combined to effectively enhance the classification accuracy. The 3D convolutional spatiotemporal network model is used to process dynamic video data to capture spatiotemporal information, thereby improving the recognition ability of moving target movements. The downsampling layer and the residual layer in the 3D convolutional spatiotemporal network model reduce the computational complexity and extract higher-level features, thereby enhancing the model's ability to process input data and thus improving the dynamic recognition ability of moving targets. When training the 3D convolutional spatiotemporal network model, a random initialization strategy and an exponential moving average method (EMA) are used to reduce network overfitting, thereby ensuring the stability of the training process and the generalization ability of the model.
[0059] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, the step of sending the motion description to the humanoid robot for motion imitation comprises:
[0060] The classification results of the static recognition and the dynamic recognition are converted into action descriptions in a standard format.
[0061] The action description in a standard format is converted into an action instruction of a humanoid robot, and sent to the humanoid robot through a communication interface.
[0062] The humanoid robot receives instructions, analyzes and executes corresponding actions and provides feedback on the execution status of the actions.
[0063] Its technical effect is: through standardized action description, adaptive control protocol, efficient communication and feedback mechanism, the accuracy of action simulation, system compatibility and user experience are improved.
[0064] A humanoid robot motion recognition and imitation system, comprising:
[0065] The static scene view extraction module is used to extract the static scene view containing the moving target to be identified from the original motion video by adopting the grayscale difference background segmentation method.
[0066] The moving target separation module is used to separate the moving target from the static scene view.
[0067] The motion trajectory tracking module is used to track the motion trajectory of the moving target using a tracking algorithm.
[0068] The feature vector extraction module is used to capture the key action frame of the moving target, extract the fused action feature of the moving target, and obtain the fused action feature vector of the moving target.
[0069] The label marking module is used to mark the key action frames with action labels, and to sort the motion target fusion action feature vectors and the corresponding action labels to obtain a training data set.
[0070] The 3D convolutional spatiotemporal network model building module is used to train the training data set using the target fusion action feature vector, and combine static recognition and dynamic recognition to obtain the action description of the moving target.
[0071] The action execution module is used to send the action description to the humanoid robot for action imitation.
[0072] The beneficial effects of the present invention are:
[0073] The present invention utilizes grayscale difference background segmentation technology to extract static scene views from original videos, and reduces the impact of instantaneous dynamic changes by dynamically updating the background model, thereby providing more accurate background and foreground separation and being able to effectively extract clear moving target areas from complex backgrounds.
[0074] The present invention applies a tracking algorithm that can track the position changes of the moving target in each frame in real time and accurately, ensuring the coherent tracking of the moving target throughout the entire action process, so that the motion trajectory of the moving target is completely recorded, which not only improves the accuracy of motion feature extraction, but also avoids data loss due to target loss or occlusion.
[0075] The present invention provides an in-depth description of the motion of a moving target by fusing geometric features and global invariant features. The geometric features capture the shape and size changes of the moving target, while the global invariant features take into account the stable characteristics of the moving target in different postures. This multi-dimensional feature fusion enables the action recognition model to understand and describe the motion characteristics of the moving target more comprehensively.
[0076] The present invention establishes a 3D convolutional spatiotemporal network model, which makes full use of the spatiotemporal correlation information between frames in the video. Through the 3D convolution operation, the model can simultaneously process information in the time and space dimensions, thereby improving the accuracy and robustness of action recognition, enabling the model to better capture and understand the dynamic changes and time series characteristics of the action. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0078] Figure 1 The figure is a flow chart of the motion recognition and imitation method of the humanoid robot of the present invention. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0080] Please refer to Figure 1The first embodiment of the present invention provides a method for motion recognition and imitation of a humanoid robot, which includes: using a grayscale difference background segmentation method to extract a static scene view containing a moving target to be identified from an original motion video; separating the moving target from the static scene view; using a tracking algorithm to track the motion trajectory of the moving target; capturing key action frames of the moving target, extracting fused action features of the moving target, and obtaining a fused action feature vector of the moving target; annotating the key action frames with action labels, and arranging the fused action feature vector of the moving target and the corresponding action labels to obtain a training data set; using the target fused action feature vector to train the training data set, combining static recognition and dynamic recognition to obtain a description of the action of the moving target; and sending the action description to the humanoid robot for motion imitation.
[0081] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the grayscale difference background segmentation method is used to extract a static scene view containing a moving target to be identified from the original motion video, which includes: extracting a number of initial image frames from the original motion video, calculating the average grayscale value of these frames as the initial background model Among them, G t (x, y) is the gray value of the tth frame, (x, y) is the pixel coordinate in each image frame, and N is the number of extracted image frames; the initial background model is updated using dynamic change smoothing to obtain an updated background model B t+1 (x,y)=(1-α)×B t (x,y)+α×G t+1 (x, y), where α is a smoothing factor, 0<α<1. The initial background model is updated by using dynamic change smoothing to make it gradually smooth and weaken the impact of instantaneous dynamic changes. A larger α value will make the background model update faster, while a smaller α value will make it more stable. For each frame, the grayscale difference between the current image frame and the initial background model is calculated to obtain the grayscale difference model ΔG t (x,y)=|G t (x,y)-B t (x, y) |; for the grayscale difference model ΔG t (x, y), distinguish the foreground area from the background area by setting the threshold P, the scene foreground mask The value of the scene foreground mask M is 1, which indicates the foreground area, and the value of the scene foreground mask M is 0, which indicates the background area. t (x,y) extracts the foreground region F from the current frame t (x,y)=G t (x,y)×M t(x, y); Calculate the average value of the cumulatively extracted foreground areas in all the initial image frames to obtain the cumulative foreground area used to represent the static scene view containing the moving target to be identified Wherein, T is the total number of frames used to extract static views.
[0082] The technical effect is as follows: by calculating the average gray value of the image frame and using dynamic changes to smoothly update the background model, it can effectively capture and update the background information of the scene, improve the accuracy of the background model, adapt to slow changes in the environment, reduce the interference of instantaneous dynamic changes, and ensure the stability and accuracy of the background model over a long period of time; through the setting of the gray difference model and the threshold P, the foreground area is clearly distinguished from the background area, the foreground area of the dynamic target is accurately extracted, and the influence of background interference on the extraction of the moving target is reduced.
[0083] In a preferred embodiment of the present invention, in the above-mentioned method for motion recognition and imitation of a humanoid robot, a smoothing filter (such as a Gaussian filter) is applied to the accumulated foreground area for smoothing.
[0084] The technical effect is that by applying a smoothing filter to smooth the accumulated foreground area, the image can be effectively smoothed, the noise introduced by the camera equipment, environmental changes or segmentation algorithms can be reduced, the smoothness of the target edge can be enhanced, the continuity of the target area can be improved, the target area can be made more coherent in space, the stability of feature extraction can be enhanced, the foreground area after smoothing has fewer details and a more uniform distribution, and the computational complexity of the system can be reduced.
[0085] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the separation of the moving target from the static scene view includes: calculating the grayscale difference between the current image frame and the cumulative foreground area from the static scene view; dynamically adjusting the threshold according to the statistical characteristics of the grayscale difference map, and generating a moving target foreground mask according to the threshold; using the moving target foreground mask to extract the moving target from the current image frame.
[0086] Its technical effects are: by calculating the grayscale difference between the current image frame and the accumulated foreground area, the difference between the moving target and the background can be more accurately identified; the threshold is dynamically adjusted by using the statistical characteristics of the grayscale difference map to adapt to different scenes and changes in lighting conditions, and the accuracy of segmentation is improved by dynamically adjusting the threshold.
[0087] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the use of a tracking algorithm to track the motion trajectory of the moving target includes: according to the initial area of the moving target, setting a tracking window with a width of Height is Wherein, (x0, y0) is the initial center coordinate of the moving target, W0 is the initial width of the tracking window, and H0 is the initial height of the tracking window; feature extraction is performed on each pixel (x, y) in the current tracking window, and the gray value I is calculated. t (x, y), calculate the relative feature weight of each pixel in the current tracking window Generate a weight distribution graph for representing the weighted feature distribution of each pixel in the current tracking window, using the weight distribution model W dist (x,y)=w(x,y)×I t (x, y) represents; the weighted center position (x) of the current window is calculated according to the weight distribution model c ,y c ),in, According to the concentration of weight distribution, adjust the current tracking window, the width is W t+1 =α·W t +(1-α)·Var(W dist ), height is H t+1 =α·H t +(1-α)·Var(H dist ), where α is the smoothing factor, Var(W dist ) and Var(H dist ) is the variance of the weight distribution map in the x and y directions; according to the position of the weighted centroid, the tracking window in each frame is updated, and the position is (x t+1 =x c ,y t+1 =y c ), with a width of W t+1 , height H t+1 ; Repeat tracking the tracking window for each frame of the original motion video until the original motion video ends or the moving target is lost.
[0088] Its technical effects are: by setting the tracking window and calculating the weighted feature distribution of pixels in each frame, the position and shape changes of the moving target can be accurately tracked, the continuity of the target between consecutive frames can be maintained, and the possibility of target loss can be reduced; the size of the tracking window can be dynamically adjusted according to the concentration of the weight distribution to adapt to changes in the target size. When the target moves, it may be enlarged or reduced. By adjusting the width and height of the tracking window, the actual size of the target can be better captured, ensuring that the tracking window always covers the target, and effectively preventing the target from leaving the tracking range.
[0089] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the capturing of the key motion frame of the moving target, extracting the fusion motion feature of the moving target, and obtaining the fusion motion feature vector of the moving target comprises: calculating the difference between each frame and the previous frame in the video of the motion trajectory of the moving target, and when the difference exceeds a set threshold, identifying the frame as a key motion frame; for each of the key motion frames, extracting the geometric morphology feature vector G j =[R j ,A r ,P r ,SM t ], the geometric morphological feature vector, where R j is the aspect ratio of the target area, A r is the ratio of the area to the area of the circumscribed rectangle, P r is the ratio of the perimeter to the perimeter of the circumscribed rectangle, SM t is the shape moment parameter, and the calculation formula is For each of the key action frames, the global invariant feature vector H is calculated by global invariant moment. j ; Merge the geometric morphology feature vector and the global invariant feature vector to obtain a fused motion feature vector V j =[G j ,H j ].
[0090] Specifically, the global invariant eigenvector H is calculated j The method comprises: for a two-dimensional image of a key action frame, the pixel value in the two-dimensional image is represented as I(x, y); the p-order moment of the two-dimensional image Wherein p and q are non-negative integers; the centering moment of the two-dimensional image in, and are the center coordinates of the two-dimensional image, Calculate the invariant moment characteristic components at specific orders p and q to obtain the invariant characteristic function Using the combination of different orders p and q, the values of (p,q) are selected as (2,0), (0,2), (1,1), and the global invariant eigenvector is calculated.
[0091] Its technical effects are: by calculating the difference between each frame and the previous frame and setting a threshold, important action change moments in the video can be effectively identified, and the captured key action frames can represent the significant action features of the target; by fusing geometric morphological features and global invariant features, the action features of the target are comprehensively described from different angles. The geometric morphological features reflect the changes in the target's morphology and geometric structure, while the global invariant features provide a description of shape invariance, rotation and scaling invariance. By merging these two types of features, the target's action state and change trend can be more accurately portrayed, forming a comprehensive description of the action.
[0092] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the key action frames are labeled with action labels, and the motion target fused action feature vectors and the corresponding action labels are sorted to obtain a training data set, including: defining the types of action labels, including at least one of walking, jumping, and waving; combining the fused action feature vectors of the motion target on each of the key action frames with the corresponding action labels to obtain a training data set.
[0093] Specifically, the key action frame F j The action label is given by the formula T j =label(v j ) determine; the obtained training data set D = {(v1, T1), (v2, T2), ..., (v n ,T n )}. Before arranging the training data set, normalization processing can be performed, which will not be repeated here. The final training sample will be composed of the pre-processed fusion action feature vector and the corresponding action label.
[0094] Its technical effect is: by annotating key action frames with action labels and organizing them with fused action feature vectors, a high-quality training data set is obtained, which helps to improve the classification accuracy of the model, enhance the representativeness of the training data set, and optimize the learning effect of the model.
[0095] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the training data set is trained by using the target fusion motion feature vector, and the motion description of the moving target is obtained by combining static recognition and dynamic recognition, which includes: for the geometric morphology feature vector, a geometric morphology linear classification output f is obtained after being processed by a linear classifier G (G j )=W G ·G j +b G , where W Gis the weight matrix of the geometric morphology feature vector, with a dimension of C×d G , d G is the geometric morphology feature vector G j The dimension of C is the number of categories, b is G is the bias term of the geometric morphological feature; for the global invariant feature vector, the global invariant linear classification output f is obtained after being processed by the linear classifier H (H j )=W H ·H j +b H , where W H is the weight matrix of the global invariant feature vector, with dimension C×d H , d H is the global invariant eigenvector H j The dimension of C is the number of categories, b is H is the bias term of the geometric morphological feature; the geometric morphological linear classification output and the global invariant linear classification output are fused by weighted averaging to obtain a classification probability distribution model for static recognition Among them, α is the weight parameter of the geometric morphological feature vector, β is the weight parameter of the global invariant feature vector, α+β=1; a 3D convolutional spatiotemporal network model is established, and the 3D convolutional spatiotemporal network model includes an input layer, a downsampling layer, a residual layer and an output layer; the input layer uses a 3D convolution kernel, the size of the 3D convolution kernel is 3×3×4, the step size is 1×2×2, and the input layer outputs a feature map of size N×T×C1×H / 4×W / 4 as the input of the residual layer, wherein N is the batch size, T is the number of frames, C1 is the number of channels, H is the height, and W is the width; the downsampling layer uses a 3D convolution kernel, the size of the 3D convolution kernel is 1×2×2, the step size is 1×2×2, and each of the residual layer The convolution blocks include a 3×3×3 convolution kernel layer, a 3×7×7 convolution kernel layer, an MLP layer, a normalization layer and a GELU activation function layer, respectively, to obtain the output feature map y=x+MLP(GELU(MLP(LN(conv3(x)+conv7(x)+conv3x3x3(x))))), where x is the input of each convolution block; the output layer uses a global average pooling layer and a fully connected layer for final classification, the feature map size of the global average pooling layer output is N×512, and the output of the fully connected layer is N×C, which is the final output of the 3D convolutional spatiotemporal network model; the 3D convolutional spatiotemporal network model is trained to obtain a classification probability distribution model for dynamic recognition. Among them, W d is the weight matrix of the fully connected layer, b dis the bias vector of the fully connected layer; the classification probability distribution model of the static recognition and the classification probability distribution model of the dynamic recognition are weightedly fused to obtain a fusion model Among them, γ s is the weight parameter of the static recognition result, γ d is the weight parameter of the dynamic recognition result, γ s +γ d =1; select the category with the highest probability The final action classification result for the action description of the moving target is obtained.
[0096] Specifically, training the 3D convolutional spatiotemporal network model includes: decomposing the original motion video into image frames, obtaining T frame images as input, downsampling the input images, and adjusting the resolution to 224×224; initializing the network parameters using a random initialization strategy, and the size of the input data is N×T×3×224×224, where N is the batch size and T is the number of clips; using the exponential moving average method (EMA) during training to reduce network overfitting, and selecting the model with the highest verification accuracy in the EMA model as the final model, and finally obtaining a classification probability distribution model for dynamic recognition.
[0097] The technical effects are as follows: the geometric morphology feature vector and the global invariant feature vector are processed by linear classifiers respectively, so that the static recognition model can accurately extract key information from each feature vector and classify it. By weighted fusion of the geometric morphology linear classification output and the global invariant linear classification output, static recognition and dynamic recognition are combined to effectively enhance the classification accuracy. The 3D convolutional spatiotemporal network model is used to process dynamic video data to capture spatiotemporal information, thereby improving the recognition ability of moving target movements. The downsampling layer and the residual layer in the 3D convolutional spatiotemporal network model reduce the computational complexity and extract higher-level features, thereby enhancing the model's ability to process input data and thus improving the dynamic recognition ability of moving targets. When training the 3D convolutional spatiotemporal network model, a random initialization strategy and an exponential moving average method (EMA) are used to reduce network overfitting, thereby ensuring the stability of the training process and the generalization ability of the model.
[0098] In a preferred embodiment of the present invention, in the above-mentioned humanoid robot motion recognition and imitation method, the sending of the action description to the humanoid robot for motion imitation includes: converting the classification results of the static recognition and the dynamic recognition into a standard format action description, the action description can be a string or an instruction in a specific format, such as JSON or XML; converting the standard format action description into a humanoid robot action instruction, if the humanoid robot supports a specific control protocol, then converting the action description into the message format of the protocol, and sending it to the humanoid robot through a communication interface, wherein common communication interfaces include serial communication, network communication and middleware interface; the humanoid robot receives the instruction, parses, executes the corresponding action and feedbacks the action execution status, and sends or feeds back the action execution status to the control system to ensure the correctness of the action execution.
[0099] Its technical effect is: through standardized action description, adaptive control protocol, efficient communication and feedback mechanism, the accuracy of action simulation, system compatibility and user experience are improved.
[0100] The second embodiment of the present invention provides a humanoid robot motion recognition and imitation system, which includes: a static scene view extraction module, which is used to extract a static scene view containing a moving target to be identified from an original motion video using a grayscale difference background segmentation method; a moving target separation module, which is used to separate the moving target from the static scene view; a motion trajectory tracking module, which is used to track the motion trajectory of the moving target using a tracking algorithm; a feature vector extraction module, which is used to capture the key action frames of the moving target, extract fused action features of the moving target, and obtain a fused action feature vector of the moving target; a label annotation module, which is used to annotate the key action frames with action labels, and organize the fused action feature vector of the moving target and the corresponding action labels to obtain a training data set; a 3D convolutional spatiotemporal network model establishment module, which is used to train the training data set using the target fused action feature vector, and combine static recognition and dynamic recognition to obtain a description of the action of the moving target; and an action execution module, which is used to send the action description to the humanoid robot for action imitation.
[0101] The computer program product of the humanoid robot motion recognition and imitation method and device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. The specific implementation can be found in the method embodiment, which will not be repeated here.
[0102] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned humanoid robot motion recognition and imitation method, thereby achieving more accurate motion description and more efficient motion imitation by integrating static recognition and dynamic recognition.
[0103] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0104] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for motion recognition and imitation of a humanoid robot, characterized in that: include: The grayscale difference background segmentation method is used to extract the static scene view containing the moving target to be identified from the original motion video; Separating a moving target from the static scene view; Tracking the motion trajectory of the moving target using a tracking algorithm; Capturing the key action frames of the moving target, extracting fused action features of the moving target, and obtaining a fused action feature vector of the moving target; Annotating the key action frames with action labels, and arranging the motion target fusion action feature vectors and the corresponding action labels to obtain a training data set; The target fusion action feature vector is used to train the training data set, and static recognition and dynamic recognition are combined to obtain the action description of the moving target; sending the action description to a humanoid robot for action imitation; The capturing of the key action frame of the moving target, extracting the fused action feature of the moving target, and obtaining the fused action feature vector of the moving target comprises: In the video of the motion trajectory of the moving target, the difference between each frame and the previous frame is calculated, and when the difference exceeds a set threshold, the frame is identified as a key action frame; For each of the key action frames, extract the geometric feature vector G j =[R j ,A r ,P r ,SM t ], the geometric morphological feature vector, where R j is the aspect ratio of the target area, A r is the ratio of the area to the area of the circumscribed rectangle, P r is the ratio of the perimeter to the perimeter of the circumscribed rectangle, SM t is the shape moment parameter, and the calculation formula is For each of the key action frames, the global invariant feature vector H is calculated by global invariant moment. j ; The geometric feature vector and the global invariant feature vector are combined to obtain a fused motion feature vector V j =[G j ,H j ]; The training data set is trained by using the target fusion action feature vector, and static recognition and dynamic recognition are combined to obtain the action description of the moving target, which includes: The geometric morphology feature vector is processed by a linear classifier to obtain a geometric morphology linear classification output f G (G j )=W G ·G j +b G , where W G is the weight matrix of the geometric morphology feature vector, with a dimension of C×d G , d G is the geometric morphology feature vector G j The dimension of C is the number of categories, b is G is the bias term of geometric features; The globally invariant feature vector is processed by a linear classifier to obtain a globally invariant linear classification output f H (H j )=W H ·H j +b H , where W H is the weight matrix of the global invariant feature vector, with dimension C×d H , d H is the global invariant eigenvector H j The dimension of C is the number of categories, b is H is the bias term of geometric features; The geometric linear classification output and the global invariant linear classification output are fused by weighted averaging to obtain a classification probability distribution model for static recognition. Wherein, α is the weight parameter of the geometric morphological feature vector, β is the weight parameter of the global invariant feature vector, and α+β=1; Establishing a 3D convolutional spatiotemporal network model, wherein the 3D convolutional spatiotemporal network model includes an input layer, a downsampling layer, a residual layer, and an output layer; The input layer uses a 3D convolution kernel with a size of 3×3×4 and a step size of 1×2×2. The input layer outputs a feature map with a size of N×T×C1×H / 4×W / 4 as the input of the residual layer, where N is the batch size, T is the number of frames, C1 is the number of channels, H is the height, and W is the width; The downsampling layer uses a 3D convolution kernel with a size of 1×2×2 and a step size of 1×2×2. Each convolution block of the residual layer includes a 3×3×3 convolution kernel layer, a 3×7×7 convolution kernel layer, an MLP layer, a normalization layer and a GELU activation function layer, respectively, to obtain an output feature map y=x+MLP(GELU(MLP(LN(conv3(x)+conv7(x)+conv3x3x3(x))))) of the residual layer, wherein x is the input of each convolution block; The output layer uses a global average pooling layer and a fully connected layer for final classification, the feature map size of the global average pooling layer output is N×512, and the output of the fully connected layer is N×C, which is the final output of the 3D convolutional spatiotemporal network model; The 3D convolutional spatiotemporal network model is trained to obtain a dynamic recognition classification probability distribution model Among them, W d is the weight matrix of the fully connected layer, b d is the bias vector of the fully connected layer; The classification probability distribution model of the static recognition and the classification probability distribution model of the dynamic recognition are weightedly fused to obtain a fusion model Among them, γ s is the weight parameter of the static recognition result, γ d is the weight parameter of the dynamic recognition result, γ s +γ d =1; Select the category with the highest probability The final action classification result for the action description of the moving target is obtained.
2. The method for motion recognition and imitation of a humanoid robot according to claim 1, characterized in that: The method of extracting a static scene view containing a moving target to be identified from an original moving video by using a grayscale difference background segmentation method includes: Extract several initial image frames from the original motion video and calculate the average grayscale value of these frames as the initial background model Among them, G t (x, y) is the grayscale value of the tth frame, (x, y) is the pixel coordinate in each of the image frames, and N is the number of the extracted image frames; The initial background model is updated using dynamic change smoothing to obtain an updated background model B t+1 (x,y)=(1-α)×B t (x,y)+α×G t+1 (x, y), where α is the smoothing factor, 0<α<1; For each frame, the grayscale difference between the current image frame and the initial background model is calculated to obtain the grayscale difference model ΔG t (x,y)=|G t (x,y)-B t (x,y)|; For the grayscale difference model ΔG t (x, y), distinguish the foreground area from the background area by setting the threshold P, the scene foreground mask The value of the scene foreground mask being 1 represents the foreground area, and the value of the scene foreground mask being 0 represents the background area; Using the scene foreground mask M t (x,y) extracts the foreground region F from the current frame t (x,y)=G t (x,y)×M t (x,y); The average value of the cumulatively extracted foreground areas in all the initial image frames is calculated to obtain a cumulative foreground area for representing a static scene view containing a moving target to be identified. Wherein, T is the total number of frames used to extract static views.
3. The method for motion recognition and imitation of a humanoid robot according to claim 2, characterized in that: For the accumulated foreground area, a smoothing filter is applied to perform smoothing processing.
4. The method for motion recognition and imitation of a humanoid robot according to claim 2, characterized in that: The step of separating the moving target from the static scene view comprises: From the static scene view, calculating the grayscale difference between the current image frame and the accumulated foreground area; dynamically adjusting a threshold value according to the statistical characteristics of the grayscale difference map, and generating a moving target foreground mask according to the threshold value; The moving object foreground mask is used to extract the moving object from the current image frame.
5. The method for motion recognition and imitation of a humanoid robot according to claim 1, characterized in that: Tracking the motion trajectory of the moving target using a tracking algorithm includes: According to the initial area of the moving target, set the tracking window with a width of Height is Wherein, (x0, y0) is the initial center coordinate of the moving target, W0 is the initial width of the tracking window, and H0 is the initial height of the tracking window; Perform feature extraction on each pixel (x, y) in the current tracking window and calculate the gray value I t (x, y), calculate the relative feature weight of each pixel in the current tracking window Generate a weight distribution graph for representing the weighted feature distribution of each pixel in the current tracking window, using the weight distribution model W dist (x,y)=w(x,y)×I t (x,y) represents; The weighted center position (x c ,y c ),in, According to the concentration of weight distribution, adjust the current tracking window, the width is W t+1 =α·W t +(1-α)·Var(W dist ), height is H t+1 =α·H t +(1-α)·Var(H dist ), where α is the smoothing factor, Var(W dist ) and Var(H dist ) is the variance of the weight distribution graph in the x and y directions; According to the position of the weighted center, the tracking window in each frame is updated, and the position is (x t+1 =x c ,y t+1 =y c ), with a width of W t+1 , height H t+1 ; The tracking window is repeatedly tracked for each frame of the original motion video until the original motion video ends or the moving target is lost.
6. The method for motion recognition and imitation of a humanoid robot according to claim 1, characterized in that: The step of labeling the key action frames with action labels and arranging the motion target fusion action feature vectors and the corresponding action labels to obtain a training data set includes: Define the type of action tag, including at least one of walking, jumping, and waving; The fused action feature vector of the moving target in each of the key action frames is combined with the corresponding action label to obtain a training data set.
7. The method for motion recognition and imitation of a humanoid robot according to claim 1, characterized in that: The sending the action description to the humanoid robot for action imitation comprises: Converting the classification results of the static recognition and the dynamic recognition into action descriptions in a standard format; Converting the action description in a standard format into an action instruction of a humanoid robot, and sending the instruction to the humanoid robot through a communication interface; The humanoid robot receives instructions, analyzes and executes corresponding actions and provides feedback on the execution status of the actions.
8. A humanoid robot motion recognition and imitation system, characterized in that: include: A static scene view extraction module is used to extract a static scene view containing a moving target to be identified from the original moving video by using a grayscale difference background segmentation method; A moving target separation module, used to separate the moving target from the static scene view; A motion trajectory tracking module, used for tracking the motion trajectory of the moving target using a tracking algorithm; A feature vector extraction module is used to capture the key action frame of the moving target, extract the fused action feature of the moving target, and obtain the fused action feature vector of the moving target; A label marking module is used to mark the key action frames with action labels, and to sort the motion target fusion action feature vectors and the corresponding action labels to obtain a training data set; A 3D convolutional spatiotemporal network model building module, used to train the training data set using the target fusion action feature vector, and combine static recognition and dynamic recognition to obtain the action description of the moving target; An action execution module, used for sending the action description to the humanoid robot for action imitation; The operations performed by the feature vector extraction module include: In the video of the motion trajectory of the moving target, the difference between each frame and the previous frame is calculated, and when the difference exceeds a set threshold, the frame is identified as a key action frame; For each of the key action frames, extract the geometric feature vector G j =[R j ,A r ,P r ,SM t ], the geometric morphological feature vector, where R j is the aspect ratio of the target area, A r is the ratio of the area to the area of the circumscribed rectangle, P r is the ratio of the perimeter to the perimeter of the circumscribed rectangle, SM t is the shape moment parameter, and the calculation formula is For each of the key action frames, the global invariant feature vector H is calculated by global invariant moment. j ; The geometric feature vector and the global invariant feature vector are combined to obtain a fused motion feature vector V j =[G j ,H j ]; The operations performed by the 3D convolutional spatiotemporal network model building module include: The geometric morphology feature vector is processed by a linear classifier to obtain a geometric morphology linear classification output f G (G j )=W G ·G j +b G , where W G is the weight matrix of the geometric morphology feature vector, with a dimension of C×d G , d G is the geometric morphology feature vector G j The dimension of C is the number of categories, b is G is the bias term of geometric features; The globally invariant feature vector is processed by a linear classifier to obtain a globally invariant linear classification output f H (H j )=W H ·H j +b H , where W H is the weight matrix of the global invariant feature vector, with dimension C×d H , d H is the global invariant eigenvector H j The dimension of C is the number of categories, b is H is the bias term of geometric features; The geometric linear classification output and the global invariant linear classification output are fused by weighted averaging to obtain a classification probability distribution model for static recognition. Wherein, α is the weight parameter of the geometric morphological feature vector, β is the weight parameter of the global invariant feature vector, and α+β=1; Establishing a 3D convolutional spatiotemporal network model, wherein the 3D convolutional spatiotemporal network model includes an input layer, a downsampling layer, a residual layer, and an output layer; The input layer uses a 3D convolution kernel with a size of 3×3×4 and a step size of 1×2×2. The input layer outputs a feature map with a size of N×T×C1×H / 4×W / 4 as the input of the residual layer, where N is the batch size, T is the number of frames, C1 is the number of channels, H is the height, and W is the width; The downsampling layer uses a 3D convolution kernel with a size of 1×2×2 and a step size of 1×2×2. Each convolution block of the residual layer includes a 3×3×3 convolution kernel layer, a 3×7×7 convolution kernel layer, an MLP layer, a normalization layer and a GELU activation function layer, respectively, to obtain an output feature map y=x+MLP(GELU(MLP(LN(conv3(x)+conv7(x)+conv3x3x3(x))))) of the residual layer, wherein x is the input of each convolution block; The output layer uses a global average pooling layer and a fully connected layer for final classification, the feature map size of the global average pooling layer output is N×512, and the output of the fully connected layer is N×C, which is the final output of the 3D convolutional spatiotemporal network model; The 3D convolutional spatiotemporal network model is trained to obtain a dynamic recognition classification probability distribution model Among them, W d is the weight matrix of the fully connected layer, b d is the bias vector of the fully connected layer; The classification probability distribution model of the static recognition and the classification probability distribution model of the dynamic recognition are weightedly fused to obtain a fusion model Among them, γ s is the weight parameter of the static recognition result, γ d is the weight parameter of the dynamic recognition result, γ s +γ d =1; Select the category with the highest probability The final action classification result for the action description of the moving target is obtained.