A 3D posture real-time recognition method, storage medium and device

By collecting and processing the pose image and position information of the target object in 3D pose recognition, and using the FasterNet neural network model for processing, the balance problem of real-time and accuracy of 3D gesture estimation in the prior art is solved, and real-time and high-precision 3D pose recognition on mobile terminals and embedded devices is realized.

CN118587285BActive Publication Date: 2025-05-16GUANGZHOU ZIWEIYUN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410739489.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-05-16
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

The existing 3D gesture estimation technology is difficult to balance in real-time and accuracy, especially on mobile and embedded devices, real-time detection cannot be achieved. The top-down method discards the position information of the target object, affecting the model's accurate judgment of the 3D pose.

Method used

By collecting the pose image of the target object, obtaining its position information and cropping, inputting cropping diagrams and position information into the neural network model for processing, obtaining feature vectors and performing key point detection, and FasterNet is used as the backbone network to improve recognition speed and accuracy.

Benefits of technology

The accuracy of 3D posture detection is improved, allowing the model to accurately predict the depth changes of the target object at different positions under the original camera coordinate system, and is suitable for mobile terminals and embedded devices, real-time recognition is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118587285B_ABST
    Figure CN118587285B_ABST
Patent Text Reader

Abstract

The present invention relates to a 3D gesture real-time recognition method, storage medium and device, which are used to solve the problem that the top-down gesture estimation method discards the position information of the target object, so that the model cannot accurately predict the subtle changes in depth at different positions in the original camera coordinate system, resulting in low detection accuracy, including: collecting a gesture image of the target object; obtaining the position information of the target object from the gesture image of the target object and cropping the gesture image of the target object to obtain a cropped image of the target object; inputting the cropped image and the position information into a neural network model for processing to obtain a feature vector; performing key point detection on the feature vector to obtain the key point coordinate information of the target object. The present invention provides global information for the model through the cropped image and position information, so that the model can accurately predict the subtle changes in the depth of the target object at different positions in the original camera coordinate system, thereby improving the accuracy of 3D gesture detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and more specifically, to a 3D gesture real-time recognition method, storage medium and device. Background Art

[0002] 3D hand gesture recognition is a technology used to detect and understand the movements and postures of human hands in three-dimensional space. It integrates 3D sensors, camera technology and other sensing devices with computer vision and machine learning algorithms and is suitable for a variety of application fields, such as virtual reality, augmented reality, gesture control, medical image processing, game development, security authentication and other fields.

[0003] Although the academic community has made significant progress in the field of 3D gesture estimation and has developed many high-precision algorithms, the optimization of these algorithms requires the use of high-computing hardware resources, and cannot meet the needs of real-time detection when actually applied to mobile terminals and embedded devices. Therefore, the 3D gesture estimation technology based on these algorithms is in an unbalanced state in terms of real-time performance and accuracy. On the other hand, the position of the gesture relative to the camera will also affect the accuracy of the model in 3D gesture recognition. These problems will lead to various problems in the use of the product and lack of generalization.

[0004] The current mainstream high-precision 3D hand posture estimation technology solution is to use the MANO model proposed by the Max Planck Institute in 2017. MANO has a very reasonable structure and a well-defined forward dynamics tree. MANO can model the anatomical structure of the human hand in detail, so the recognition accuracy of this model is very high. However, this model is very resource-intensive and is generally applicable to GPU environments. It cannot achieve real-time detection on the embedded side, so related product solutions are difficult to implement.

[0005] On the other hand, although the top-down real-time posture estimation strategy has alleviated resource pressure to a certain extent and achieved a balance between speed and accuracy, it still has limitations. This method usually includes two stages: first detecting the hand object and then performing posture estimation. Among them, the Heatmap-based method is intuitive and easy to understand, but in order to improve the accuracy of key point positioning, complex upsampling technology is required, which undoubtedly increases the computational burden. In contrast, although the direct regression method is more efficient and can achieve fast detection under limited resources, the detection effect of this method may decline in complex scenes, such as abnormal scenes with large lighting changes and occlusion.

[0006] The top-down approach can separate gesture posture recognition from detection, but the position information is discarded at the beginning of the first stage of detection and cropping, making it impossible for the model to accurately predict the subtle changes in depth caused by different positions in the original camera coordinate system, thereby affecting the precise judgment of gesture depth and camera orientation. Summary of the invention

[0007] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art and provide a real-time 3D posture recognition method to solve the problem that the top-down posture estimation method discards the position information of the target object, making the model unable to accurately predict the subtle changes in depth caused by different positions in the original camera coordinate system, resulting in reduced accuracy of 3D posture detection and making it impossible to determine whether the predicted target object is facing the camera.

[0008] The technical solution adopted by the present invention is a 3D posture real-time recognition method, comprising:

[0009] Collecting posture images of target objects;

[0010] Acquiring position information of the target object from the target object posture image and cropping the target object posture image to obtain a cropped image of the target object;

[0011] Inputting the cropped image and the position information into a neural network model for processing to obtain a feature vector;

[0012] Key point detection is performed on the feature vector to obtain key point coordinate information of the target object.

[0013] The present invention provides global information for the model by inputting a cropped image of the target object's posture image and position information into a neural network model for processing, thereby improving the accuracy of the model's estimation of the target object's posture, having low hardware requirements, and being suitable for mobile terminals or embedded devices.

[0014] Furthermore, the neural network model uses FasterNet as the backbone network.

[0015] FasterNet runs very fast and is suitable for use in mobile and embedded devices, improving the recognition speed of the model.

[0016] Furthermore, the FasterNet includes several levels, each level is composed of several residual module structures, and a convolution layer Conv with a convolution kernel of 3x3 is embedded in front of each level; the several levels are connected with a pooling layer, and the pooling layer is connected with a fully connected layer FC1, and the fully connected layer FC1 is respectively connected with 3 parallel fully connected layers.

[0017] Furthermore, there are 4 levels, among which the first level stage1 has at least 1 residual module structure, the number of residual module structures of the second level stage2 is twice that of the first level, the number of residual module structures of the third level stage3 is 4 times that of the second level, and the number of residual module structures of the fourth level stage4 is the same as the number of residual module structures of the second level.

[0018] Furthermore, the target object is a gesture, and the method also includes a model loss function, which is used to constrain the learning of the neural network, and the calculation of the model loss function is based on a 2D key point loss function, a 3D key point loss function and a bone length loss function.

[0019] Furthermore, the key point coordinate information of the target object includes predicted 2D key points; the function calculation formula of the 2D key point loss is Loss 2d =|y pt -y gt |, where y pt is the predicted 2D key point, y gt are real 2D key points.

[0020] Furthermore, the key point coordinate information of the target object also includes the predicted 3D key point depth, and the predicted 2D key point and the predicted 3D key point depth are converted to obtain the predicted 3D key point; the loss function formula of the 3D key point is: Loss 3d =|y pt_3d -y gt_3d |, where y pt_3d is the predicted 3D key point, y gt_3d is the real 3D key point.

[0021] Furthermore, the formula of the bone length loss function is: Where ε is the number of hand bones, and are the i-th and j-th predicted 3D key points respectively, and are the i-th and j-th real 3D key points respectively.

[0022] On the other hand, the present invention provides a storage medium having a program stored thereon, wherein the program, when executed, implements the above-mentioned 3D gesture real-time recognition method.

[0023] On the other hand, the present invention provides a 3D gesture real-time recognition device, the device comprising:

[0024] An image and position information acquisition module is used to acquire the posture image of the target object;

[0025] An image preprocessing module, used for acquiring the position information of the target object from the target object posture image and cropping the target object posture image to obtain a cropped image of the target object;

[0026] An image processing module, used for inputting the cropped image and the position information into a neural network model for processing to obtain a feature vector;

[0027] The key point detection module is used to perform key point detection on the feature vector to obtain key point coordinate information of the target object.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] (1) By inputting the cropped image and position information of the target object into the neural network model for processing, the model is provided with global information, so that the model can accurately predict the subtle changes in depth caused by the different positions of the target object in the original camera coordinate system, thereby improving the accuracy of 3D pose detection.

[0030] (2) The neural network model uses FasterNet as the backbone network, which can make the model's reasoning more suitable for mobile terminals and embedded devices, thereby achieving the effect of improving the recognition speed.

[0031] (3) The 2D key point information and depth information are detected through the key point detection module, and the learning of the model is constrained by the specified loss function, thereby improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Schematic diagram of an application scenario of an embodiment of the present invention.

[0033] Figure 2 The present invention is a flowchart of a 3D posture real-time recognition method according to an embodiment of the present invention.

[0034] Figure 3 It is a schematic diagram of the structure of the neural network model backbone network FasterNet according to an embodiment of the present invention.

[0035] Figure 4 Schematic diagram of the structure of the residual module according to an embodiment of the present invention.

[0036] Figure 5 This is a structural diagram of a 3D posture real-time recognition device according to Example 3 of the present invention. DETAILED DESCRIPTION

[0037] The drawings of the present invention are only for illustrative purposes and should not be construed as limiting the present invention. In order to better illustrate the following embodiments, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; it is understandable to those skilled in the art that some well-known structures and their descriptions in the drawings may be omitted.

[0038] Example 1

[0039] This embodiment provides a technical solution that can solve the above-mentioned problem. The specific implementation methods of this application are described in detail below in conjunction with the accompanying drawings.

[0040] For example, Figure 1 A schematic diagram of an application scenario of a 3D posture real-time recognition method provided in an embodiment of the present application. Figure 1 As shown, the application scenario includes at least a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has an image processing function and can also have a data transmission function for video streams and audio streams; the terminal device 200 has a streaming media playback function and can also have an image processing function.

[0041] It is understandable that the server 100 can be an independent electronic device or a cluster of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a car terminal, etc., but is not limited thereto.

[0042] In one practicable manner, the server 100 and the terminal 200 may respectively and separately execute the 3D posture real-time recognition method provided in the embodiment of the present application, or, optionally, the 3D posture real-time recognition method provided in the embodiment of the present application is partially executed in the server 100 and partially executed in the terminal 200.

[0043] like Figure 2 As shown, a 3D gesture real-time recognition method of this embodiment specifically includes the following steps:

[0044] S1. Collect the posture image of the target object.

[0045] In this embodiment, the posture image of the target object is collected, which can be implemented by using any monocular camera or other high-end cameras.

[0046] The target objects of gesture recognition in this embodiment include, but are not limited to, human hand gestures, head gestures, and overall human gestures.

[0047] S2. Acquire the position information of the target object from the posture image of the target object and crop the posture image of the target object to obtain a cropped image of the target object.

[0048] A cropped image is cropped from the collected posture image of the target object according to preset rules, and the cropped image is normalized as the input image of the subsequent model. At the same time, the position information of the cropped image in the original image is collected, and then the position information is normalized to obtain the normalized position information bbox. The normalized cropped image and its position information are used as the input of the subsequent model. The cropped image and its position information can be feature fused in the model, so that the model can learn global information. The representation formula of bbox is as follows:

[0049]

[0050] Among them, (c x ,c y ) is the position information of the cropped image relative to the center of the entire posture image, (b w ,b h ) is the width and height of the target object crop image, and (W,H) is the width and height of the entire pose image.

[0051] S3. Input the cropped image and the position information into a neural network model for processing to obtain a feature vector.

[0052] The neural network model of this embodiment uses FasterNet as a lightweight backbone network, which can make the reasoning speed of the neural network model meet the requirements of real-time reasoning and consume less resources. In specific applications, it can be used Figure 1 The server 100 can implement the method of this embodiment by using Figure 1 The terminal 200 shown implements the method of this embodiment.

[0053] The FasterNet includes several levels, each of which is composed of several residual module structures. Each level embeds a convolution layer Conv with a convolution kernel of 3x3. The several levels are connected with a pooling layer Global Pool. The pooling layer Global Pool is connected with a fully connected layer FC1. The fully connected layer FC1 is respectively connected with three parallel fully connected layers FC2DX, FC2DY, and FC3DZ.

[0054] like Figure 3As shown, according to some embodiments, FasterNet has a total of 4 levels, and each level is composed of several residual module structures FasterNetBlock. Among them, the first level stage1 has at least 1 residual module structure, the second level stage2 has twice the number of residual module structures of the first level, the third level stage3 has 4 times the number of residual module structures of the second level, and the fourth level stage4 has the same number of residual module structures as the second level. For example, in this embodiment, the first level stage1 is composed of 1 FasterNetBlock, the second level stage2 is composed of 2 FasterNetBlocks, the third level stage3 is composed of 8 FasterNetBlocks, and the fourth level stage4 is composed of 2 FasterNetBlocks.

[0055] like Figure 4 As shown in the figure, in FasterNetBlock, the input passes through a PConv convolution with a convolution kernel of 3x3, and then connects two depth-wise separable convolution layers PWConv2d. A normalization layer BN and an activation layer are placed between the two depth-wise separable convolution layers PWConv2d. The output of the last depth-wise separable convolution layer PWConv2d will be weighted fused Add with the input to obtain the output output.

[0056] In FasterNetBlock, the activation layer uses the GELU activation function instead of the original activation function. GELU is a smooth activation function whose output is differentiable over the entire real number range. In addition, the GELU activation function can use the backpropagation algorithm to effectively update the weights, so the GELU activation function is particularly suitable for the training of deep neural networks. Compared with traditional activation functions such as Sigmoid and Tanh, the GELU activation function is less sensitive to the gradient vanishing problem, so it is easier to train in deep networks. In addition, the GELU activation function can help the neural network better capture the nonlinear relationship in the input data, thereby improving the performance of the model. The formula for the GELU activation function is:

[0057]

[0058] Wherein, tanh is the hyperbolic tangent function, π is the ratio of pi, and x is the input value. In this embodiment, the input value x is the output value of the normalization layer.

[0059] After processing through the four layers of the neural network model FasterNet, the feature map Z can be obtained. The feature map Z is globally pooled using the pooling layer Global Pool. Then, the feature vector after global pooling is concatenated with the input position information Cat to obtain a comprehensive feature vector, thereby improving the neural network model's ability to learn image depth information. Next, the comprehensive feature vector is converted and fused through a fully connected layer FC1 to obtain a feature vector Finally, they pass through three fully connected layers FC2DX, FC2DY, and FC3DZ respectively. Specifically, the feature vector Input to the fully connected layer FC2DX, we can get the new feature X, the matrix size of X is (n,b w ) ; eigenvector Input to the fully connected layer FC2DY, we can get the new feature Y, the matrix size of Y is (n,b h ), where n represents the number of key points, (b w ,b h ) are the width and height of the cropped image input to the model.

[0060] S4. Perform key point detection on the feature vector to obtain key point coordinate information of the target object.

[0061] Key point recognition is performed on the new features X and Y to obtain the 3D key point information corresponding to the target object posture. Specifically, FC2DX is used to perform argmax processing on the new feature X to output the x-axis coordinate information of n key points. x FC2DY is used to perform argmax processing on the new feature Y to output the y-axis coordinate information O of n key points y ; The feature vector Z is input to the fully connected layer FC3DZ to directly regress the z-axis depth information and output the z-axis coordinate depth information of n 3D key points Specifically, the i-th 2D keypoint The predicted coordinates are calculated as follows:

[0062]

[0063]

[0064] It can be understood by those skilled in the art that the above argmax processing is to find the index position corresponding to the maximum value. For example, the argmax processing of the new feature X is to find the index position corresponding to the maximum value along the second dimension of the matrix of the new feature X. Direction, for the i-th row of the matrix (i.e. the i-th key point), find the column index of the element with the largest value in the row. The column index of the largest element represents the coordinate position of the i-th key point in the X-axis direction. The process of performing argmax processing on the new feature Y is similar and will not be repeated here.

[0065] The predicted depth coordinate of the i-th 3D keypoint on the z axis As shown below:

[0066]

[0067] In specific applications, when the target object is a gesture, in this embodiment, the total loss function Loss of the key point detection module is all for:

[0068] Loss all =Loss 2d +Loss 3d +Loss bone

[0069] Among them, Loss 2d is the 2D key point loss, Loss 3d is the 3D key point loss, Loss bone Loss of bone length.

[0070] The calculation of 2D key point loss is performed by combining the predicted 2D key points output by the model with the actual 2D key points. The loss function of the 2D key point loss is expressed as:

[0071] Loss 2d =|y pt -y gt |

[0072] Among them, y pt To predict 2D key points, y gt are real 2D key points.

[0073] The 3D keypoint loss is calculated by predicting the 2D keypoint y output by the model pt By predicting the 2D key point y pt The x-axis coordinate value x in the image coordinate system can be obtained 2d and the y-axis coordinate value y 2d , x 2d The x-axis coordinate information of the key point can be x Converted to the coordinates of the target object posture image, y 2d The y-axis coordinate information of the key point can be y The x-axis coordinate value x of the 3D key point in the camera coordinate system can be obtained by converting the image coordinate system to the camera coordinate system. 3d and the y-axis coordinate value y3d , the conversion formula is as follows:

[0074]

[0075] where z 3d is the depth information predicted by the model, (f x ,f y ) is the focal length of the camera. In this embodiment, the camera focal length value is approximated as Where (W, H) is the width and height of the entire posture image, from which the predicted 3D key point y in camera coordinates can be obtained pt_3d and the real 3D key point y gt_3d . Perform L1 loss calculation on 3D key points, 3D key point loss Loss 3d The loss function expression is:

[0076] Loss 3d =|y pt_3d -y gt_3d |

[0077] For the calculation of bone length loss, the predicted 3D key point y can be combined with the number of hand bones ε pt_3d and the real 3D key point y gt_3d Calculate L1 loss and bone length loss bone The loss function is defined as follows:

[0078]

[0079] in and are the i-th and j-th predicted 3D key points respectively, and are the i-th and j-th real 3D key points respectively.

[0080] Bone length loss can constrain the spatial relationship between each key point, which is equivalent to constraining the length of each bone, avoiding predicting some postures that violate biological structure movements and improving the learning of 3D postures.

[0081] Example 2

[0082] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed, a 3D gesture real-time recognition method as described in Embodiment 1 is implemented.

[0083] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or the like.

[0084] Example 3

[0085] This embodiment provides a 3D posture real-time recognition device 300, including an image and position information acquisition module 301, an image preprocessing module 302, an image processing module 303 and a key point detection module 304. Among them:

[0086] The image and position information acquisition module 301 is used to acquire the posture image of the target object;

[0087] The image and position information acquisition module 301 can be used to perform Figure 2 As shown in step S1, for the specific description of the image and position information acquisition module 301, reference may be made to the description of step S1.

[0088] The image preprocessing module 302 is used to obtain the position information of the target object from the target object posture image and to crop the target object posture image to obtain a cropped image of the target object. The image preprocessing module 302 can be used to perform Figure 2 Regarding the step S2, the specific description of the image preprocessing module 302 can refer to the description of the step S2;

[0089] The image processing module 303 is used to input the cropped image and the position information into the neural network model for processing to obtain a feature vector. The image processing module 303 can be used to perform Figure 2 Regarding the step S3, the specific description of the image processing module 303 can refer to the description of the step S3;

[0090] The key point detection module 304 is used to perform key point detection on the feature vector to obtain key point coordinate information of the target object. The key point detection module 304 can be used to perform Figure 2 Regarding step S4, the detailed description of the key point detection module 304 may refer to the description of step S4.

[0091] It can be understood that the above-mentioned device embodiments and the above-mentioned method embodiments can correspond to each other, and similar descriptions of the device embodiments can refer to the method embodiments. To avoid repetition, it will not be repeated here. A 3D posture real-time recognition device provided in an embodiment of the present application can execute a 3D posture real-time recognition method provided in any embodiment of the present application, and has functional modules and beneficial effects corresponding to the execution method. The functional modules of the 3D posture real-time recognition device can be implemented in the form of hardware, can be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules.

[0092] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A 3D gesture real-time recognition method, characterized in that: include: Collecting posture images of target objects; Acquiring position information of the target object from the target object posture image and cropping the target object posture image to obtain a cropped image of the target object; Inputting the cropped image and the position information into a neural network model for processing to obtain a feature vector; The neural network model uses FasterNet as the backbone network; the FasterNet includes several layers, each layer is composed of several residual module structures, and a convolution layer Conv is embedded in front of each layer; the several layers are connected with a pooling layer, and the pooling layer is connected with a fully connected layer FC1, and the fully connected layer FC1 is respectively connected with three parallel fully connected layers FC2DX, FC2DY, and FC3DZ; After processing through the four layers of the neural network model FasterNet, the feature map Z can be obtained, and the feature map Z is globally pooled using the pooling layer; the feature vector after global pooling is concatenated with the input position information to obtain a comprehensive feature vector; The comprehensive feature vector is transformed and fused by the fully connected layer FC1 to obtain the feature vector, and finally passed through three fully connected layers FC2DX, FC2DY, and FC3DZ respectively; Performing key point detection on the feature vector to obtain key point coordinate information of the target object; Key point recognition is performed on the new features X and Y to obtain the 3D key point coordinate information corresponding to the target object posture; specifically, FC2DX is used to perform argmax processing on the new feature X to output the x-axis coordinate information Ox of n key points, and FC2DY is used to perform argmax processing on the new feature Y to output the y-axis coordinate information Oy of n key points; the feature vector Z is input into the fully connected layer FC3DZ to directly regress the z-axis depth information and output the z-axis coordinate depth information of n 3D key points.

2. A 3D gesture real-time recognition method according to claim 1, characterized in that: There are 4 levels, among which the first level stage1 has at least 1 residual module structure, the second level stage2 has twice the number of residual module structures of the first level, the third level stage3 has four times the number of residual module structures of the second level, and the fourth level stage4 has the same number of residual module structures as the second level.

3. A 3D gesture real-time recognition method according to any one of claims 1 or 2, characterized in that: The target object is a gesture, and the method also includes a model loss function, which is used to constrain the learning of the neural network. The calculation of the model loss function is based on a 2D key point loss function, a 3D key point loss function, and a bone length loss function.

4. A 3D gesture real-time recognition method according to claim 3, characterized in that: The key point coordinate information of the target object includes predicted 2D key points; the function calculation formula of the 2D key point loss is Loss 2d =|y pt -y gt |, where y pt is the predicted 2D key point, y gt is the real 2D key point.

5. A 3D gesture real-time recognition method according to claim 4, characterized in that: The key point coordinate information of the target object also includes the predicted 3D key point depth. The predicted 2D key point and the predicted 3D key point depth are converted to obtain the predicted 3D key point. The loss function formula of the 3D key point is: Loss 3d =|y pt _ 3d -y gt _ 3d |, where y pt _ 3d is the predicted 3D key point, y gt _ 3d is the real 3D key point.

6. A 3D gesture real-time recognition method according to claim 5, characterized in that: The formula of the bone length loss function is: Where ε is the number of hand bones, and are the i-th and j-th predicted 3D key points respectively, and are the i-th and j-th real 3D key points respectively.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, a 3D gesture real-time recognition method as described in any one of claims 1 to 6 is implemented.

8. A 3D gesture real-time recognition device, characterized in that: include: An image and position information acquisition module is used to acquire the posture image of the target object; An image preprocessing module, used for acquiring the position information of the target object from the target object posture image and cropping the target object posture image to obtain a cropped image of the target object; An image processing module, used for inputting the cropped image and the position information into a neural network model for processing to obtain a feature vector; The neural network model uses FasterNet as the backbone network; the FasterNet includes several layers, each layer is composed of several residual module structures, and a convolution layer Conv is embedded in front of each layer; the several layers are connected with a pooling layer, and the pooling layer is connected with a fully connected layer FC1, and the fully connected layer FC1 is respectively connected with three parallel fully connected layers FC2DX, FC2DY, and FC3DZ; After processing through the four layers of the neural network model FasterNet, the feature map Z can be obtained, and the feature map Z is globally pooled using the pooling layer; the feature vector after global pooling is concatenated with the input position information to obtain a comprehensive feature vector; The comprehensive feature vector is transformed and fused by the fully connected layer FC1 to obtain the feature vector, and finally passed through three fully connected layers FC2DX, FC2DY, and FC3DZ respectively; A key point detection module, used to perform key point detection on the feature vector to obtain key point coordinate information of the target object; Key point recognition is performed on the new features X and Y to obtain the 3D key point coordinate information corresponding to the target object posture; specifically, FC2DX is used to perform argmax processing on the new feature X to output the x-axis coordinate information Ox of n key points, and FC2DY is used to perform argmax processing on the new feature Y to output the y-axis coordinate information Oy of n key points; the feature vector Z is input into the fully connected layer FC3DZ to directly regress the z-axis depth information and output the z-axis coordinate depth information of n 3D key points.

Citation Information

Patent Citations

  • Training method of posture recognition model and image recognition method and device

    CN110020633A

  • Double-flow multi-scale hand posture estimation method based on single RGB image

    CN113052030A

  • Lip language recognition method based on partial convolution and multi-scale feature extraction

    CN116978115A