Driving method of 3D gesture model and electronic equipment
Through prediction and feature extraction of hand posture data, combined with a dual-time series model to match the target standard hand posture, the problem of inaccurate pose estimation caused by occlusion in 3D gesture interaction is solved, and the stability and accuracy of the interaction are significantly improved.
Patent Information
- Application Number
- CN202311475427.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-13
AI Technical Summary
During the 3D gesture interaction, the finger part is blocked by both hands or self-blocking with one hand, resulting in low accuracy in hand posture estimation, affecting the stability and accuracy of the interaction.
By using the posture data of each hand joint node in the hand image with the specified number of frames in the hand image of the current frame hand image for prediction, combined with the pre-trained dual-time sequence model for feature extraction and matching, the target standard hand posture is obtained to drive the 3D gesture model.
It significantly improves the stability and accuracy of gesture interaction, reduces the fluctuations in the detection results of hand joint nodes of adjacent frames, and ensures the consistency of the posture timing process in the case of occlusion.
Smart Images

Figure CN119987528A_ABST
Abstract
Description
Background Art
[0002] As people pay more and more attention to the concept of the metaverse, the development of the AR (Augmented Reality) / VR (Virtual Reality) industry is gradually heating up. Among them, 3D (three dimensional) gesture interaction technology has become a key technology for AR / VR device interaction, with a high technical threshold and extremely high importance. The stability and accuracy of gesture interaction are the main factors affecting the gesture interaction experience. The common method in the industry is to build a motion capture system to collect high-precision data sets to improve the accuracy of the algorithm. This method is very costly and has limited effect on the accuracy of gestures with severe occlusion.
[0003] At present, in the actual gesture interaction process, in order to solve the problem that the fingers are often in the visible and invisible state due to occlusion of both hands or self-occlusion of one hand, the posture of the occluded hand is estimated. However, the hand posture estimated under occlusion is usually quite different from the actual hand posture. Therefore, the accuracy of 3D gesture interaction is low. Summary of the invention
[0004] The present application provides a driving method of a 3D gesture model and an electronic device, which are used to improve the accuracy of 3D gesture interaction.
[0005] In a first aspect, an embodiment of the present application provides a method for driving a 3D gesture model, the method comprising:
[0006] Using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image, the posture of each hand joint point in the current frame hand image is predicted to obtain the predicted posture data of each hand joint point in the current frame hand image;
[0007] Based on the predicted posture data of each hand joint point in the hand image of the current frame, a hand timing diagram corresponding to the hand image of the current frame is obtained, wherein the target area of each hand joint point is marked in the hand timing diagram;
[0008] Using a pre-trained dual-time series model, based on a time series feature map obtained by extracting features from a target area in the hand time series map and a hand feature map obtained by extracting features from the hand image of the current frame, intermediate posture data of each hand joint point in the hand image of the current frame are obtained;
[0009] If the intermediate posture data of the designated hand joint points among the hand joint points meet the specified conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain a target standard hand posture that matches the current frame hand image;
[0010] Inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image;
[0011] The preset 3D gesture model is driven by using the target posture data of each hand joint point in the current frame hand image.
[0012] A second aspect of the present application provides an electronic device, including a processor and a memory, wherein the processor and the memory are connected via a bus;
[0013] The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program:
[0014] Using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image, the posture of each hand joint point in the current frame hand image is predicted to obtain the predicted posture data of each hand joint point in the current frame hand image;
[0015] Based on the predicted posture data of each hand joint point in the hand image of the current frame, a hand timing diagram corresponding to the hand image of the current frame is obtained, wherein the target area of each hand joint point is marked in the hand timing diagram;
[0016] Using a pre-trained dual-time series model, based on a time series feature map obtained by extracting features from a target area in the hand time series map and a hand feature map obtained by extracting features from the hand image of the current frame, intermediate posture data of each hand joint point in the hand image of the current frame are obtained;
[0017] If the intermediate posture data of the designated hand joint points among the hand joint points meet the specified conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain a target standard hand posture that matches the current frame hand image;
[0018] Inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image;
[0019] The preset 3D gesture model is driven by using the target posture data of each hand joint point in the current frame hand image.
[0020] According to a third aspect provided by an embodiment of the present invention, a computer storage medium is provided, wherein the computer storage medium stores a computer program, and the computer program is used to execute the method as described in the first aspect.
[0021] In the above-mentioned embodiments of the present application, the posture data of each hand joint point in the hand image of the current frame is used to predict the posture of each hand joint point in the hand image of the current frame, so as to obtain the predicted posture data of each hand joint point in the hand image of the current frame; and based on the predicted posture data of each hand joint point in the hand image of the current frame, a hand timing diagram corresponding to the hand image of the current frame is obtained, and then a pre-trained dual timing model is used to extract features from the target area in the hand timing diagram to obtain a timing feature diagram and a hand feature diagram obtained by extracting features from the hand image of the current frame to obtain the hand joint point in the hand image of the current frame. The intermediate posture data of the node; if the intermediate posture data of the designated hand joint points in the hand joint points meet the specified conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain the target standard hand posture matching the current frame hand image; the hand timing diagram of the current frame hand image and the target standard hand posture is input into the dual timing model to obtain the target posture data of the hand joint points in the current frame hand image, and the preset 3D gesture model is driven by the target posture data of the hand joint points in the current frame hand image. Therefore, in the embodiment of the present application, the predicted posture data of the current frame hand image obtained according to the historical frame prediction is extracted by the dual timing model, and the intermediate posture data of the hand joint points is obtained by combining the feature map obtained by implementing the attention mechanism to guide the model to detect the target area in the input hand timing diagram. Therefore, the fluctuation of the detection results of the hand joint points in adjacent frames is reduced in the embodiment of the present application, and the stability of gesture interaction is significantly improved. In addition, in the embodiment of the present application, by constructing posture data of each preset standard hand posture, when the intermediate posture data output by the detection model has a high degree of similarity with the standard hand posture, the target standard hand posture is obtained, and the hand timing diagram corresponding to the target standard hand posture is used as the data to guide the joint point detection and input into the dual timing model. In this way, the consistency of the posture timing process is guaranteed in the presence of occlusion, thereby improving the accuracy of gesture model driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0023] Figure 1 One of the application scenario schematic diagrams provided by the embodiment of the present application is exemplarily shown;
[0024] Figure 2 The second schematic diagram of the application scenario provided by the embodiment of the present application is exemplarily shown;
[0025] Figure 3 One of the flow charts of the driving method of the 3D gesture model provided in the embodiment of the present application is exemplarily shown;
[0026] Figure 4 A schematic diagram of hand joints provided in an embodiment of the present application is exemplarily shown;
[0027] Figure 5 The schematic diagram of the process of determining the hand timing diagram corresponding to the hand image of the current frame provided by the embodiment of the present application is exemplarily shown;
[0028] Figure 6 A schematic diagram exemplarily shows the target area of the hand joint points provided in the embodiment of the present application;
[0029] Figure 7 A schematic diagram exemplarily shows a hand timing diagram corresponding to a hand image of a current frame provided by an embodiment of the present application;
[0030] Figure 8 The schematic diagram of the structure of the dual timing model provided by the embodiment of the present application is exemplarily shown;
[0031] Fig. 9 An exemplary diagram showing a comparison of visualization effects of time series data provided in an embodiment of the present application is shown;
[0032] Fig.10 The second flowchart of the expression driving method of a 3D digital human provided in an embodiment of the present application is exemplarily shown;
[0033] Fig.11 The schematic diagram of the expression driving device of the 3D digital human provided in the embodiment of the present application is exemplarily shown;
[0034] Fig.12 The hardware structure diagram of the electronic device provided in the embodiment of the present application is exemplified. DETAILED DESCRIPTION
[0035] In order to make the purpose, implementation mode and advantages of the present application clearer, the exemplary implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0036] Based on the exemplary embodiments described in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the claims attached to this application. In addition, although the disclosure in this application is introduced according to one or several exemplary examples, it should be understood that each aspect of the disclosure can also constitute a complete implementation method separately.
[0037] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0038] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a series of components is not necessarily limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0039] The term "module" as used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0040] The following is an overview of the concepts of the embodiments of the present application.
[0041] At present, in the actual gesture interaction process, in order to solve the problem that the fingers are often in the visible and invisible state due to occlusion of both hands or self-occlusion of one hand, the posture of the occluded hand is estimated. However, the hand posture estimated under occlusion is usually quite different from the actual hand posture. Therefore, the accuracy of 3D gesture interaction is low.
[0042] Based on the problem of low accuracy of 3D gesture interaction in the prior art, an embodiment of the present application provides a driving method for a 3D gesture model, which predicts the posture of each hand joint point in the current frame hand image by using the posture data of each hand joint point in the hand image located a specified number of frames before the current frame hand image, and obtains the predicted posture data of each hand joint point in the current frame hand image; and based on the predicted posture data of each hand joint point in the current frame hand image, obtains a hand timing diagram corresponding to the current frame hand image, and then uses a pre-trained dual timing model to extract features from the target area in the hand timing diagram to obtain a timing feature diagram and a hand timing diagram obtained by feature extraction of the current frame hand image. The hand feature map is used to obtain the intermediate posture data of each hand joint point in the hand image of the current frame; if the intermediate posture data of the designated hand joint point among the hand joint points meets the specified conditions, the intermediate posture data of each hand joint point in the hand image of the current frame is matched with the posture data of each preset standard hand posture to obtain the target standard hand posture matching the hand image of the current frame; the hand timing map of the hand image of the current frame and the target standard hand posture is input into the dual timing model to obtain the target posture data of each hand joint point in the hand image of the current frame, and the preset 3D gesture model is driven by the target posture data of each hand joint point in the hand image of the current frame. Therefore, in the embodiment of the present application, the predicted posture data of the hand image of the current frame obtained according to the historical frame prediction is extracted by the dual timing model, and the intermediate posture data of the hand joint point is obtained by combining the feature map obtained by implementing the attention mechanism to guide the model to detect the target area in the input hand timing map. Therefore, the fluctuation of the detection results of the hand joint points in adjacent frames is reduced in the embodiment of the present application, and the stability of gesture interaction is significantly improved. In addition, in the embodiment of the present application, by constructing posture data of each preset standard hand posture, when the intermediate posture data output by the detection model has a high degree of similarity with the standard hand posture, the target standard hand posture is obtained, and the hand timing diagram corresponding to the target standard hand posture is used as the data to guide the joint point detection and input into the dual timing model. In this way, the consistency of the posture timing process is guaranteed in the presence of occlusion, thereby improving the accuracy of gesture model driving.
[0043] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0044] Figure 1 The following is a schematic diagram of an application scenario of a 3D gesture model driven by an embodiment of the present application; Figure 1The application scenario shown is described by taking the electronic device as a server as an example. The application scenario includes a VR device 110 and a server 120. The server 120 can be implemented by a single server or by multiple servers. The server 130 can be implemented by a physical server or by a virtual server.
[0045] In a possible application scenario, the server 120 uses the posture data of each hand joint point in the hand image located a specified number of frames before the current frame hand image to predict the posture of each hand joint point in the current frame hand image, and obtains the predicted posture data of each hand joint point in the current frame hand image; then the server 120 obtains the hand timing diagram corresponding to the current frame hand image based on the predicted posture data of each hand joint point in the current frame hand image, wherein the target area of each hand joint point is marked in the hand timing diagram; then the server 120 uses a pre-trained dual timing model to extract the target area in the hand timing diagram to obtain the timing feature diagram and the hand feature diagram obtained by extracting the features of the current frame hand image. The server 120 obtains the intermediate posture data of each hand joint point in the current frame hand image; if the server 120 determines that the intermediate posture data of the specified hand joint point among the hand joint points meets the specified conditions, the intermediate posture data of each hand joint point in the current frame hand image is matched with the posture data of each preset standard hand posture to obtain the target standard hand posture matching the current frame hand image; and the hand timing diagram of the current frame hand image and the target standard hand posture is input into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image; finally, the server 120 drives the preset 3D gesture model using the target posture data of each hand joint point in the current frame hand image, and displays it in the VR device 110.
[0046] In a possible application scenario, Figure 2As shown, the application scenario includes a VR device 110 and a memory 130. The VR device 110 predicts the posture of each hand joint point in the current frame hand image by using the posture data of each hand joint point in the hand image located in the previous specified number of frames of the current frame hand image, and obtains the predicted posture data of each hand joint point in the current frame hand image; then the VR device 110 obtains the hand timing diagram corresponding to the current frame hand image based on the predicted posture data of each hand joint point in the current frame hand image, wherein the target area of each hand joint point is marked in the hand timing diagram; then the VR device 110 uses the pre-trained dual timing model to extract the timing feature diagram obtained by extracting the target area in the hand timing diagram and the hand feature diagram obtained by extracting the feature of the current frame hand image. Figure, obtain the intermediate posture data of each hand joint point in the current frame hand image; if the intermediate posture data of the specified hand joint point among the hand joint points meets the specified conditions, then the intermediate posture data of each hand joint point in the current frame hand image is matched with the posture data of each standard hand posture preset in the memory 130 to obtain the target standard hand posture matching the current frame hand image; and the hand timing diagram of the current frame hand image and the target standard hand posture is input into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image; finally, the VR device 110 uses the target posture data of each hand joint point in the current frame hand image to drive the preset 3D gesture model.
[0047] Among them, the description in this application only details a single VR device 110, a single server 120, and a single storage 130. However, it should be understood by those skilled in the art that the VR device 110, server 120, and storage 130 shown are intended to represent the operations of the VR device 110, server 120, and storage 130 involved in the technical solution of this application. It does not imply any restrictions on the number, type, or location of the VR device 110, server 120, and storage 130. It should be noted that if additional modules are added to the illustrated environment or individual modules are removed from it, the underlying concepts of the example embodiments of this application will not be changed.
[0048] It should be noted that the driving method of the 3D gesture model proposed in this application is not only applicable to Figure 1 as well as Figure 2 The application scenario shown is also applicable to any drive device with a 3D gesture model.
[0049] The following describes an exemplary driving method of a 3D gesture model in the present application in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the methods and principles of the present application, and the implementation methods of the present application are not limited in this regard.
[0050] like Figure 3 FIG. 1 is a flow chart of a method for driving a 3D gesture model, which may include the following steps:
[0051] Step 301: using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image, predicting the posture of each hand joint point in the current frame hand image, and obtaining the predicted posture data of each hand joint point in the current frame hand image;
[0052] All posture data in the embodiments of the present application are positions, that is, the predicted posture data are predicted positions, the intermediate posture data are intermediate positions, the target posture data are target positions, etc. The implementation of the present application will not be described in detail here.
[0053] like Figure 4 As shown, this is a schematic diagram of the hand joints. Figure 4 The number of hand joints shown in the figure is 21, but the number of hand joints and the corresponding positions are not limited in the embodiments of the present application. The number and positions of hand joints are not limited in the embodiments of the present application.
[0054] Among them, the hand images of the first specified number of frames in the embodiment of the present application are the hand images of the first two frames of the hand image of the current frame.
[0055] In one embodiment, step 301 may be specifically implemented as follows: obtaining the predicted posture data of any hand joint point in the hand image of the current frame by formula (1):
[0056] P t =P t-1 +(P t-1 -P t-2 )+σ......(1);
[0057] Among them, P t is the predicted posture data of any hand joint point in the current frame hand image t, P t-1 is the posture data of any hand joint point in the hand image of frame t-1, P t-2 is the posture data of any hand joint point in the t-2th frame of the hand image, and σ is a preset noise value.
[0058] It should be noted that the noise value in the embodiment of the present application can be set according to actual conditions, and the embodiment of the present application does not limit the specific value of the noise value.
[0059] Step 302: based on the predicted posture data of each hand joint point in the hand image of the current frame, obtaining a hand timing diagram corresponding to the hand image of the current frame, wherein the target area of each hand joint point is marked in the hand timing diagram;
[0060] Next, a method for determining the hand timing diagram corresponding to the hand image of the current frame is described. Figure 5 As shown, it is a schematic diagram of a process of determining a hand timing diagram corresponding to a hand image of a current frame, comprising the following steps:
[0061] Step 501: projecting the predicted posture data of each hand joint point using a preset camera internal parameter to obtain two-dimensional posture data of each hand joint point in the hand image of the current frame;
[0062] In one embodiment, step 501 can be specifically implemented as follows: for any hand joint point, multiply the matrix corresponding to the camera intrinsic parameter by the matrix corresponding to the predicted posture data of the hand joint point to obtain the two-dimensional posture data of the hand joint point in the hand image of the current frame. The two-dimensional posture data of the hand joint point in the hand image of the current frame can be obtained by formula (2):
[0063] p t =AP t ......(2);
[0064] Among them, p t is the two-dimensional posture data of any hand joint point in the hand image t of the current frame, A is the matrix corresponding to the camera intrinsic parameter, P t It is the predicted posture data of any hand joint point in the hand image t of the current frame.
[0065] Step 502: for any hand joint point, according to the two-dimensional posture data of the hand joint point, obtain the intermediate hand timing diagram corresponding to the hand joint point;
[0066] In one embodiment, step 502 can be specifically implemented as follows: based on the two-dimensional posture data of the hand joint points, the positions of the hand joint points in a preset standard timing diagram are obtained, and after the target area of the hand joint points is obtained with the positions of the hand joint points in the standard timing diagram as the center, the pixel values of each pixel point in the target area are set to a first specified pixel value, and the pixel values of other pixel points in the standard timing diagram except the pixel points in the target area are set to a second specified pixel value, so as to obtain an intermediate hand timing diagram of the hand joint points.
[0067] In the embodiment of the present application, the posture data of the hand joint point is the position, and the two-dimensional posture data is the two-dimensional position coordinate. In addition, the target area in the embodiment of the present application is an area of a first specified size. In addition, the standard timing diagram in the embodiment of the present application is an image of a second specified size, and the first specified size is smaller than the second specified size. Figure 6 As shown, the left side is a schematic diagram of the standard timing diagram. Figure 6 It can be seen that if the dark gray area a is the position of any hand joint point, the rectangle formed by the corresponding light gray area and the dark gray area is the target area of the hand joint point.
[0068] It should be noted that the size of the target area in the embodiment of the present application is 7*7, and the size of the standard timing diagram is 32*32. In addition, the first designated pixel value in the embodiment of the present application is 1, and the second designated pixel value is 0. However, the embodiment of the present application does not limit the specific values of the size of the target area, the size of the standard timing diagram, the first designated pixel value, and the second designated pixel value. However, the specific values of the size of the target area, the size of the standard timing diagram, the first designated pixel value, and the second designated pixel value in the embodiment of the present application can be set according to actual conditions.
[0069] Step 503: Obtain the hand timing diagram corresponding to the hand image of the current frame through the intermediate hand timing diagrams corresponding to the hand joint points.
[0070] In one embodiment, the intermediate hand timing diagrams corresponding to the hand joint points are combined into a hand timing diagram corresponding to the hand image of the current frame. Figure 7 As shown, it is a schematic diagram of the hand timing diagram corresponding to the hand image of the current frame, from Figure 7 It can be seen that the hand timing diagram corresponding to the hand image of the current frame is a 21-channel image, wherein the image of each channel corresponds to an intermediate hand timing diagram of a hand joint point.
[0071] Step 303: using a pre-trained dual-time series model to obtain intermediate posture data of each hand joint point in the hand image of the current frame based on a time series feature map obtained by extracting features from a target area in the hand time series map and a hand feature map obtained by extracting features from the hand image of the current frame;
[0072] Before introducing the intermediate posture data of each hand joint point in the current hand image in the embodiment of the present application, first, the dual temporal model in the embodiment of the present application is introduced. In one embodiment, the dual temporal model in the embodiment of the present application includes a temporal backbone network, a first image backbone network, at least one second image backbone network, a third image backbone network and a plurality of fusion modules; the number of the fusion modules is the sum of the number of the first image backbone network and the number of the at least one second image backbone network; wherein:
[0073] The input end of any one of the multiple fusion modules is respectively connected to the timing backbone network and any one of the image backbone networks, and the output end of any one of the fusion modules is connected to the next layer of the image backbone network located at any one of the image backbone networks, wherein the any one of the image backbone networks is one of the first image backbone network, the at least one second image backbone network and the third image backbone network; and the head second image backbone network in the at least one second image backbone network is connected to the first image backbone network, and the tail second image backbone network in the at least one second image backbone network is connected to the third backbone network.
[0074] like Figure 8 As shown, the structural diagram of the dual timing model in the embodiment of the present application is shown. Figure 8 The dual-time series model 800 in the embodiment includes a time series backbone network 801, a first image backbone network 802, a second image backbone network 803, a third image backbone network 804, a fusion module 8051 and a fusion module 8052; the number of the fusion modules is the sum of the number of the first image backbone network and the at least one second image backbone network; wherein:
[0075] The input end of the fusion module 8051 is connected to the time sequence backbone network 801 and the first image backbone network 802, respectively, and the output end of the fusion module 8052 is connected to the second image backbone network 803. The input end of the fusion module 8052 is connected to the time sequence backbone network 801 and the second image backbone network 803, respectively, and the output end of the fusion module 8052 is connected to the third image backbone network 804.
[0076] It should be noted that: Figure 8The example in which the number of the second image backbone network is one and the number of the fusion modules is two is used for explanation. However, the embodiment of the present application does not limit the number of the second image backbone network and the number of the fusion modules. The number of the second image backbone network and the number of the fusion modules in the embodiment of the present application can be set according to the actual situation.
[0077] Next, the method of determining the intermediate posture data of each hand joint point in the current hand image in step 303 is explained. In one embodiment, step 303 can be specifically implemented as follows: using the timing backbone network to extract features of the target area of the hand timing graph corresponding to the current frame hand image to obtain a first timing feature graph; and using the first image backbone network to extract features of the current frame hand image to obtain a first shallow hand feature graph; using the fusion module to fuse the first shallow hand feature with the first timing feature graph to obtain a first intermediate feature graph; for any second image backbone network, The present invention relates to a network, wherein the first target feature map is extracted by the second image backbone network to obtain a first deep hand feature map, wherein the first target feature map is the first intermediate feature map or the second intermediate feature map output by the fusion module connected to the second image backbone network; the first deep hand feature map is fused with the first temporal feature map by the fusion module to obtain a second intermediate feature map; the third image backbone network is used to extract features from the second intermediate feature map output by the fusion module connected to the third image backbone network to obtain intermediate posture data of each hand joint in the hand image of the current frame.
[0078] by Figure 8 The structure of the dual temporal model in is used as an example to illustrate: the temporal backbone network 801 is used to extract features from the target area of the hand timing graph corresponding to the current frame hand image to obtain a first temporal feature graph; and the first image backbone network 802 is used to extract features from the current frame hand image to obtain a first shallow hand feature graph; the fusion module 8051 is used to fuse the first shallow hand feature with the first temporal feature graph to obtain a first intermediate feature graph; the second image backbone network 803 is used to extract features from the first intermediate feature graph to obtain a first deep hand feature graph; the fusion module 8052 is used to fuse the first deep hand feature graph with the first temporal feature graph to obtain a second intermediate feature graph; the third image backbone network 804 is used to extract features from the second intermediate feature graph output by the fusion module 8052 connected to the third image backbone network to obtain intermediate posture data of each hand joint in the current frame hand image.
[0079] It should be noted that the backbone network (time series backbone network, first image backbone network, second backbone network and third image backbone network) in the embodiment of the present application can be resnet (Residual Network), NAS series network and darknet series network, etc. However, the embodiment of the present application does not limit the specific structure of the backbone network. The structure of the backbone network can be set according to the actual situation. The structures of the backbone networks can be the same or different, and can be set according to the actual situation.
[0080] The training method for the dual time series model in the embodiment of the present application is as follows: first, the training data of the input branch of the time series backbone network in the dual time series model is set to all 0s, the number of training iterations is 200, and the basic model weights without superimposed time series data are obtained. The detection results of the hand joints are inferred only by the features of the conventional hand images. Secondly, the two branches of the dual time series model are input with the corresponding training data for training, and the number of training iterations is set to 200, and the learning rate is 1e-4. The weights of the dual time series model generated after training have the effect of time series reasoning, which can significantly improve the stability of the hand joint detection timing.
[0081] Step 304: if the intermediate posture data of the designated hand joint points among the hand joint points meet the designated conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain a target standard hand posture matching the current frame hand image;
[0082] The designated hand joint points are the index finger tip joint points and the thumb finger tip joint points; and the intermediate posture data are image position coordinates. The standard hand posture in the embodiment of the present application can be a click posture, a pinch posture, a grab posture, etc., which can be set according to actual conditions.
[0083] In one embodiment, whether the intermediate posture data of a specified hand joint point among the hand joint points meets the specified condition is determined in the following manner:
[0084] Based on the image position coordinates of the index finger joint point and the image position coordinates of the thumb joint point, the distance between the index finger joint point and the thumb joint point is obtained; if the distance is less than the specified distance, it is determined that the intermediate posture data of the specified hand joint point meets the specified condition; otherwise, it is determined that the intermediate posture data of the specified hand joint point does not meet the specified condition. The distance between the index finger joint point and the thumb joint point can be obtained by formula (3):
[0085]
[0086] Where d is the distance between the index finger joint and the thumb joint, x 1 is the image horizontal position coordinate of the index finger joint point, y 1 is the image position ordinate of the index finger joint point, x 2 is the image position abscissa of the thumb fingertip joint, y 2 is the image position ordinate of the thumb fingertip joint.
[0087] It should be noted that the specified distance in the embodiment of the present application is 2 cm, but the embodiment of the present application does not limit the specified distance. The specified distance in the embodiment of the present application can be set according to actual conditions.
[0088] Next, the specific method of determining the target standard hand posture matching the hand image of the current frame in step 304 is described. Fig. 9 As shown, a flowchart for determining a target standard hand posture may include the following steps:
[0089] Step 901: Based on the intermediate posture data of each hand joint point, the intermediate posture data of each hand joint point is subjected to posture alignment and normalization processing to obtain the normalized intermediate posture data of each hand joint point; wherein the normalized intermediate posture data of any hand joint point can be obtained by formula (4):
[0090]
[0091] Where D is the translation matrix that translates the intermediate posture data of the target hand joint point to the origin of the image coordinate system. is the normalized intermediate posture data of any hand joint point, P 中 is the intermediate posture data of any hand joint point, R is the rotation matrix corresponding to any joint point, I 0_9 is the unit vector corresponding to hand joint point 0 and hand joint point 9.
[0092] It should be noted that the target hand joint point in the embodiment of the present application is the joint point at the base of the palm, that is, Figure 4 The 0th joint in the.
[0093] In the embodiments of this application in, is the horizontal position coordinate of the intermediate posture data of the target hand joint point, is the vertical position coordinate of the intermediate posture data of the target hand joint point, It is the vertical position coordinate of the intermediate posture data of the target hand joint point. in, is the vector between the hand joint No. 0 and the hand joint No. 9. In the embodiment of the present application, the hand joint No. 0 is the joint at the base of the palm, and the hand joint No. 9 is the joint at the base of the middle finger of the palm. The vector is obtained based on the intermediate posture data of the hand joint No. 0 and the intermediate posture data of the hand joint No. 9. It is the modulus of the vector between hand joint point 0 and hand joint point 9. in, is the horizontal position coordinate of the middle posture data of hand joint point No. 9, is the horizontal position coordinate of the middle posture data of hand joint point 0, is the vertical position coordinate of the middle posture data of hand joint point No. 9, It is the vertical position coordinate of the middle posture data of hand joint point 0.
[0094] The rotation matrix R corresponding to any joint point in the embodiment of the present application can be obtained by formula (5):
[0095] R=R 0,9 ×R 17,5 ……(5);
[0096] Among them, R is the rotation matrix corresponding to any joint point, R 0,9 is the first intermediate rotation matrix, R 17,5 is the second intermediate rotation matrix.
[0097] In the embodiment of the present application, the first intermediate rotation matrix can be obtained by formula (6):
[0098]
[0099] Among them, θ 1 is the first rotation angle, θ 1 =arc cosI Y ⊙I 0_9 , I Y =(0,1,0),k 1,x is the horizontal position coordinate of the first normalized rotation axis, k 1,y is the vertical position coordinate of the first normalized rotation axis, k 1,z is the vertical position coordinate of the first normalized rotation axis;
[0100] In the embodiment of the present application, the second intermediate rotation matrix can be obtained by formula (7):
[0101]
[0102] Among them, R 17,5is the second intermediate rotation matrix, θ 2 is the second rotation angle, θ 2 =arc cos I x ⊙I 17_5 , I x =(1,0,0),I 17_5 is the unit vector corresponding to the hand joint point 17 and the hand joint point 5. In the embodiment of the present application, the hand joint point 17 is the joint point at the base of the little finger of the palm, and the hand joint point 5 is the joint point at the base of the index finger of the palm. Figure 4 As shown. 17_5 The determination method is the same as I 0_9 The determination method is the same as that of k, and will not be repeated in the implementation of this application. 2,x is the horizontal position coordinate of the second normalized rotation axis, k 2,y is the vertical position coordinate of the second normalized rotation axis, k 2,z is the vertical position coordinate of the second normalized rotation axis; wherein, the second normalized rotation axis k 2 =I 0_9 .
[0103] Step 902: obtaining the similarity between each hand joint point in the hand image of the current frame and each standard hand posture according to the normalized intermediate posture data of each hand joint point and the posture data corresponding to each standard hand posture;
[0104] The similarity between each hand joint point in the current frame hand image and each standard hand posture can be obtained by formula (8):
[0105]
[0106] Wherein, L is the similarity between each hand joint point in the hand image of the current frame and the standard hand posture i, M is the total number of hand joint points, and in the embodiment of the present application, M=21, is the normalized intermediate posture data of the hand joint point j in the current frame hand image, It is the normalized standard posture data of hand joint point j in standard hand posture i.
[0107] The standard hand posture in the embodiment of the present application includes standard posture data of each hand joint point.
[0108] Step 903: Determine the posture data of the standard hand posture with the maximum similarity value with each hand joint point in the current frame image as the target standard hand posture.
[0109] In one embodiment, before executing step 305, the target standard hand posture is restored to a three-dimensional coordinate system to obtain a restored target standard hand posture, and the restored target standard hand posture is determined as the target standard hand posture. The restored target standard hand posture can be determined by formula (9):
[0110]
[0111] in, is the target standard hand posture, P selected This is the target standard hand posture after restoration.
[0112] Step 305: inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image;
[0113] The hand timing diagram of the target standard hand posture in the embodiment of the present application is the same as the method of determining the hand timing diagram corresponding to the hand image of the current frame as described above, and the embodiment of the present application will not be repeated here.
[0114] In one embodiment, step 303 can be specifically implemented as follows: using the timing backbone network to perform feature extraction on the target area of the hand timing graph of the target standard hand posture to obtain a second timing feature graph; and using the first image backbone network to perform feature extraction on the current frame hand image to obtain a second shallow hand feature graph; using the fusion module to perform image fusion on the second shallow hand feature and the second timing feature graph to obtain a third intermediate feature graph; for any second image backbone network, using the second image backbone network to perform feature extraction on the second target feature graph to obtain a second deep hand feature graph, wherein the second target feature graph is the third intermediate feature graph or the fourth intermediate feature graph output by the fusion module connected to the second image backbone network; using the fusion module to fuse the second deep hand feature graph with the second timing feature graph to obtain a fourth intermediate feature graph; using the third image backbone network to perform feature extraction on the fourth intermediate feature graph output by the fusion module connected to the third image backbone network to obtain target posture data of each hand joint in the current frame hand image.
[0115] by Figure 8The structure of the dual-temporal model in is used as an example to illustrate: the timing backbone network 801 is used to extract features from the target area of the hand timing diagram of the target standard hand posture to obtain a second timing feature diagram; and the first image backbone network 802 is used to extract features from the hand image of the current frame to obtain a second shallow hand feature diagram; the fusion module 8051 is used to fuse the second shallow hand features with the second timing feature diagram to obtain a third intermediate feature diagram; the second image backbone network 803 is used to extract features from the third intermediate feature diagram to obtain a second deep hand feature diagram; the fusion module 8052 is used to fuse the second deep hand feature diagram with the second timing feature diagram to obtain a fourth intermediate feature diagram; the third image backbone network 804 is used to extract features from the fourth intermediate feature diagram output by the fusion module 8052 connected to the third image backbone network to obtain target posture data of each hand joint in the hand image of the current frame.
[0116] like Fig. 9 As shown in the figure, it can be seen that the visualization difference between the timing data in the conventional hand timing diagram and the hand posture corresponding to the hand joint point when overlapping is relatively large. However, in the embodiment of the present application, the timing data in the second timing feature diagram obtained by extracting the features of the target area of the hand timing diagram of the target standard hand posture using a dual timing model has a relatively small visualization difference when overlapping with the hand posture corresponding to the hand joint point. Therefore, the accuracy of the target posture data of each hand joint point obtained in the embodiment of the present application is relatively high.
[0117] Step 306: using the target posture data of each hand joint point in the hand image of the current frame to drive the preset 3D gesture model.
[0118] The method of driving the preset 3D gesture model in the embodiment of the present application is the method in the embodiment of the present application, and the embodiment of the present application will not be repeated here.
[0119] In order to further connect the technical solutions in this application, Fig.10 A detailed description may include the following steps:
[0120] Step 1001: using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image, predicting the posture of each hand joint point in the current frame hand image, and obtaining the predicted posture data of each hand joint point in the current frame hand image;
[0121] Step 1002: Projecting the predicted posture data of each hand joint point using a preset camera internal parameter to obtain two-dimensional posture data of each hand joint point in the hand image of the current frame;
[0122] Step 1003: for any hand joint point, according to the two-dimensional posture data of the hand joint point, obtain the intermediate hand timing diagram corresponding to the hand joint point;
[0123] Step 1004: Obtain the hand timing diagram corresponding to the hand image of the current frame through the intermediate hand timing diagrams corresponding to the hand joint points;
[0124] Step 1005: using the pre-trained temporal backbone network in the dual temporal model to perform feature extraction on the target area of the hand temporal graph corresponding to the hand image of the current frame, to obtain a first temporal feature graph;
[0125] Step 1006: extracting features of the hand image of the current frame using the first image backbone network to obtain a first shallow hand feature map;
[0126] It should be noted that the execution sequence of step 1005 and step 1006 is not limited in the embodiment of the present application. Step 1005 may be executed first, and then step 1006; step 1006 may be executed first, and then step 1005; or step 1005 and step 1006 may be executed simultaneously.
[0127] Step 1007: using the fusion module to perform image fusion on the first shallow hand feature and the first temporal feature map to obtain a first intermediate feature map;
[0128] Step 1008: for any second image backbone network, use the second image backbone network to perform feature extraction on the first target feature map to obtain a first deep hand feature map, wherein the first target feature map is the first intermediate feature map or the second intermediate feature map output by the fusion module connected to the second image backbone network;
[0129] Step 1009: using a fusion module to fuse the first deep hand feature map with the first temporal feature map to obtain a second intermediate feature map;
[0130] Step 1010: using the third image backbone network to perform feature extraction on the second intermediate feature map output by the fusion module connected to the third image backbone network, to obtain intermediate posture data of each hand joint point in the hand image of the current frame;
[0131] Step 1011: if the intermediate posture data of the designated hand joint points among the hand joint points meet the designated conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain a target standard hand posture matching the current frame hand image;
[0132] Step 1012: inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image;
[0133] Step 1013: using the target posture data of each hand joint in the current frame hand image to drive the preset 3D gesture model.
[0134] Based on the same inventive concept, the driving method of the 3D gesture model disclosed above can also be implemented by a driving device of the 3D gesture model. The effect of the driving device of the 3D gesture model is similar to that of the aforementioned method, which will not be described in detail here.
[0135] Fig.11 The figure is a schematic diagram of the structure of a driving device of a 3D gesture model according to an embodiment of the present disclosure.
[0136] like Fig.11 As shown, the driving device 1100 of the 3D gesture model of the present disclosure may include a posture prediction module 1110, a hand timing diagram determination module 1120, an intermediate posture data determination module 1130, a matching module 1140, a target posture data determination module 1150 and a driving module 1160.
[0137] The posture prediction module 1110 is used to predict the posture of each hand joint point in the hand image of the current frame by using the posture data of each hand joint point in the hand image of the previous specified number of frames, and obtain the predicted posture data of each hand joint point in the hand image of the current frame;
[0138] A hand timing diagram determining module 1120 is used to obtain a hand timing diagram corresponding to the hand image of the current frame based on the predicted posture data of each hand joint point in the hand image of the current frame, wherein the target area of each hand joint point is marked in the hand timing diagram;
[0139] The intermediate posture data determination module 1130 is used to obtain the intermediate posture data of each hand joint point in the hand image of the current frame by using a pre-trained dual-time series model based on a time series feature map obtained by extracting features from a target area in the hand time series map and a hand feature map obtained by extracting features from the hand image of the current frame;
[0140] A matching module 1140 is configured to obtain a target standard hand posture matching the hand image of the current frame by matching the intermediate posture data of each hand joint point in the hand image of the current frame with the posture data of each preset standard hand posture if the intermediate posture data of the designated hand joint point among the hand joint points meets a specified condition;
[0141] The target posture data determination module 1150 is used to input the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image;
[0142] The driving module 1160 is used to drive a preset 3D gesture model using the target posture data of each hand joint point in the hand image of the current frame.
[0143] In one embodiment, the hand image of the first specified number of frames is the hand image of the first two frames of the hand image of the current frame; the posture prediction module 1110 is specifically used for:
[0144] The predicted posture data of any hand joint point in the hand image of the current frame is obtained by the following formula:
[0145] P t =P t-1 +(P t-1 -P t-2 )+σ;
[0146] Among them, P t is the predicted posture data of any hand joint point in the current frame hand image t, P t-1 is the posture data of any hand joint point in the hand image of frame t-1, P t-2 is the posture data of any hand joint point in the t-2th frame of the hand image, and σ is a preset noise value.
[0147] In one embodiment, the hand timing diagram determination module 1120 is specifically used for:
[0148] Projecting the predicted posture data of each hand joint point using a preset camera internal parameter to obtain two-dimensional posture data of each hand joint point in the hand image of the current frame;
[0149] For any hand joint point, according to the two-dimensional posture data of the hand joint point, an intermediate hand timing diagram corresponding to the hand joint point is obtained;
[0150] The hand timing diagram corresponding to the hand image of the current frame is obtained through the intermediate hand timing diagrams corresponding to the hand joint points.
[0151] In one embodiment, the hand timing diagram determining module 1120 is further configured to:
[0152] For any hand joint point, multiply the matrix corresponding to the camera intrinsic parameter by the matrix corresponding to the predicted posture data of the hand joint point to obtain the two-dimensional posture data of the hand joint point in the hand image of the current frame;
[0153] According to the two-dimensional posture data of the hand joint points, the positions of the hand joint points in the preset standard timing diagram are obtained, and after the target area of the hand joint points is obtained with the positions of the hand joint points in the standard timing diagram as the center, the pixel values of each pixel point in the target area are set to the first specified pixel value, and the pixel values of other pixel points in the standard timing diagram except the pixel points in the target area are set to the second specified pixel value, so as to obtain the intermediate hand timing diagram of the hand joint points.
[0154] In one embodiment, the dual temporal model includes a temporal backbone network, a first image backbone network, at least one second image backbone network, a third image backbone network and a plurality of fusion modules; the number of the fusion modules is the sum of the number of the first image backbone network and the number of the at least one second image backbone network; wherein:
[0155] An input end of any one of the plurality of fusion modules is connected to the temporal backbone network and any one of the image backbone networks, respectively, and an output end of any one of the fusion modules is connected to an image backbone network at a next layer of any one of the image backbone networks, wherein any one of the image backbone networks is one of the first image backbone network, the at least one second image backbone network, and the third image backbone network; and
[0156] The head second image backbone network in the at least one second image backbone network is connected to the first image backbone network, and the tail second image backbone network in the at least one second image backbone network is connected to the third backbone network.
[0157] In one embodiment, the intermediate posture data determination module 1130 is specifically used to:
[0158] Using the temporal backbone network to extract features from a target area of a hand temporal graph corresponding to the hand image of the current frame, to obtain a first temporal feature graph; and using the first image backbone network to extract features from the hand image of the current frame, to obtain a first shallow hand feature graph;
[0159] Using the fusion module to perform image fusion on the first shallow hand feature and the first temporal feature map to obtain a first intermediate feature map;
[0160] For any second image backbone network, use the second image backbone network to perform feature extraction on the first target feature map to obtain a first deep hand feature map, wherein the first target feature map is the first intermediate feature map or the second intermediate feature map output by the fusion module connected to the second image backbone network;
[0161] Using a fusion module to fuse the first deep hand feature map with the first temporal feature map to obtain a second intermediate feature map;
[0162] The third image backbone network is used to perform feature extraction on the second intermediate feature map output by the fusion module connected to the third image backbone network to obtain intermediate posture data of each hand joint point in the current frame hand image.
[0163] In one embodiment, the target posture data determination module 1150 is specifically used to:
[0164] Using the timing backbone network to perform feature extraction on a target area of the hand timing graph of the target standard hand posture to obtain a second timing feature graph; and using the first image backbone network to perform feature extraction on the hand image of the current frame to obtain a second shallow hand feature graph;
[0165] Using the fusion module to perform image fusion on the second shallow hand feature and the second temporal feature map to obtain a third intermediate feature map;
[0166] For any second image backbone network, use the second image backbone network to perform feature extraction on a second target feature map to obtain a second deep hand feature map, wherein the second target feature map is the third intermediate feature map or a fourth intermediate feature map output by a fusion module connected to the second image backbone network;
[0167] Using a fusion module to fuse the second deep hand feature map with the second temporal feature map to obtain a fourth intermediate feature map;
[0168] The third image backbone network is used to perform feature extraction on the fourth intermediate feature map output by the fusion module connected to the third image backbone network to obtain target posture data of each hand joint point in the current frame hand image.
[0169] In one embodiment, the designated hand joints are the index finger tip joints and the thumb finger tip joints; and the intermediate posture data are image position coordinates; the matching module 1140 is specifically used to:
[0170] Determine whether the intermediate posture data of the specified hand joint point among the hand joint points meets the specified condition in the following manner:
[0171] Based on the image position coordinates of the index fingertip joint point and the image position coordinates of the thumb fingertip joint point, obtaining the distance between the index fingertip joint point and the thumb fingertip joint point;
[0172] If the distance is less than the specified distance, determining that the intermediate posture data of the specified hand joint point satisfies the specified condition;
[0173] Otherwise, it is determined that the intermediate posture data of the specified hand joint point does not meet the specified condition.
[0174] In one embodiment, the posture data corresponding to any one of the standard hand postures includes the labeled image position coordinates of a plurality of hand joint points;
[0175] The matching module 1140 is specifically used for:
[0176] Based on the intermediate posture data of each hand joint point, posture alignment and normalization processing are performed on the intermediate posture data of each hand joint point to obtain the normalized intermediate posture data of each hand joint point;
[0177] According to the normalized intermediate posture data of each hand joint point and the posture data corresponding to each standard hand posture, obtaining the similarity between each hand joint point in the hand image of the current frame and each standard hand posture;
[0178] The posture data of the standard hand posture having the largest similarity value with each hand joint point in the current frame image is determined as the target standard hand posture.
[0179] After introducing a method and apparatus for driving a 3D gesture model according to an exemplary embodiment of the present invention, next, an electronic device according to another exemplary embodiment of the present invention is introduced.
[0180] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as a system, method or program product. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as a "circuit", "module" or "system".
[0181] In some possible implementations, the electronic device according to the present invention may include at least one processor and at least one computer storage medium. The computer storage medium stores program code, and when the program code is executed by the processor, the processor executes the steps of the driving method of the 3D gesture model according to various exemplary embodiments of the present invention described above in this specification. For example, the processor may execute the following steps: Figure 3 Steps 301-306 shown in .
[0182] Refer to the following Fig.12 The electronic device 1200 according to this embodiment of the present invention is described. Fig.12 The electronic device 1200 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0183] like Fig.12 As shown, the electronic device 1200 is in the form of a general electronic device. The components of the electronic device 1200 may include but are not limited to: the at least one processor 1201, the at least one computer storage medium 1202, and a bus 1203 connecting different system components (including the computer storage medium 1202 and the processor 1201).
[0184] Bus 1203 represents one or more of several types of bus structures, including a computer storage media bus or computer storage media controller, a peripheral bus, a processor, or a local bus using any of a variety of bus architectures.
[0185] Computer storage media 1202 may include readable media in the form of volatile computer storage media, such as random access computer storage media (RAM) 1221 and / or cache storage media 1222 , and may further include read-only computer storage media (ROM) 1223 .
[0186] The computer storage medium 1202 may also include a program / utility 1225 having a set (at least one) of program modules 1224, such program modules 1224 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0187] The electronic device 1200 may also communicate with one or more external devices 1204 (e.g., keyboards, pointing devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1200, and / or communicate with any device that enables the electronic device 1200 to communicate with one or more other electronic devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1205. Furthermore, the electronic device 1200 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1206. As shown, the network adapter 1206 communicates with other modules for the electronic device 1200 via a bus 1203. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1200, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0188] In some possible implementations, various aspects of a driving method for a 3D gesture model provided by the present invention may also be implemented in the form of a program product, which includes a program code. When the program product runs on a computer device, the program code is used to enable the computer device to execute the steps of the driving method for a 3D gesture model according to various exemplary embodiments of the present invention described above in this specification.
[0189] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A driving method for a 3D gesture model, characterized in that: The method comprises: Using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image, the posture of each hand joint point in the current frame hand image is predicted to obtain the predicted posture data of each hand joint point in the current frame hand image; Based on the predicted posture data of each hand joint point in the hand image of the current frame, a hand timing diagram corresponding to the hand image of the current frame is obtained, wherein the target area of each hand joint point is marked in the hand timing diagram; Using a pre-trained dual-time series model, based on a time series feature map obtained by extracting features from a target area in the hand time series map and a hand feature map obtained by extracting features from the hand image of the current frame, intermediate posture data of each hand joint point in the hand image of the current frame are obtained; If the intermediate posture data of the designated hand joint points among the hand joint points meet the specified conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain a target standard hand posture that matches the current frame hand image; Inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image; The preset 3D gesture model is driven by using the target posture data of each hand joint point in the current frame hand image.
2. The method according to claim 1, characterized in that The hand images of the first specified number of frames are the hand images of the first two frames of the hand image of the current frame; The method of using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image to predict the posture of each hand joint point in the current frame hand image to obtain the predicted posture data of each hand joint point in the current frame hand image includes: The predicted posture data of any hand joint point in the hand image of the current frame is obtained by the following formula: P t =P t-1 +(P t-1 -P t-2 )+σ; Among them, P t is the predicted posture data of any hand joint point in the current frame hand image t, P t-1 is the posture data of any hand joint point in the hand image of frame t-1, P t-2 is the posture data of any hand joint point in the t-2th frame of the hand image, and σ is a preset noise value.
3. The method according to claim 1, characterized in that The step of obtaining the hand timing diagram corresponding to the hand image of the current frame based on the predicted posture data of each hand joint point in the hand image of the current frame includes: Projecting the predicted posture data of each hand joint point using a preset camera internal parameter to obtain two-dimensional posture data of each hand joint point in the hand image of the current frame; For any hand joint point, according to the two-dimensional posture data of the hand joint point, an intermediate hand timing diagram corresponding to the hand joint point is obtained; The hand timing diagram corresponding to the hand image of the current frame is obtained through the intermediate hand timing diagrams corresponding to the hand joint points.
4. The method according to claim 3, characterized in that The projecting the predicted posture data of each hand joint point by using the preset camera internal parameter to obtain the two-dimensional posture data of each hand joint point in the hand image of the current frame includes: For any hand joint point, multiply the matrix corresponding to the camera intrinsic parameter by the matrix corresponding to the predicted posture data of the hand joint point to obtain the two-dimensional posture data of the hand joint point in the hand image of the current frame; For any hand joint point, according to the two-dimensional posture data of the hand joint point, obtaining the intermediate hand timing diagram corresponding to the hand joint point includes: According to the two-dimensional posture data of the hand joint points, the positions of the hand joint points in the preset standard timing diagram are obtained, and after the target area of the hand joint points is obtained with the positions of the hand joint points in the standard timing diagram as the center, the pixel values of each pixel point in the target area are set to the first specified pixel value, and the pixel values of other pixel points in the standard timing diagram except the pixel points in the target area are set to the second specified pixel value, so as to obtain the intermediate hand timing diagram of the hand joint points.
5. The method according to claim 1, characterized in that: The dual-time series model includes a time series backbone network, a first image backbone network, at least one second image backbone network, a third image backbone network and a plurality of fusion modules; the number of the fusion modules is the sum of the number of the first image backbone network and the number of the at least one second image backbone network; wherein: An input end of any one of the plurality of fusion modules is connected to the temporal backbone network and any one of the image backbone networks, respectively, and an output end of any one of the fusion modules is connected to an image backbone network at a next layer of any one of the image backbone networks, wherein any one of the image backbone networks is one of the first image backbone network, the at least one second image backbone network, and the third image backbone network; and The head second image backbone network in the at least one second image backbone network is connected to the first image backbone network, and the tail second image backbone network in the at least one second image backbone network is connected to the third backbone network.
6. The method according to claim 5, characterized in that The pre-trained dual-time series model is used to extract the time series feature map obtained by extracting the features of the target area in the hand time series map and the hand feature map obtained by extracting the features of the hand image of the current frame to obtain the intermediate posture data of each hand joint point in the hand image of the current frame, including: Using the temporal backbone network to extract features from a target area of a hand temporal graph corresponding to the hand image of the current frame, to obtain a first temporal feature graph; and using the first image backbone network to extract features from the hand image of the current frame, to obtain a first shallow hand feature graph; Using the fusion module to perform image fusion on the first shallow hand feature and the first temporal feature map to obtain a first intermediate feature map; For any second image backbone network, use the second image backbone network to perform feature extraction on the first target feature map to obtain a first deep hand feature map, wherein the first target feature map is the first intermediate feature map or the second intermediate feature map output by the fusion module connected to the second image backbone network; Using a fusion module to fuse the first deep hand feature map with the first temporal feature map to obtain a second intermediate feature map; The third image backbone network is used to perform feature extraction on the second intermediate feature map output by the fusion module connected to the third image backbone network to obtain intermediate posture data of each hand joint point in the current frame hand image.
7. The method according to claim 5, characterized in that Inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image, including: Using the timing backbone network to perform feature extraction on a target area of the hand timing graph of the target standard hand posture to obtain a second timing feature graph; and using the first image backbone network to perform feature extraction on the hand image of the current frame to obtain a second shallow hand feature graph; Using the fusion module to perform image fusion on the second shallow hand feature and the second temporal feature map to obtain a third intermediate feature map; For any second image backbone network, use the second image backbone network to perform feature extraction on a second target feature map to obtain a second deep hand feature map, wherein the second target feature map is the third intermediate feature map or a fourth intermediate feature map output by a fusion module connected to the second image backbone network; Using a fusion module to fuse the second deep hand feature map with the second temporal feature map to obtain a fourth intermediate feature map; The third image backbone network is used to perform feature extraction on the fourth intermediate feature map output by the fusion module connected to the third image backbone network to obtain target posture data of each hand joint point in the current frame hand image.
8. The method according to claim 1, characterized in that: The designated hand joint points are the index finger tip joint points and the thumb finger tip joint points; and the intermediate posture data are image position coordinates; Determine whether the intermediate posture data of the specified hand joint point among the hand joint points meets the specified condition in the following manner: Based on the image position coordinates of the index fingertip joint point and the image position coordinates of the thumb fingertip joint point, obtaining the distance between the index fingertip joint point and the thumb fingertip joint point; If the distance is less than the specified distance, determining that the intermediate posture data of the specified hand joint point satisfies the specified condition; Otherwise, it is determined that the intermediate posture data of the specified hand joint point does not meet the specified condition.
9. The method according to claim 1, characterized in that: The posture data corresponding to any one of the standard hand postures includes the labeled image position coordinates of a plurality of hand joint points; The step of matching the intermediate posture data of each hand joint point in the current frame hand image with the posture data of each preset standard hand posture to obtain a target standard hand posture matching the current frame hand image includes: Based on the intermediate posture data of each hand joint point, posture alignment and normalization processing are performed on the intermediate posture data of each hand joint point to obtain the normalized intermediate posture data of each hand joint point; According to the normalized intermediate posture data of each hand joint point and the posture data corresponding to each standard hand posture, obtaining the similarity between each hand joint point in the hand image of the current frame and each standard hand posture; The posture data of the standard hand posture having the largest similarity value with each hand joint point in the current frame image is determined as the target standard hand posture.
10. An electronic device, characterized in that: comprising a processor and a memory, wherein the processor and the memory are connected via a bus; The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program: Using the posture data of each hand joint point in the hand image located at a specified number of frames before the current frame hand image, the posture of each hand joint point in the current frame hand image is predicted to obtain the predicted posture data of each hand joint point in the current frame hand image; Based on the predicted posture data of each hand joint point in the hand image of the current frame, a hand timing diagram corresponding to the hand image of the current frame is obtained, wherein the target area of each hand joint point is marked in the hand timing diagram; Using a pre-trained dual-time series model, based on a time series feature map obtained by extracting features from a target area in the hand time series map and a hand feature map obtained by extracting features from the hand image of the current frame, intermediate posture data of each hand joint point in the hand image of the current frame are obtained; If the intermediate posture data of the designated hand joint points among the hand joint points meet the specified conditions, the intermediate posture data of the hand joint points in the current frame hand image are matched with the posture data of the preset standard hand postures to obtain a target standard hand posture that matches the current frame hand image; Inputting the hand timing diagram of the current frame hand image and the target standard hand posture into the dual timing model to obtain the target posture data of each hand joint point in the current frame hand image; The preset 3D gesture model is driven by using the target posture data of each hand joint point in the current frame hand image.
Citation Information
Cited By
Method, apparatus, and computer program product for closed eye defect repair
CN122554723A