Hand data set generation method and electronic equipment
By generating hand datasets using motion capture systems and 3D synthesis technology, the problems of low accuracy and efficiency of hand datasets in existing technologies have been solved, achieving high-precision and high-efficiency automated generation of hand datasets.
Patent Information
- Application Number
- CN202410841685.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-02
AI Technical Summary
The existing hand datasets have low accuracy and low generation efficiency, mainly due to the low annotation accuracy of human hand joints and the long time required for manual annotation.
By utilizing the motion trajectory of the motion capture system and combining it with a 3D synthesized 3D human hand model, a hand dataset is generated. The motion trajectory of the camera and hand joints is restored using transformation relationships, and the hand dataset is automatically generated, avoiding errors caused by finger occlusion.
It improves the accuracy and generation efficiency of hand datasets, ensuring accurate annotation of 2D and 3D point positions even in complex hand movements, reducing the need for manual processing.
Smart Images

Figure CN121259902A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method for generating a hand dataset and an electronic device. Background Technology
[0002] Gesture recognition, or hand pose estimation, refers to predicting the three-dimensional coordinates of hand joints in an image containing a human hand. Thanks to the rapid development of deep neural networks, significant progress has been made in hand pose estimation based on two-dimensional images. However, these deep neural network-based methods heavily rely on large training datasets.
[0003] In existing technologies, the data set creation method involves acquiring images recorded by real devices, and then manually or semi-automatically annotating the 2D positions of each joint of the human hand on the images to obtain a hand dataset. Because the human hand has a very complex joint structure and is prone to severe occlusion during complex hand gestures, the accuracy of manual annotation of hand joints is low. This results in a low-precision hand dataset, and the long time required for manual annotation further reduces the efficiency of hand dataset generation. Summary of the Invention
[0004] This application provides a method and electronic device for generating a hand dataset, which uses a 3D synthesized human hand model to ensure that the annotation of hand joints is almost error-free. It can generate not only 2D annotations of human hand joints, but also 3D annotation information. Even for complex hand movements, there will be no errors due to finger occlusion, thus improving the accuracy of the hand dataset. Furthermore, in the embodiments of this application, the hand dataset is automatically generated without manual processing, which improves efficiency.
[0005] In a first aspect, embodiments of this application provide a method for generating a hand dataset, the method comprising:
[0006] The motion trajectory of the head-mounted display and the motion trajectory of each joint in the hand are received from the motion capture system; wherein the motion trajectory is composed of multiple poses in the world coordinate system;
[0007] The motion trajectory of the head-mounted display is transformed using each first transformation relationship to obtain the motion trajectory of each camera mounted on the head-mounted display; wherein, any one of the first transformation relationships is the transformation relationship between the head-mounted display and any one of the cameras on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras;
[0008] Based on the motion trajectories of each camera mounted on the head-mounted display, the motion trajectory of the virtual camera is obtained, and based on the motion trajectories of each hand joint, the motion trajectory of the three-dimensional hand model is obtained.
[0009] The motion trajectory of the three-dimensional hand model is used to drive the three-dimensional hand model, and the motion trajectory of the virtual camera is used to drive the virtual camera to take pictures of the three-dimensional hand model, thereby obtaining multiple green screen images of the human hand, wherein the multiple green screen images of the human hand are the foreground images of the human hand used to generate the hand dataset;
[0010] The multiple green screen images of the human hand and the multiple preset background images are combined to obtain multiple target human hand images; wherein, the multiple background images are obtained by taking pictures with multiple cameras of the same model;
[0011] Based on the multiple target hand images, the 3D position coordinates of each hand joint in the multiple target hand images, and the image position coordinates of each hand joint in the multiple target hand images, a hand dataset is obtained. The 3D position coordinates and image position coordinates of each hand joint are obtained based on the motion trajectory of each hand joint.
[0012] A second aspect of this application provides an electronic device, including a processor and a memory, wherein the processor and the memory are connected via a bus;
[0013] The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program:
[0014] The motion trajectory of the head-mounted display and the motion trajectory of each joint in the hand are received from the motion capture system; wherein the motion trajectory is composed of multiple poses in the world coordinate system;
[0015] The motion trajectory of the head-mounted display is transformed using each first transformation relationship to obtain the motion trajectory of each camera mounted on the head-mounted display; wherein, any one of the first transformation relationships is the transformation relationship between the head-mounted display and any one of the cameras on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras;
[0016] Based on the motion trajectories of each camera mounted on the head-mounted display, the motion trajectory of the virtual camera is obtained, and based on the motion trajectories of each hand joint, the motion trajectory of the three-dimensional hand model is obtained.
[0017] The motion trajectory of the three-dimensional hand model is used to drive the three-dimensional hand model, and the motion trajectory of the virtual camera is used to drive the virtual camera to take pictures of the three-dimensional hand model, thereby obtaining multiple green screen images of the human hand, wherein the multiple green screen images of the human hand are the foreground images of the human hand used to generate the hand dataset;
[0018] The multiple green screen images of the human hand and the multiple preset background images are combined to obtain multiple target human hand images; wherein, the multiple background images are obtained by taking pictures with multiple cameras of the same model;
[0019] Based on the multiple target hand images, the 3D position coordinates of each hand joint in the multiple target hand images, and the image position coordinates of each hand joint in the multiple target hand images, a hand dataset is obtained. The 3D position coordinates and image position coordinates of each hand joint are obtained based on the motion trajectory of each hand joint.
[0020] According to a third aspect of the present invention, a computer storage medium is provided, the computer storage medium storing a computer program for performing the method as described in the first aspect.
[0021] In the above embodiments of this application, the motion trajectory of the head-mounted display sent by the motion capture system is transformed using various first transformation relationships to obtain the motion trajectory of each camera installed on the head-mounted display; based on the motion trajectory of each camera installed on the head-mounted display, the motion trajectory of the virtual camera is obtained, and based on the motion trajectory of each hand joint, the motion trajectory of the three-dimensional hand model is obtained; the motion trajectory of the three-dimensional hand model is used to drive the three-dimensional hand model, and the motion trajectory of the virtual camera is used to drive the virtual camera to take pictures of the three-dimensional hand model, obtaining multiple green screen images of the human hand; the multiple green screen images of the human hand and multiple preset background images are combined to obtain multiple target human hand images; based on the multiple target human hand images, the 3D position coordinates of each hand joint and the multiple target human hand images, and the image position coordinates of each hand joint in the multiple target human hand images, a human hand dataset is obtained. Thus, in the embodiments of this application, the actual motion trajectory of the camera restored based on the rigid body motion trajectory of the motion capture system and the hand posture captured in the actual device restored based on the motion trajectory of the hand joints are used to make the final synthesized target hand image conform to the imaging angle of the hand image observed during the actual head-mounted display movement. Furthermore, the 3D synthesized three-dimensional human hand model ensures that the annotation of hand joints is almost error-free. It can generate not only 2D annotations of hand joints, but also 3D annotation information. Even for complex hand movements, there will be no errors due to finger occlusion, thus improving the accuracy of the hand dataset. In addition, the hand dataset is automatically generated in this embodiment without manual processing, which improves efficiency. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 One of the flowcharts illustrating the method for generating a hand dataset provided in this application is shown as an example.
[0024] Figure 2 An exemplary schematic diagram of the process for determining the first transformation relationship provided in an embodiment of this application is shown;
[0025] Figure 3 An exemplary schematic diagram of the calibration board provided in an embodiment of this application is shown;
[0026] Figure 4 An exemplary schematic diagram illustrates the process of filtering multiple green screen images of human hands provided in an embodiment of this application;
[0027] Figure 5 An exemplary schematic diagram illustrates the process of determining the position coordinates of each hand joint point in the camera coordinate system according to an embodiment of this application;
[0028] Figure 6 An exemplary schematic diagram illustrates the process of determining a target human hand image provided in an embodiment of this application;
[0029] Figure 7 An exemplary schematic diagram illustrates the process of determining multiple target human hand images provided in an embodiment of this application;
[0030] Figure 8 An exemplary flowchart of a method for generating a hand dataset provided in an embodiment of this application is shown.
[0031] Figure 9 An exemplary diagram shows a hand dataset generation apparatus provided in an embodiment of this application;
[0032] Figure 10 An exemplary structural diagram of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0033] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.
[0034] Based on the exemplary embodiments described in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the appended claims. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation on its own.
[0035] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0036] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to be omnipresent but not exclusive; for example, a product or device comprising a series of components is not necessarily limited to those explicitly listed, but may include other components not explicitly listed or inherent to such product or device.
[0037] As used in this application, the term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.
[0038] The following is an overview of the ideas behind the embodiments of this application.
[0039] In existing technologies, the data set creation method involves acquiring images recorded by real devices, and then manually or semi-automatically annotating the 2D positions of each joint of the human hand on the images to obtain a hand dataset. Because the human hand has a very complex joint structure and is prone to severe occlusion during complex hand gestures, the accuracy of manual annotation of hand joints is low. This results in a low-precision hand dataset, and the long time required for manual annotation further reduces the efficiency of hand dataset generation.
[0040] Addressing the issues of low accuracy and low generation efficiency in existing hand datasets, this application proposes a method for generating hand datasets. This method involves transforming the motion trajectory of a head-mounted display (HUD) sent by a motion capture system using first transformation relationships to obtain the motion trajectories of cameras mounted on the HUD; obtaining the motion trajectory of a virtual camera based on the motion trajectories of the cameras on the HUD; and obtaining the motion trajectory of a 3D hand model based on the motion trajectories of each hand joint. The method then uses the motion trajectory of the 3D hand model to drive the 3D hand model and uses the motion trajectory of the virtual camera to drive the virtual camera to take pictures of the 3D hand model, resulting in multiple green screen images of the hand. These green screen images are then combined with multiple preset background images to obtain multiple target hand images. Finally, based on these target hand images, the 3D position coordinates of each hand joint corresponding to the multiple target hand images, and the image position coordinates of each hand joint within the multiple target hand images, a hand dataset is generated. Therefore, in this embodiment, the actual motion trajectory of the camera, based on the rigid body motion trajectory reconstruction of the motion capture system, and the motion trajectory reconstruction of the hand joints, are used to recreate the hand posture captured in the actual device. This ensures that the final synthesized target hand image matches the imaging angle of the hand image observed during actual head-mounted display movement. Furthermore, the 3D synthesized three-dimensional hand model ensures that the annotation of the hand joints is almost error-free. It can generate not only 2D annotations of hand joints but also 3D annotation information. Even for complex hand movements, there will be no errors due to finger occlusion, thus improving the accuracy of the hand dataset. In addition, in this embodiment, the hand dataset is automatically generated without manual processing, improving efficiency.
[0041] The method for generating the hand dataset in the embodiments of this application will be described below with reference to the accompanying drawings. Figure 1 The diagram shown illustrates the process of generating a hand dataset, which may include the following steps:
[0042] Step 101: Receive the motion trajectory of the head-mounted display and the motion trajectory of each joint of the hand sent by the motion capture system; wherein the motion trajectory is composed of multiple poses in the world coordinate system;
[0043] In this embodiment, multiple fixed cameras are mounted on the head-mounted display to capture external images. To capture the motion trajectory of the head-mounted display in a motion capture system, multiple marker points (i.e., small balls capable of reflecting infrared light) need to be attached to the head-mounted display. In the motion capture system, these points are connected together to form a rigid body, thereby obtaining an overall motion trajectory, i.e., the motion trajectory of the head-mounted display.
[0044] Furthermore, in this embodiment, after marking points (i.e., small balls capable of reflecting infrared light) are affixed to the joints of the hand, the motion trajectory of each hand joint is captured based on the motion capture system. In this embodiment, each hand joint has a corresponding motion trajectory. The motion trajectory of any hand joint includes multiple poses of that hand joint in the world coordinate system, that is, multiple 3D position coordinates and directions of that hand joint in the world coordinate system. Moreover, in this embodiment, the number of poses in the motion trajectory of the head-mounted display is the same as the number of poses of any single hand joint.
[0045] Step 102: Transform the motion trajectory of the head-mounted display using each of the first transformation relationships to obtain the motion trajectory of each camera mounted on the head-mounted display; wherein, any one of the first transformation relationships is the transformation relationship between the head-mounted display and any one of the cameras on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras;
[0046] Before describing the method in step 102 of this application embodiment, the method for determining the first transformation relationship in this application embodiment will be explained first, such as... Figure 2 The diagram shown illustrates the process for determining the first transformation relationship, which may include the following steps:
[0047] Step 201: For any camera installed on the head-mounted display, acquire images of the calibration board taken by the camera at any two different times;
[0048] In this embodiment, a stationary calibration plate is used. While the motion capture system captures the motion trajectory of the head-mounted display (HUD), the HUD is positioned at multiple different stationary angles to capture images of the calibration plate using a camera. Because the calibration plate is stationary, its pose in the world coordinate system remains unchanged. Figure 3 The diagram shown is a schematic of a calibration plate, which includes multiple calibration points. Figure 3 The image of the calibration board in the illustration is for illustrative purposes only and does not limit the calibration board in the embodiments of this application.
[0049] Wherein, formula (1) represents the world position coordinates of the calibration plate at time i.
[0050]
[0051] in, Let be the matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time i. This is the matrix corresponding to the first transformation relationship between the nth camera on the headset and the headset itself. This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time i.
[0052] Formula (2) gives the world position coordinates of the calibration plate at time j.
[0053]
[0054] in, Let be the matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time j. This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time j.
[0055] Step 202: Use the images of the calibration board taken by any one of the cameras at any two different times to calibrate the camera, and obtain the second transformation relationships between the camera and the calibration board at any two different times, wherein any one of the second transformation relationships is used to represent the external parameter transformation relationship between the camera and the calibration board at any time.
[0056] In one embodiment, step 202 may be specifically implemented as follows: for an image of a calibration board captured by any camera at any time, wherein the image of the calibration board includes each calibration point in the calibration board; using the image position coordinates of each calibration point in the image of the calibration board, as well as the pre-set transformation relationship between the image coordinate system and the camera coordinate system, and the transformation relationship between the image coordinate system and the calibration board, the position coordinates of each calibration point in the camera coordinate system are obtained; using the position coordinates of each calibration point in the camera coordinate system and the position coordinates of each calibration point in the calibration board coordinate system, the second transformation relationship is obtained.
[0057] In this embodiment, the transformation relationship between the image coordinate system and the calibration plate can be obtained using the least squares method. Since the least squares method is existing technology, this embodiment does not limit its application. The second transformation relationship can be obtained using formula (3):
[0058]
[0059] in, This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time i. Let be the coordinates of calibration point m in the coordinate system of the calibration plate at time i. Let be the image position coordinates of calibration point m in the image coordinate system at time i, and T1 be the transformation relationship between the image coordinate system and the camera coordinate system.
[0060] In this embodiment, each calibration point at any given time has the corresponding equation in the above formula (3), and the second transformation relationship can be obtained by solving the system of equations. This embodiment will not be elaborated further here.
[0061] Step 203: Based on the 3D position coordinates of the head-mounted display in the world coordinate system at any two different times and the second transformation relationships, obtain the first transformation relationship between any camera and the head-mounted display, wherein the 3D position coordinates of the head-mounted display in the world coordinate system at any time are obtained from the motion trajectory of the head-mounted display.
[0062] Based on the fact that the pose of the calibration plate in the world coordinate system is invariant as described above, the first transformation relationship can be obtained through formula (4):
[0063]
[0064] in, Let be the matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time i. This is the matrix corresponding to the first transformation relationship between the nth camera on the headset and the headset itself. This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time i. Let be the matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time j. This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time j.
[0065] After introducing the method of the first transformation relationship, the method of determining the motion trajectory of each camera installed on the head-mounted display in step 102 of the embodiment of this application will be introduced. In one embodiment, step 102 can be specifically implemented as follows: for any camera, multiply the first transformation relationship corresponding to the arbitrary camera by multiple poses in the motion trajectory of the head-mounted display to obtain multiple poses of the arbitrary camera; and based on the multiple poses of the arbitrary camera, obtain the motion trajectory of the arbitrary camera.
[0066] Step 103: Based on the motion trajectories of each camera mounted on the head-mounted display, obtain the motion trajectory of the virtual camera, and based on the motion trajectories of each hand joint, obtain the motion trajectory of the three-dimensional hand model.
[0067] In this embodiment, the motion trajectories of each camera are input into Unity to construct a moving virtual camera, thus obtaining the motion trajectory of the virtual camera. Similarly, the motion trajectories of each hand joint are input into Unity to obtain the motion trajectory of the 3D hand model. That is, the 3D hand model is driven to move to specified positions based on the motion trajectories of each hand joint, and corresponding hand gesture shapes exist at each specified position.
[0068] Step 104: Drive the 3D hand model using the motion trajectory of the 3D hand model, and drive the virtual camera to take pictures of the 3D hand model using the motion trajectory of the virtual camera, to obtain multiple green screen images of the human hand, wherein the multiple green screen images of the human hand are the foreground images of the human hand used to generate the hand dataset;
[0069] In this embodiment, a green screen background is used, so multiple green screen images of human hands are obtained, that is, multiple images of a three-dimensional hand model against a green screen background captured by a virtual camera.
[0070] To further improve the accuracy of the hand dataset, multiple green screen images of human hands need to be filtered, such as... Figure 4 The diagram illustrates the process of filtering multiple green screen images of human hands, which may include the following steps:
[0071] Step 401: For any one of the multiple green screen images of a human hand, based on the 3D position coordinates of each hand joint corresponding to the green screen image of a human hand in the world coordinate system and the pose of the camera corresponding to the green screen image of a human hand, obtain the position coordinates of each hand joint in the camera coordinate system, wherein the 3D position coordinates of each hand joint corresponding to the green screen image of a human hand in the world coordinate system are obtained based on the motion trajectory of each hand joint;
[0072] In this embodiment, the green screen image of the hand, the 3D position coordinates of each hand joint in the world coordinate system, and the camera pose all include timestamps. Therefore, the green screen images of the hand with the same timestamp, the 3D position coordinates of each hand joint in the world coordinate system, and the camera pose are determined as a set of data, and the position coordinates of each hand joint in the camera coordinate system are determined based on this set of data.
[0073] Any pose in any hand joint trajectory in the embodiments of this application includes the 3D position coordinates of the hand joint in the world coordinate system.
[0074] like Figure 5 The diagram shown illustrates the process of determining the position coordinates of each hand joint point in the camera coordinate system in step 401, which may include the following steps:
[0075] Step 501: Based on the 3D position coordinates of each hand joint corresponding to any green screen image of a human hand in the world coordinate system, obtain the target pose matrix corresponding to each hand joint.
[0076] In one embodiment, step 501 can be specifically implemented as follows: for any hand joint, the 3D position coordinates of the hand joint in the world coordinate system and the preset values are used to form a target pose matrix corresponding to the hand joint.
[0077] For example, the 3D position coordinates corresponding to hand joint 1 are (x1, y1, z1). Taking the preset value as 1, the intermediate pose matrix corresponding to hand joint 1 is [x1, y1, z1, 1].
[0078] It should be noted that the preset value in this application embodiment is 1, but this does not limit the preset value in this application embodiment. The preset value in this application embodiment can be set according to the actual situation.
[0079] Step 502: Multiply the matrix formed by the pose of the camera corresponding to any green screen image of a human hand with the target pose matrix corresponding to each hand joint to obtain the position coordinates of each hand joint in the camera coordinate system.
[0080] In this embodiment, the matrix of camera poses is four rows and one column.
[0081] Step 402: Determine whether the position coordinates of each hand joint point on the camera coordinate system meet the specified conditions. If yes, proceed to step 403; otherwise, proceed to step 406.
[0082] In one embodiment, step 402 can be specifically implemented as follows: if at least one of the hand joints has a vertical coordinate in the camera coordinate system that is less than a specified value, then it is determined that the position coordinates of the hand joints in the camera coordinate system do not meet the specified condition; or if none of the hand joints has a vertical coordinate in the camera coordinate system that is less than the specified value, then it is determined that the position coordinates of the hand joints in the camera coordinate system meet the specified condition.
[0083] The specified value in this embodiment is 0, but the specific value of the specified value in this embodiment is not limited. The specified value in this embodiment can be set according to the actual situation.
[0084] Step 403: Project the image position coordinates of each hand joint point using the intrinsic parameters of the camera;
[0085] The method of projecting hand joints using camera intrinsic parameters in this embodiment is prior art, and will not be described in detail here.
[0086] Step 404: Based on the image position coordinates of each hand joint, determine whether each hand joint is located inside any green screen image of a human hand. If yes, proceed to step 405; otherwise, proceed to step 406.
[0087] Step 405: Retain any one of the person's hand green screen images;
[0088] Step 406: Delete any of the green screen images of human hands.
[0089] Therefore, in this embodiment of the application, the green screen image of the human hand needs to be filtered before image synthesis to ensure that the hand joints in the final green screen image of the human hand are complete and clearly visible, so as to further improve the accuracy of the hand dataset.
[0090] Step 105: Combine the multiple green screen images of the human hand with multiple preset background images to obtain multiple target human hand images; wherein, the multiple background images are obtained by shooting with multiple cameras of the same model;
[0091] The following describes the method for determining multiple target hand images, such as... Figure 6 The diagram shown illustrates the process of determining a target hand image, which may include the following steps:
[0092] Step 601: For any green screen image of a human hand, replace the green screen background of the green screen image of the human hand with a transparent background to obtain a human hand image with a transparent background;
[0093] In this embodiment, based on the image processing library OpenCV, the green screen image of a human hand can be converted from the RGB image space to the HSV image space, and pixels with H values within a preset range can be deleted to obtain a human hand image with a transparent background.
[0094] The preset range in this embodiment is the green hue range, but this embodiment does not limit the preset range. The specific value of the preset range in this embodiment can be set according to the actual situation.
[0095] Step 602: Perform distortion processing on the transparent background hand image using preset distortion parameters to obtain a distorted transparent background hand image;
[0096] In this embodiment, a distorted image of a human hand with a transparent background is processed using preset distortion parameters, and the joints of each hand in the image are also distorted to obtain a distorted image of a human hand with a transparent background.
[0097] To further improve accuracy, in one embodiment, the position coordinates of each person's hand joint after distortion are used to determine whether each person's hand joint after distortion is located in the distorted transparent background image. If so, the distorted transparent background image is retained; otherwise, the distorted transparent background image is deleted.
[0098] Step 603: Combine the distorted transparent background hand image with the multiple background images to obtain the multiple target hand images, wherein the distorted transparent background hand image has the same size as any of the background images.
[0099] Therefore, in this embodiment of the application, the distortion parameters obtained by calibration of the actual head-mounted display device are used to add the same distortion to the synthesized image during image synthesis. This makes the foreground image of the arm closer to the imaging effect of the actual camera, and further improves the accuracy of the generated hand dataset.
[0100] This application also provides another method for determining multiple target hand images. In one embodiment, for any green screen image of a hand, the green screen background of the green screen image of the hand is replaced with a transparent background to obtain a transparent background hand image. The transparent background hand image and the multiple background images are then combined to obtain the multiple target hand images.
[0101] The following explains the method for determining multiple target hand images, such as... Figure 7 The diagram illustrates the process of determining multiple target hand images, which may include the following steps:
[0102] Step 701: For any one of the multiple background images, adjust the brightness of the distorted transparent background hand image according to the median brightness value of the first target region in the any one background image to obtain a first intermediate hand image, wherein the first target region is the region in the any one background image that has the same position coordinates as the hand region in the distorted transparent background hand image.
[0103] In one embodiment, step 701 may be specifically implemented as: adjusting the brightness of the distorted transparent background hand image to be the same as the median brightness value.
[0104] In this embodiment, the overall brightness of the distorted transparent background hand image is adjusted using OpenCV's `convertScaleAbs` function. Specifically, the brightness of the distorted transparent background hand image is scaled to match the median brightness of the first target region in any given background image, thus adapting the distorted transparent background hand image to the lighting environment of the background image. This reduces the unnatural blending of the distorted transparent background hand image with the background, making it closer to the effect of actual camera capture and further improving the accuracy of the generated hand dataset.
[0105] If the multiple target hand images in this application embodiment are obtained by combining the transparent background hand image with the multiple background images respectively, then the distorted transparent background hand image in step 701 is the transparent background hand image.
[0106] Step 702: Based on the transparency layer of the first intermediary hand image, obtain a second intermediary hand image, wherein the second intermediary hand image is an image containing the edge region of the hand;
[0107] In this embodiment, the transparency layer of the first intermediary hand image is extracted separately, its edge region is obtained using the Canny edge detection function of OpenCV, and the edge region is dilated using the erode function of OpenCV to obtain the second intermediary hand image.
[0108] Step 703: Combine the second intermediate hand image with any of the background images to obtain a composite image, and perform linear interpolation on the edge region of the hand in the composite image to obtain the target hand image corresponding to any of the background images.
[0109] In this embodiment, the edge region in the second intermediate hand image is linearly interpolated and filled using the inpaint function of OpenCV to obtain the target hand image.
[0110] Linear interpolation fill creates a smooth, filtered boundary, reducing the unnatural blending of the image with the background and making it closer to the effect of an actual camera shot. Since the background image itself is taken with a real camera, no additional distortion needs to be added.
[0111] Step 106: Based on the multiple target hand images, the 3D position coordinates of each hand joint corresponding to the multiple target hand images, and the image position coordinates of each hand joint in the multiple target hand images, a hand dataset is obtained, wherein the 3D position coordinates and image position coordinates of each hand joint are obtained based on the motion trajectory of each hand joint.
[0112] In this embodiment, the image position coordinates of each hand joint in the multiple target hand images are the distorted image position coordinates of each hand joint in the transparent background hand image in step 602 described above.
[0113] In this embodiment, the 3D position coordinates of any hand joint point corresponding to any target human hand image are obtained based on the timestamp of the target human hand image. That is, the 3D position coordinates of any hand joint point in the motion trajectory with the same timestamp as the target human hand image are determined as the 3D position coordinates of any hand joint point corresponding to the target human hand image.
[0114] To further understand the methods in this application, such as Figure 8 The diagram shown is a flowchart illustrating the method for generating the hand dataset in this application, which may include the following steps:
[0115] Step 801: Receive the motion trajectory of the head-mounted display and the motion trajectory of each joint of the hand sent by the motion capture system; wherein the motion trajectory is composed of multiple poses in the world coordinate system;
[0116] Step 802: Transform the motion trajectory of the head-mounted display using each of the first transformation relationships to obtain the motion trajectory of each camera mounted on the head-mounted display; wherein, any one of the first transformation relationships is the transformation relationship between the head-mounted display and any one of the cameras on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras;
[0117] Step 803: Based on the motion trajectories of each camera mounted on the head-mounted display, obtain the motion trajectory of the virtual camera, and based on the motion trajectories of each hand joint, obtain the motion trajectory of the three-dimensional hand model;
[0118] Step 804: Drive the 3D hand model using the motion trajectory of the 3D hand model, and drive the virtual camera to take pictures of the 3D hand model using the motion trajectory of the virtual camera, to obtain multiple green screen images of the human hand, wherein the multiple green screen images of the human hand are the foreground images of the human hand used to generate the hand dataset;
[0119] Step 805: For any one of the multiple green screen images of a human hand, based on the 3D position coordinates of each hand joint corresponding to the green screen image of a human hand in the world coordinate system and the pose of the camera corresponding to the green screen image of a human hand, obtain the position coordinates of each hand joint in the camera coordinate system, wherein the 3D position coordinates of each hand joint corresponding to the green screen image of a human hand in the world coordinate system are obtained based on the motion trajectory of each hand joint;
[0120] Step 806: Determine whether the position coordinates of each hand joint point on the camera coordinate system meet the specified conditions. If yes, proceed to step 807; otherwise, proceed to step 810.
[0121] Step 807: Project the image position coordinates of each hand joint point using the intrinsic parameters of the camera;
[0122] Step 808: Based on the image position coordinates of each hand joint, determine whether each hand joint is located inside any green screen image of a human hand. If yes, proceed to step 809; otherwise, proceed to step 810.
[0123] Step 809: Retain any one of the person's hand green screen images;
[0124] Step 810: Delete any of the green screen images of a human hand;
[0125] Step 811: For any green screen image of a human hand, replace the green screen background of the green screen image of the human hand with a transparent background to obtain a human hand image with a transparent background;
[0126] Step 812: Perform distortion processing on the transparent background hand image using preset distortion parameters to obtain a distorted transparent background hand image;
[0127] Step 813: For any one of the multiple background images, adjust the brightness of the distorted transparent background hand image according to the median brightness value of the first target region in the any one background image to obtain a first intermediate hand image, wherein the first target region is the region in the any one background image that has the same position coordinates as the hand region in the distorted transparent background hand image.
[0128] Step 814: Based on the transparency layer of the first intermediate hand image, obtain a second intermediate hand image, wherein the second intermediate hand image is an image containing the edge region of the hand;
[0129] Step 815: Combine the second intermediate hand image with any of the background images to obtain a composite image, and perform linear interpolation on the edge region of the hand in the composite image to obtain the target hand image corresponding to any of the background images.
[0130] Step 816: Based on the multiple target hand images, the 3D position coordinates of each hand joint corresponding to the multiple target hand images, and the image position coordinates of each hand joint in the multiple target hand images, a hand dataset is obtained, wherein the 3D position coordinates and image position coordinates of each hand joint are obtained based on the motion trajectory of each hand joint.
[0131] Based on the same inventive concept, the method for generating hand datasets as described above can also be implemented by a hand dataset generation device. The effect of this hand dataset generation device is similar to that of the aforementioned method, and will not be repeated here.
[0132] Figure 9 This is a schematic diagram of a device for generating a hand dataset according to an embodiment of the present disclosure.
[0133] like Figure 9 As shown, the hand dataset generation apparatus 900 of this disclosure may include a receiving module 910, a camera motion trajectory determination module 920, a virtual trajectory determination module 930, a human hand green screen image determination module 940, an image synthesis module 950, and a human hand dataset determination module 960.
[0134] The receiving module 910 is used to receive the motion trajectory of the head-mounted display and the motion trajectory of each joint of the hand sent by the motion capture system; wherein, the motion trajectory is composed of multiple poses in the world coordinate system;
[0135] The camera motion trajectory determination module 920 is used to transform the motion trajectory of the head-mounted display using each first transformation relationship to obtain the motion trajectory of each camera mounted on the head-mounted display; wherein, any one of the first transformation relationships is the transformation relationship between the head-mounted display and any one of the cameras on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras.
[0136] The virtual trajectory determination module 930 is used to obtain the motion trajectory of the virtual camera based on the motion trajectory of each camera installed on the head-mounted display, and to obtain the motion trajectory of the three-dimensional hand model based on the motion trajectory of each hand joint.
[0137] The green screen image determination module 940 is used to drive the three-dimensional hand model using the motion trajectory of the three-dimensional hand model, and drive the virtual camera to take pictures of the three-dimensional hand model using the motion trajectory of the virtual camera, thereby obtaining multiple green screen images of the hand, wherein the multiple green screen images of the hand are foreground images of the hand used to generate the hand dataset;
[0138] The image compositing module 950 is used to combine the multiple green screen images of the human hand with multiple preset background images to obtain multiple target human hand images; wherein, the multiple background images are obtained by shooting with multiple cameras of the same model;
[0139] The hand dataset determination module 960 is used to obtain a hand dataset based on the multiple target hand images, the 3D position coordinates of each hand joint point corresponding to the multiple target hand images, and the image position coordinates of each hand joint point in the multiple target hand images, wherein the 3D position coordinates and image position coordinates of each hand joint point are obtained based on the motion trajectory of each hand joint point.
[0140] In one embodiment, the camera motion trajectory determination module 920 is specifically used for:
[0141] For any given camera, the first transformation relationship corresponding to that camera is multiplied by multiple poses in the motion trajectory of the head-mounted display to obtain multiple poses of that given camera; and,
[0142] Based on the multiple poses of any one camera, the motion trajectory of any one camera is obtained.
[0143] In one embodiment, the apparatus further includes:
[0144] The calibration board image acquisition module 970 is used to acquire an image of the calibration board taken by any one camera at any two different times for any one camera installed on the head-mounted display before transforming the motion trajectory of the head-mounted display using each first transformation relationship to obtain the motion trajectory of each camera installed on the head-mounted display.
[0145] The second transformation relationship determination module 980 is used to calibrate the camera using images of the calibration board taken by the camera at any two different times, and to obtain each second transformation relationship between the camera and the calibration board at any two different times. The second transformation relationship is used to represent the external parameter transformation relationship between the camera and the calibration board at any time.
[0146] The first transformation relationship determination module 990 is used to obtain the first transformation relationship between any camera and the head-mounted display based on the 3D position coordinates of the head-mounted display in the world coordinate system at any two different times and the second transformation relationships, wherein the 3D position coordinates of the head-mounted display in the world coordinate system at any time are obtained from the motion trajectory of the head-mounted display.
[0147] In one embodiment, the first transformation relationship determination module 990 is specifically used for:
[0148] The first transformation relationship between any one of the cameras and the head-mounted display can be obtained through the following formula:
[0149]
[0150] in, Let be the matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time i. This is the matrix corresponding to the first transformation relationship between the nth camera on the headset and the headset itself. This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time i. Let be the matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time j. This is the matrix corresponding to the second transformation relationship between the nth camera on the head-mounted display and the calibration board at time j.
[0151] In one embodiment, the apparatus further includes:
[0152] The position coordinate determination module 991 is used to determine the position coordinates of each hand joint in the camera coordinate system before combining the multiple green screen images of the human hand and the preset multiple background images to obtain multiple target human hand images. This is done for any one of the green screen images of the human hand, based on the 3D position coordinates of each hand joint in the world coordinate system corresponding to the green screen image of the human hand and the pose of the camera corresponding to the green screen image of the human hand. The 3D position coordinates of each hand joint in the world coordinate system corresponding to the green screen image of the human hand are obtained based on the motion trajectory of each hand joint.
[0153] The deletion module 992 is used to delete any green screen image of a human hand if the position coordinates of each hand joint point in the camera coordinate system do not meet the specified conditions.
[0154] The first judgment module 993 is used to, if the position coordinates of each hand joint in the camera coordinate system satisfy the specified conditions, project each hand joint using the intrinsic parameters of the camera to obtain the image position coordinates of each hand joint; and determine whether each hand joint is located inside any green screen image of a human hand based on the image position coordinates of each hand joint. If so, retain any green screen image of a human hand; otherwise, delete any green screen image of a human hand.
[0155] In one embodiment, the position coordinate determination module 991 is specifically used for:
[0156] Based on the 3D position coordinates of each hand joint corresponding to any green screen image of a human hand in the world coordinate system, the target pose matrix corresponding to each hand joint is obtained.
[0157] Multiply the matrix formed by the pose of the camera corresponding to any green screen image of a human hand with the target pose matrix corresponding to each hand joint to obtain the position coordinates of each hand joint in the camera coordinate system.
[0158] In one embodiment, the apparatus further includes:
[0159] The second judgment module 994 is used to determine whether the position coordinates of each hand joint point in the camera coordinate system meet the specified conditions by means of the following method:
[0160] If at least one of the hand joints has a vertical coordinate in the camera coordinate system that is less than a specified value, then the position coordinates of the hand joints in the camera coordinate system are determined not to meet the specified condition; or,
[0161] If none of the hand joints has a vertical coordinate in the camera coordinate system that is less than the specified value, then the position coordinates of the hand joints in the camera coordinate system are determined to satisfy the specified condition.
[0162] In one embodiment, the image synthesis module 950 is specifically used for:
[0163] For any green screen image of a human hand, replace the green screen background of the green screen image with a transparent background to obtain a human hand image with a transparent background;
[0164] The transparent background hand image is distorted using preset distortion parameters to obtain a distorted transparent background hand image.
[0165] The distorted transparent background hand image is combined with the multiple background images to obtain the multiple target hand images, wherein the distorted transparent background hand image has the same size as any of the background images.
[0166] In one embodiment, the image synthesis module 950 performs image synthesis by combining the distorted transparent background hand image with the multiple background images to obtain the multiple target hand images, specifically for:
[0167] For any one of the multiple background images, the brightness of the distorted transparent background hand image is adjusted according to the median brightness value of the first target region in the any one background image to obtain a first intermediate hand image, wherein the first target region is the region in the any one background image that has the same position coordinates as the hand region in the distorted transparent background hand image.
[0168] Based on the transparency layer of the first intermediate hand image, a second intermediate hand image is obtained, wherein the second intermediate hand image is an image containing the edge region of the hand;
[0169] The second intermediate hand image is combined with any of the background images to obtain a composite image. The edge region of the hand in the composite image is linearly interpolated to obtain the target hand image corresponding to any of the background images.
[0170] After introducing a method and apparatus for generating a hand dataset according to an exemplary embodiment of the present invention, the following describes an electronic device according to another exemplary embodiment of the present invention.
[0171] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as "circuit", "module", or "system".
[0172] In some possible implementations, the electronic device according to the present invention may include at least one processor and at least one computer storage medium. The computer storage medium stores program code that, when executed by the processor, causes the processor to perform the steps in the method for generating hand datasets according to various exemplary embodiments of the present invention described above. For example, the processor may perform actions such as... Figure 1 Steps 101-106 are shown in the diagram.
[0173] The following reference Figure 10 To describe an electronic device 1000 according to this embodiment of the present invention. Figure 10 The electronic device 1000 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0174] like Figure 10 As shown, the electronic device 1000 is manifested in the form of a general electronic device. The components of the electronic device 1000 may include, but are not limited to: at least one processor 1001, at least one computer storage medium 1002, and a bus 1003 connecting different system components (including the computer storage medium 1002 and the processor 1001).
[0175] Bus 1003 represents one or more of several bus structures, including computer storage media bus or computer storage media controller, peripheral bus, processor, or local bus using any of the various bus structures.
[0176] Computer storage medium 1002 may include readable media in the form of volatile computer storage media, such as random access computer storage medium (RAM) 1021 and / or cache storage medium 1022, and may further include read-only computer storage medium (ROM) 1023.
[0177] The computer storage medium 1002 may also include a program / utility 1025 having a set (at least one) of program modules 1024, such program modules 1024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0178] Electronic device 1000 can also communicate with one or more external devices 1004 (e.g., keyboard, pointing device, etc.), one or more devices that enable a user to interact with electronic device 1000, and / or any device that enables electronic device 1000 to communicate with one or more other AR devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1005. Furthermore, electronic device 1000 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1006. As shown, network adapter 1006 communicates with other modules used in electronic device 1000 via bus 1003. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1000, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0179] In some possible implementations, various aspects of the method for generating a hand dataset provided by the present invention can also be implemented in the form of a program product, which includes program code that, when the program product is run on a computer device, causes the computer device to perform the steps in the method for generating a hand dataset according to various exemplary embodiments of the present invention described above.
[0180] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method of generating a hand dataset, the method comprising: The method comprises: receiving a motion trajectory of a head-mounted display and a motion trajectory of each hand joint in a human hand sent by a motion capture system; wherein the motion trajectory is composed of a plurality of poses in a world coordinate system; transforming the motion trajectory of the head-mounted display by using each first transformation relationship to obtain a motion trajectory of each camera mounted on the head-mounted display; wherein any one first transformation relationship is a transformation relationship between the head-mounted display and any one camera on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras; obtaining a motion trajectory of a virtual camera according to the motion trajectory of each camera mounted on the head-mounted display, and obtaining a motion trajectory of a three-dimensional hand model according to the motion trajectory of each hand joint; driving the three-dimensional hand model by using the motion trajectory of the three-dimensional hand model, and driving the virtual camera by using the motion trajectory of the virtual camera to take photos of the three-dimensional hand model to obtain a plurality of hand green screen images of a human hand, wherein the plurality of hand green screen images are foreground images of a human hand for generating a hand dataset; performing image synthesis on the plurality of hand green screen images and a plurality of preset background images to obtain a plurality of target hand images; wherein the plurality of background images are obtained by a plurality of cameras of the same model; obtaining a hand dataset based on the plurality of target hand images, 3D position coordinates of each hand joint corresponding to the plurality of target hand images, and image position coordinates of each hand joint in the plurality of target hand images, wherein the 3D position coordinates and the image position coordinates of each hand joint are obtained based on the motion trajectory of each hand joint.
2. The method of claim 1, wherein, The method further comprises: for any one camera, multiplying the first transformation relationship corresponding to the any one camera and each pose in the motion trajectory of the head-mounted display to obtain a plurality of poses of the any one camera; and obtaining the motion trajectory of the any one camera based on the plurality of poses of the any one camera.
3. The method according to claim 1 or 2, characterized in that, Before the transformation of the motion trajectory of the head-mounted display by using each first transformation relationship to obtain the motion trajectory of each camera mounted on the head-mounted display, the method further comprises: for any one camera mounted on the head-mounted display, acquiring images of a calibration board captured by the any one camera at any two different time instants; calibrating the camera by using the images of the calibration board captured by the any one camera at any two different time instants to obtain each second transformation relationship between the any one camera and the calibration board at the any two different time instants, wherein any one second transformation relationship is used to represent an extrinsic conversion relationship between the camera and the calibration board at any one time instant; Based on the 3D position coordinates of the head-mounted display in the world coordinate system at any two different times and the second transformation relationships, the first transformation relationship between any camera and the head-mounted display is obtained, wherein the 3D position coordinates of the head-mounted display in the world coordinate system at any time are obtained from the motion trajectory of the head-mounted display.
4. The method of claim 3, wherein, The step of obtaining the first transformation relationship between any camera and the head-mounted display based on the 3D position coordinates of the head-mounted display in the world coordinate system at any two different times and the second transformation relationships includes: The first transformation relationship between any one of the cameras and the head-mounted display can be obtained through the following formula: wherein, is a matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time i, is a matrix corresponding to the first transformation relationship between the n th camera on the head-mounted display and the head-mounted display, is a matrix corresponding to the second transformation relationship between the n th camera on the head-mounted display and the calibration board at time i, is a matrix corresponding to the 3D position coordinates of the head-mounted display in the world coordinate system at time j, is a matrix corresponding to the second transformation relationship between the n th camera on the head-mounted display and the calibration board at time j.
5. The method of claim 1, wherein, Before combining the multiple green screen images of human hands with multiple preset background images to obtain multiple target human hand images, the method further includes: For any one of the multiple green screen images of a human hand, based on the 3D position coordinates of each hand joint in the world coordinate system corresponding to the green screen image of a human hand and the pose of the camera corresponding to the green screen image of a human hand, the position coordinates of each hand joint in the camera coordinate system are obtained, wherein the 3D position coordinates of each hand joint in the world coordinate system corresponding to the green screen image of a human hand are obtained based on the motion trajectory of each hand joint; If the position coordinates of each hand joint point in the camera coordinate system do not meet the specified conditions, then any one of the green screen images of the human hand will be deleted. If the position coordinates of each hand joint in the camera coordinate system satisfy the specified conditions, then the camera's intrinsic parameters are used to project each hand joint to obtain the image position coordinates of each hand joint; and based on the image position coordinates of each hand joint, it is determined whether each hand joint is located inside any one of the green screen images of a human hand. If so, the green screen image of the human hand is retained; otherwise, the green screen image of the human hand is deleted.
6. The method of claim 5, wherein, The step of obtaining the position coordinates of each hand joint in the camera coordinate system based on the 3D position coordinates of each hand joint in the world coordinate system corresponding to any given green screen image of a human hand and the pose of the camera corresponding to that given green screen image of a human hand includes: Based on the 3D position coordinates of each hand joint corresponding to any green screen image of a human hand in the world coordinate system, the target pose matrix corresponding to each hand joint is obtained. Multiply the matrix formed by the pose of the camera corresponding to any green screen image of a human hand with the target pose matrix corresponding to each hand joint to obtain the position coordinates of each hand joint in the camera coordinate system.
7. The method of claim 5, wherein, The position coordinates of each hand joint in the camera coordinate system are determined to meet the specified conditions using the following method: If at least one of the hand joints has a vertical coordinate in the camera coordinate system that is less than a specified value, then it is determined that the position coordinates of the hand joints in the camera coordinate system do not meet the specified condition. or, If a vertical coordinate of a position coordinate of each hand joint node in the camera coordinate system is less than the specified value, it is determined that the position coordinate of each hand joint node in the camera coordinate system satisfies the specified condition.
8. The method of claim 1, wherein, The image synthesis of the plurality of hand green screen images and the plurality of preset background images comprises: For any one of the hand green screen images, the green screen background of the any one of the hand green screen images is replaced by a transparent background to obtain a transparent background hand image; The transparent background hand image is subjected to distortion processing by using preset distortion parameters to obtain a distorted transparent background hand image; The distorted transparent background hand image and the plurality of background images are subjected to image synthesis respectively to obtain the plurality of target hand images, wherein the distorted transparent background hand image has the same size as any one of the background images.
9. The method of claim 8, wherein, The image synthesis of the plurality of hand green screen images and the plurality of preset background images comprises: For any one of the plurality of background images, the brightness of the distorted transparent background hand image is adjusted according to the brightness median value of a first target region in the any one of the background images to obtain a first intermediate hand image, wherein the first target region is a region in the any one of the background images having the same position coordinates as a hand region in the distorted transparent background hand image; Based on a transparency layer of the first intermediate hand image, a second intermediate hand image is obtained, wherein the second intermediate hand image is an image containing an edge region of the hand; The second intermediate hand image and the any one of the background images are subjected to image synthesis to obtain a synthesis image, and a linear interpolation is performed on an edge region of the hand in the synthesis image to obtain a target hand image corresponding to the any one of the background images.
10. An electronic device, comprising: The device comprises a processor and a memory, and the processor and the memory are connected through a bus; The memory stores a computer program, and the processor is configured to execute the following operations based on the computer program: receive the motion trajectory of the head-mounted display and the motion trajectory of each hand joint node in the hand sent by the motion capture system; wherein the motion trajectory is composed of a plurality of poses in the world coordinate system; transform the motion trajectory of the head-mounted display by using each first transformation relationship to obtain the motion trajectory of each camera installed on the head-mounted display; wherein any one of the first transformation relationships is the transformation relationship between the head-mounted display and any one of the cameras on the head-mounted display, and the number of the first transformation relationships is the same as the number of the cameras; obtain the motion trajectory of the virtual camera according to the motion trajectory of each camera installed on the head-mounted display, and obtain the motion trajectory of the three-dimensional hand model according to the motion trajectory of each hand joint node; The three-dimensional hand model is driven by a motion trajectory of the three-dimensional hand model, and the three-dimensional hand model is photographed by the virtual camera driven by a motion trajectory of the virtual camera to obtain a plurality of hand green screen images, wherein the plurality of hand green screen images are hand foreground images of a generated hand dataset; The plurality of hand green screen images and a plurality of preset background images are image synthesized to obtain a plurality of target hand images; wherein the plurality of background images are obtained by a plurality of cameras of the same model. Based on the plurality of target hand images, 3D position coordinates corresponding to the plurality of target hand images and the plurality of target hand images, and image position coordinates of the plurality of hand joint nodes in the plurality of target hand images, a hand dataset is obtained, wherein the 3D position coordinates and the image position coordinates of the plurality of hand joint nodes are obtained based on a motion trajectory of the plurality of hand joint nodes.