Method for generating hand image, method for training hand pose estimation model
By replacing parts of hand images with corresponding parts from differently posed images, the method addresses inefficiencies in hand image collection, enriching datasets and improving hand pose estimation model precision.
Patent Information
- Application Number
- CN202310098242.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-20
AI Technical Summary
The prior art is difficult to quickly expand the amount of image data during the hand image collection process, and the training data of the hand posture estimation model is insufficient.
By determining the first finger image from the first hand image and determining the corresponding second finger image from the second hand image, a third hand image is generated based on the second hand image and a third hand image is generated, and the generated image is used to train the hand posture estimation model.
It reduces the difficulty of collecting hand images, quickly expands the amount of image data, and improves the training data richness and prediction accuracy of the hand posture estimation model.
Smart Images

Figure CN116052217B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and particularly to technologies such as computer vision, augmented reality, virtual reality, deep learning, etc., and can be applied to scenarios such as the metaverse, virtual digital humans, etc. Background Art
[0002] The collection of hand images depends on manually using a multi-view color (rgb) camera or a stereo color (rgbd) camera to photograph a person's hand. Summary of the Invention
[0003] The present disclosure provides a method for generating hand images, a method for training a hand pose estimation model, an apparatus, a device, and a storage medium.
[0004] According to one aspect of the present disclosure, there is provided a method for generating hand images, including:
[0005] Determining a first finger part image from a first hand image; wherein, the first finger part image at least includes a partial finger region of a human hand;
[0006] Determining a second finger part image corresponding to the first finger part image from a second hand image; wherein, the postures of the human hands in the first finger part image and the second finger part image are different; and
[0007] Based on the second hand image, replacing the second finger part image with the first finger part image to generate a third hand image.
[0008] According to another aspect of the present disclosure, there is provided a method for training a hand pose estimation model, including:
[0009] Using the third hand image generated by any method in the embodiments of the present disclosure to train a first hand pose estimation model to obtain a second hand pose estimation model.
[0010] According to another aspect of the present disclosure, there is provided a device for generating hand images, including:
[0011] A first image determination device for determining a first finger part image from a first hand image; wherein, the first finger part image at least includes a partial finger region of a human hand;
[0012] A second image determination device for determining a second finger part image corresponding to the first finger part image from a second hand image; wherein, the postures of the human hands in the first finger part image and the second finger part image are different; and
[0013] A generation module for replacing the second finger part image with the first finger part image based on the second hand image to generate a third hand image.
[0014] According to another aspect of the present disclosure, there is provided a training apparatus for a hand pose estimation model, including:
[0015] A training module, configured to train a first hand pose estimation model by using the third hand image generated by any one of the apparatuses in the embodiments of the present disclosure, so as to obtain a second hand pose estimation model.
[0016] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any one of the methods in the embodiments of the present disclosure.
[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any one of the methods in the embodiments of the present disclosure.
[0021] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements any one of the methods in the embodiments of the present disclosure.
[0022] In the embodiments of the present disclosure, generating a new hand image by using an existing hand image can reduce the difficulty of collecting hand images and quickly expand the amount of image data.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0025] Figure 1 is a schematic flowchart of a method for generating a hand image according to an embodiment of the present disclosure;
[0026] Figure 2a is a schematic diagram of a second hand image according to an embodiment of the present disclosure;
[0027] Figure 2b is a schematic diagram of a first hand image according to an embodiment of the present disclosure;
[0028] Figure 2cIt is a schematic diagram of a third hand image according to an embodiment of the present disclosure;
[0029] Figure 3 It is a schematic flowchart of a method for training a hand pose estimation model according to an embodiment of the present disclosure;
[0030] Figure 4 It is a schematic structural diagram of a generating device for hand images according to an embodiment of the present disclosure;
[0031] Figure 5 It is a schematic structural diagram of a method for training a hand pose estimation model according to an embodiment of the present disclosure;
[0032] Figure 6 It is a block diagram of an electronic device for implementing the method of the embodiment of the present disclosure. Detailed implementation manners
[0033] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0034] Figure 1 It is a schematic flowchart of a method for generating hand images according to an embodiment of the present disclosure. The method may include the following steps:
[0035] S101. Determine a first finger image from the first hand image. The first finger image at least includes a partial finger region of a human hand.
[0036] S102. Determine a second finger image corresponding to the first finger image from the second hand image. The postures of the human hands in the first finger image and the second finger image are different.
[0037] In the embodiments of the present disclosure, the first hand image and the second hand image can be understood as existing hand images, and the first hand image and the second hand image include a human hand region. The first hand image and the second hand image may include depth information. The first hand image can also be understood as an image for local replacement, and correspondingly, the second hand image can also be understood as an image to be replaced.
[0038] The first finger part image and the second finger part image at least include partial finger regions of a human hand. The partial finger regions can be several fingers or several finger joints. The partial finger regions included in the first finger part image and the second finger part image are the same to ensure that the resulting human hand after local replacement is still normal. For example, the first finger part image includes the image corresponding to the entire middle finger in the first hand image. Correspondingly, the second finger part image includes the image corresponding to the entire middle finger in the second hand image. Another example is that the first finger part image includes the images corresponding to the end joints of the entire index finger and thumb in the first hand image. Correspondingly, the second finger part image includes the images corresponding to the index finger part and the end joint of the thumb in the second hand image.
[0039] S103. Based on the second hand image, replace the second finger part image with the first finger part image to generate a third hand image.
[0040] In the embodiments of the present disclosure, based on the second hand image, the second finger part image is replaced with the first finger part image to generate a third hand image. The third hand image can be understood as a new hand image generated by local replacement according to the existing first hand image and second hand image. Since the postures of the human hands in the first finger part image and the second finger part image are different, after local replacement, the posture of the human hand in the obtained third hand image is different from the postures of the human hands in the first hand image and the second hand image.
[0041] In one example, as Figure 2a shown, the second hand image includes a human hand in a fist posture (five fingers retracted and closed together). As Figure 2b shown, the first hand image includes a human hand with five fingers spread apart. The first finger part image is the image of the index finger and middle finger of the first hand image. Similarly, the second finger part image is the image of the index finger and middle finger in the second hand image. Based on the second hand image, replacing the second finger part image with the first finger part image to generate a third hand image means replacing the retracted and closed index finger and middle finger (second finger part image) with the spread index finger and middle finger (first finger part image) on the human hand in the fist posture (second hand image). The generated third hand image is a human hand with only the index finger and middle finger extended and spread apart, as Figure 2c shown, that is, the gesture of making the number 2. Thus, a third hand image with a new hand posture is obtained.
[0042] It should be noted that the first hand image, the second hand image, and the third hand image can be used for model training. For example, training a hand posture estimation model.
[0043] According to the solution of the embodiments of the present disclosure, using existing hand images to generate new hand images can reduce the difficulty of collecting hand images and quickly expand the data volume of images.
[0044] In a possible implementation, step S101, determining a first finger image from a first hand image, may further include the steps of:
[0045] S1011. Determine a first hand image from a hand image dataset.
[0046] S1012. Determine a partial image of at least one finger or a partial image of at least one finger joint from the first hand image to obtain a first finger image.
[0047] In the embodiments of the present disclosure, the hand image dataset includes multiple first hand images. The manner of determining a first hand image from the hand image dataset may be random. The manner of determining a partial image of at least one finger or a partial image of at least one finger joint from the first hand image may also be random.
[0048] It should be noted that the second hand image may also be randomly selected from the hand image dataset. However, in order to ensure that the postures of the human hands in the first finger image and the second finger image of the second hand image are different, therefore, after the randomly determined second hand image, if the posture of the second finger image is the same as that of the first finger image, the second hand image is randomly selected again.
[0049] According to the solution of the embodiments of the present disclosure, partial fingers or partial finger joints of a human hand can be replaced, so that the hand postures of the generated third hand image are more diverse.
[0050] In a possible implementation, step S1012, determining a partial image of at least one finger or a partial image of at least one finger joint from the first hand image to obtain a first finger image, includes:
[0051] Perform image segmentation on the first hand image to obtain partial images of multiple fingers of the first hand image.
[0052] Determine a partial image of at least one finger or a partial image of at least one finger joint from the partial images of multiple fingers to obtain a first finger image.
[0053] In the embodiments of the present disclosure, an image segmentation model may be used to perform image segmentation on the first hand image to obtain partial images of multiple fingers of the first hand image. The image segmentation model can be trained using conventional model training methods and training samples. The training samples can be obtained by annotating each region of the human hand in some hand images in the hand image dataset.
[0054] It should be noted that the partial images of multiple fingers of the first hand image can be stored in the hand image dataset.
[0055] According to the solution of the embodiment of the present disclosure, the first hand image can be pre-segmented, which is convenient for randomly selecting a first finger image from the local images of multiple fingers included in the first hand image, thereby improving the efficiency of generating a hand image.
[0056] In a possible implementation manner, step S103, based on the second hand image, replacing the second finger image with the first finger image to generate a third hand image may further include the steps of:
[0057] S1031. Determine the first hand pose information of the first hand image and the second hand pose information of the second hand image.
[0058] In the embodiment of the present disclosure, the first hand pose information and the second hand pose information can be understood to at least include the rotation angles of each joint of the human hand. It may also include the positions of each joint and the lengths of each finger phalanx.
[0059] The first hand pose information can be obtained by predicting the first hand image through a human hand pose estimation model, and the second hand pose information is the same. It is also possible to manipulate the hand three-dimensional model generated in the renderer to make its pose consistent with the human hand pose in the image, and then record the rotation angles of each joint.
[0060] S1032. Perform pose fusion on the first hand pose information and the second hand pose information to obtain third hand pose information.
[0061] Specifically, it may be to replace the pose information corresponding to the second finger image in the second hand pose information with the pose information corresponding to the first finger image in the first hand pose information to obtain the third hand pose information.
[0062] S1033. According to the third hand pose information and the second hand image, replace the second finger image with the first finger image to generate a third hand image.
[0063] In the embodiment of the present disclosure, according to the third hand pose information, on the basis of the second hand image, the second finger image is replaced with the first finger image. Since the third hand pose information includes the pose of the first finger image, when the first finger image is replaced on the second hand image, its original hand pose is retained.
[0064] According to the solution of the embodiment of the present disclosure, when replacing the finger image, the pose of the first finger image is retained, thereby changing the hand pose of the second hand image and generating a third hand image with a hand pose different from that of the first hand image and the second hand image, improving the richness of the hand pose of the hand image.
[0065] In a possible implementation, step S1032, performing pose fusion on the first hand pose information and the second hand pose information to obtain third hand pose information, includes:
[0066] Generating a three-dimensional hand model with a preset orientation according to the second hand pose information.
[0067] Updating the three-dimensional hand model according to the pose information corresponding to the first finger image in the first hand pose information to achieve pose fusion.
[0068] Determining the third hand pose information according to the updated three-dimensional hand model.
[0069] In the embodiments of the present disclosure, in a renderer, a three-dimensional hand model with a preset orientation can be generated according to the second hand pose information. The preset orientation can be that the palm direction is forward or backward, so as to facilitate the observation of the position and pose of the five fingers when artificial participation is involved. Specifically, it can be achieved by setting the wrist rotation angle in the second hand pose information.
[0070] It should be noted that when updating the three-dimensional hand model according to the pose information corresponding to the first finger image in the first hand pose information, the first hand pose can also be set to the preset orientation to facilitate the fusion of the two hand pose information.
[0071] According to the solution of the embodiments of the present disclosure, the fusion of the first hand pose information and the second hand pose information is achieved.
[0072] In a possible implementation, the method of the embodiments of the present disclosure further includes the steps of:
[0073] Determining the invisible area information of the occluded finger according to the front-back relationship of the fingers of the three-dimensional hand model in the target orientation.
[0074] Setting a mask on the first hand image and / or the second hand image according to the invisible area information.
[0075] In the embodiments of the present disclosure, the target orientation can be the orientation corresponding to the original shooting angle of the second hand image.
[0076] In an example, after replacing the second finger image with the first finger image, in the target orientation, the thumb changes from the original position (not occluding other fingers) to in front of the index finger, thus occluding a part of the index finger area. Therefore, when generating the third hand image, a mask needs to be set on the second hand image so that the occluded part of the index finger is invisible in the third hand image.
[0077] According to the solution of the embodiments of the present disclosure, by setting a mask on the original hand image, it is ensured that the front-back relationship between the fingers shown in the third hand image is correct.
[0078] In a possible implementation, step S1033, replacing the second finger image with the first finger image according to the third hand pose information and the second hand image to generate a third hand image, further includes the steps of:
[0079] S1033a. When it is determined according to the third hand pose information that finger conflict will occur, changing the first rotation angle of at least one finger joint corresponding to the finger conflict in the third hand pose information to obtain a second rotation angle capable of eliminating the finger conflict. Wherein, finger conflict includes that the target area image of the first hand image coincides with other hand area images except the target area image of the second hand image in three-dimensional space.
[0080] Finger conflict can be understood as that at least two fingers penetrate each other in three-dimensional space. For example, the first hand image is a gesture with the index finger and middle finger forming a V shape (the index finger is tilted towards the thumb direction, and the middle finger is tilted towards the ring finger direction), and the second hand image is a posture with five fingers side by side and extended. When replacing the middle finger of the first hand image with the second hand image, since the middle finger of the first hand image is tilted towards the ring finger direction, it is very likely to coincide with the ring finger in the second hand image in three-dimensional space.
[0081] In the case of penetration, the posture of the penetrated finger can be adjusted. Specifically, by setting the hand pose information, the penetrated finger can be rotated by a certain angle in a certain direction, so that the two fingers are separated in three-dimensional space.
[0082] S1033b. Updating the third hand pose information according to the second rotation angle to obtain a fourth hand pose information.
[0083] S1033c. Replacing the second finger image with the first finger image according to the fourth hand pose information and the second hand image to generate a third hand image.
[0084] It should be noted that limit conditions can be set for the rotation angle of each joint according to the human body structure to avoid finger bending angles that ordinary people cannot achieve.
[0085] According to the solution of the embodiment of the present disclosure, by correcting the posture of the penetrated finger, it effectively avoids the hand postures that do not exist in reality when generating the third hand image.
[0086] In a possible implementation, step S1033a. When it is determined according to the third hand pose information that finger conflict will occur, changing the first rotation angle of at least one finger phalanx corresponding to the finger conflict in the third hand pose information to obtain a second rotation angle capable of eliminating the finger conflict, includes:
[0087] In the case where finger conflicts are determined according to the third hand gesture information, determine the fingers to be adjusted from at least two fingers where finger conflicts occur.
[0088] On the hand three-dimensional model generated according to the third hand gesture information, change the first rotation angle of at least one finger joint of the finger to be adjusted to obtain a second rotation angle that can eliminate finger conflicts.
[0089] In the embodiments of the present disclosure, the fingers to be adjusted can be determined from at least two fingers where finger conflicts occur. The finger to be adjusted can be one, or all the fingers where finger conflicts occur. For example, in the direction of the palm, the root joint of one of the fingers where penetration occurs can be rotated forward or backward. It can also be that the two fingers where finger conflicts occur are rotated forward and backward respectively. For another example, if the finger joints at the outermost ends of two fingers have penetration, only the rotation angle of this finger joint can be adjusted.
[0090] The rotation angle of the finger can be adjusted by using the hand three-dimensional model. Using the hand three-dimensional model, it is possible to determine whether the fingers overlap according to the three-dimensional volume of the hand three-dimensional model.
[0091] According to the solution of the embodiments of the present disclosure, correct the front-back relationship of the fingers where penetration occurs to obtain a hand image that correctly expresses depth.
[0092] The third hand image obtained by the generation method according to any of the above embodiments can be used in various scenarios, such as the training of a hand gesture estimation model, the training of a gesture recognition model, or the training of a hand image segmentation model, etc. The following takes the training of a hand gesture estimation model as an example for illustration.
[0093] Figure 3 It is a schematic flowchart of a method for training a hand gesture estimation model according to an embodiment of the present disclosure. As Figure 3 shown, the method may include:
[0094] S301. Use the third hand image generated by the hand image generation method according to any embodiment to train the first hand gesture estimation model to obtain a second hand gesture estimation model.
[0095] According to the solution of the embodiments of the present disclosure, the third hand image obtained by the hand image generation method increases the data volume on the basis of the existing hand images, can provide rich hand gestures including abnormal hand data, and helps to improve the prediction accuracy of the hand gesture estimation model.
[0096] Figure 4 It is a schematic structural diagram of a hand image generation device according to an embodiment of the present disclosure. AsFigure 4 As shown, the device may include:
[0097] A first image determination device 401, configured to determine a first finger part image from a first hand image, where the first finger part image at least includes a partial finger area of a human hand.
[0098] A second image determination device 402, configured to determine a second finger part image corresponding to the first finger part image from a second hand image. Among them, the postures of the human hands in the first finger part image and the second finger part image are different. And
[0099] A generation module 403, configured to generate a third hand image by replacing the second finger part image with the first finger part image based on the second hand image.
[0100] In a possible implementation, the first image determination device 401 is configured to:
[0101] A first determination sub-module, configured to determine a first hand image from a hand image dataset.
[0102] A second determination sub-module, configured to determine a local image of at least one finger or a local image of at least one finger joint from the first hand image to obtain the first finger part image.
[0103] In a possible implementation, the second determination sub-module is configured to:
[0104] Perform image segmentation on the first hand image to obtain local images of multiple fingers of the first hand image.
[0105] Determine a local image of at least one finger or a local image of at least one finger joint from the local images of multiple fingers to obtain the first finger part image.
[0106] In a possible implementation, the generation module 403 includes:
[0107] A posture determination sub-module, configured to determine first hand posture information of the first hand image and second hand posture information of the second hand image.
[0108] A posture fusion sub-module, configured to perform posture fusion on the first hand posture information and the second hand posture information to obtain third hand posture information.
[0109] A first replacement sub-module, configured to replace the second finger part image with the first finger part image according to the third hand posture information and the second hand image to generate a third hand image.
[0110] In a possible implementation, the posture fusion sub-module is configured to:
[0111] Generate a three-dimensional hand model in a preset orientation according to the second hand pose information.
[0112] Update the three-dimensional hand model according to the pose information corresponding to the first finger image in the first hand pose information to achieve pose fusion.
[0113] Determine the third hand pose information according to the updated three-dimensional hand model.
[0114] In a possible implementation manner, the device further includes:
[0115] An occlusion determination module, configured to determine the invisible area information of the occluded finger according to the front-back relationship of the fingers of the three-dimensional hand model in the target orientation.
[0116] A masking module, configured to set a mask on the first hand image and / or the second hand image according to the invisible area information.
[0117] In a possible implementation manner, the generation module 403 further includes:
[0118] An adjustment sub-module, configured to change the first rotation angle of at least one finger joint corresponding to the finger conflict in the third hand pose information to obtain a second rotation angle that can eliminate the finger conflict when it is determined according to the third hand pose information that a finger conflict will occur. Wherein, the finger conflict includes that the target region image of the first hand image coincides with other hand region images of the second hand image except the target region image in three-dimensional space.
[0119] An update sub-module, configured to update the third hand pose information according to the second rotation angle to obtain the fourth hand pose information.
[0120] A second replacement sub-module, configured to replace the second finger image with the first finger image according to the fourth hand pose information and the second hand image to generate a third hand image.
[0121] In a possible implementation manner, the adjustment sub-module is configured to:
[0122] Determine the finger to be adjusted from at least two fingers where the finger conflict occurs when it is determined according to the third hand pose information that a finger conflict will occur.
[0123] On the three-dimensional hand model generated according to the third hand pose information, change the first rotation angle of at least one finger joint of the finger to be adjusted to obtain a second rotation angle that can eliminate the finger conflict.
[0124] For the specific functions and example descriptions of the modules and sub-modules of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0125] Figure 5 It is a schematic structural diagram of a training device for a hand gesture estimation model according to an embodiment of the present disclosure. As Figure 5 shown, the device may include:
[0126] A training module 501, configured to train a first hand gesture estimation model by using a third hand image generated by the hand image generation device according to any one of the above embodiments, so as to obtain a second hand gesture estimation model.
[0127] For the specific functions and examples of the modules and sub-modules of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0128] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0129] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0130] Figure 6 Fig. shows a schematic block diagram of an exemplary electronic device 600 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0131] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0132] Multiple components in device 600 are connected to I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0133] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for generating a hand image or the method for training a hand gesture estimation model. For example, in some embodiments, the method for generating a hand image or the method for training a hand gesture estimation model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for generating a hand image or the method for training a hand gesture estimation model described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for generating a hand image or the method for training a hand gesture estimation model in any other suitable manner (e.g., by means of firmware).
[0134] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0135] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0136] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0137] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0138] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0139] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0140] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0141] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating a hand image, comprising: Determining a first finger part image from a first hand image; wherein, the first finger part image at least includes a partial finger area of a human hand; Determining a second finger part image corresponding to the first finger part image from a second hand image; wherein, the postures of the human hands in the first finger part image and the second finger part image are different; Determining first hand posture information of the first hand image and second hand posture information of the second hand image; Generating a hand three-dimensional model with a preset orientation according to the second hand posture information; Updating the hand three-dimensional model according to the posture information corresponding to the first finger part image in the first hand posture information to achieve posture fusion; Determining third hand posture information according to the updated hand three-dimensional model; Replacing the second finger part image with the first finger part image according to the third hand posture information and the second hand image to generate a third hand image.
2. The method according to claim 1, wherein Determining a first finger part image from a first hand image includes: Determining a first hand image from a hand image dataset; Determining a local image of at least one finger or a local image of at least one finger joint from the first hand image to obtain a first finger part image.
3. The method according to claim 2, wherein Determining a local image of at least one finger or a local image of at least one finger joint from the first hand image to obtain a first finger part image, including: Performing image segmentation on the first hand image to obtain local images of multiple fingers of the first hand image; Determining a local image of at least one finger or a local image of at least one finger joint from the local images of the multiple fingers to obtain a first finger part image.
4. The method according to claim 1, further comprising: Determining invisible area information of an occluded finger according to the finger front-back relationship of the hand three-dimensional model in a target orientation; Setting a mask on the first hand image and / or the second hand image according to the invisible area information.
5. The method according to claim 1, wherein Replacing the second finger part image with the first finger part image according to the third hand posture information and the second hand image to generate a third hand image, including: In the case where it is determined according to the third hand posture information that finger conflicts will occur, changing a first rotation angle of at least one finger joint corresponding to the finger conflicts in the third hand posture information to obtain a second rotation angle capable of correcting the finger conflicts; wherein, the finger conflicts include that the target area image of the first hand image coincides with other hand area images of the second hand image in three-dimensional space; Updating the third hand posture information according to the second rotation angle to obtain fourth hand posture information; Replacing the second finger part image with the first finger part image according to the fourth hand posture information and the second hand image to generate a third hand image.
6. The method according to claim 5, wherein, When it is determined according to the third hand gesture information that finger conflicts will occur, changing the first rotation angle of at least one finger joint corresponding to the finger conflicts in the third hand gesture information to obtain a second rotation angle capable of correcting the finger conflicts, including: When it is determined according to the third hand gesture information that finger conflicts will occur, determining the fingers to be adjusted from at least two fingers where the finger conflicts occur; On the hand three-dimensional model generated according to the third hand gesture information, changing the first rotation angle of at least one finger joint of the finger to be adjusted to obtain a second rotation angle capable of correcting the finger conflicts.
7. A method for training a hand gesture estimation model, including: Training a first hand gesture estimation model by using the third hand image generated by the method according to any one of claims 1 to 6 to obtain a second hand gesture estimation model.
8. A device for generating a hand image, including: A first image determination device, configured to determine a first finger image from a first hand image; wherein, the first finger image at least includes a partial finger region of a human hand; A second image determination device, configured to determine a second finger image corresponding to the first finger image from a second hand image; wherein, the postures of the human hands in the first finger image and the second finger image are different; and A generation module, including: A gesture determination sub-module, configured to determine the first hand gesture information of the first hand image and the second hand gesture information of the second hand image; A gesture fusion sub-module, configured to generate a hand three-dimensional model in a preset orientation according to the second hand gesture information; Updating the hand three-dimensional model according to the gesture information corresponding to the first finger image in the first hand gesture information to achieve gesture fusion; Determining third hand gesture information according to the updated hand three-dimensional model; A first replacement sub-module, configured to replace the second finger image with the first finger image according to the third hand gesture information and the second hand image to generate a third hand image.
9. The device according to claim 8, wherein, The first image determination device is used for: A first determination sub-module, configured to determine a first hand image from a hand image dataset; A second determination sub-module, configured to determine a local image of at least one finger or a local image of at least one finger joint from the first hand image to obtain a first finger image.
10. The device according to claim 9, wherein, The second determination sub-module is used for: Performing image segmentation on the first hand image to obtain local images of multiple fingers of the first hand image; Determining a local image of at least one finger or a local image of at least one finger joint from the local images of the multiple fingers to obtain a first finger image.
11. The device according to claim 10, further including: An occlusion determination module, configured to determine invisible region information of the occluded finger according to the front-back relationship of the fingers when the hand three-dimensional model is in the target orientation; A masking module, configured to set a mask on the first hand image and / or the second hand image according to the invisible region information.
12. The apparatus according to claim 9, wherein, The generation module, including: An adjustment sub-module, configured to change a first rotation angle of at least one finger joint corresponding to a finger conflict in the third hand pose information to obtain a second rotation angle capable of correcting the finger conflict when it is determined according to the third hand pose information that a finger conflict will occur; wherein, the finger conflict includes that the target region image of the first hand image coincides with other hand region images of the second hand image except the target region image in three-dimensional space; An update sub-module, configured to update the third hand pose information according to the second rotation angle to obtain a fourth hand pose information; A second replacement sub-module, configured to replace the second finger image with the first finger image according to the fourth hand pose information and the second hand image to generate a third hand image.
13. The apparatus according to claim 12, wherein, The adjustment sub-module is configured to: When it is determined according to the third hand pose information that a finger conflict will occur, determine a finger to be adjusted from at least two fingers with finger conflicts; On the hand three-dimensional model generated according to the third hand pose information, change a first rotation angle of at least one finger phalanx of the finger to be adjusted to obtain a second rotation angle capable of correcting the finger conflict.
14. A training device for a hand pose estimation model, comprising: A training module, configured to train a first hand pose estimation model by using the third hand image generated by the device according to any one of claims 8 to 13 to obtain a second hand pose estimation model.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
17. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Face transformation method and device, equipment, storage medium and product
CN113658035A
Hand posture recognition method and training method and device of hand posture recognition model
CN115223248A