Image processing method, device, electronic device and storage medium

By obtaining and optimizing the posture detection data of virtual images, combining the posture detection data of body and head, the problem of unnatural virtual image driving effect is solved, and a more natural and stable virtual image driving effect is achieved, reducing costs.

CN115138063BActive Publication Date: 2025-09-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210770420.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-09-02
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

In the prior art, the driving effect of virtual images is not natural and reasonable enough, especially in the posture migration of multiple parts.

Method used

By obtaining posture detection data of multiple parts of the target object, using body and head posture detection data combined with associated data, the posture driving data of each part is determined, and the virtual image is driven based on these data. The two-norm weighting and minimization method is used to optimize the posture driving data to improve the accuracy and stability of posture detection.

Benefits of technology

The naturalness and accuracy of the driving effect of virtual images are improved, the jitter of the driving effect is reduced, the jitter of the driving effect is high, the requirements for the capture device are reduced, and the cost-effective driving of virtual images is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115138063B_ABST
    Figure CN115138063B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method, apparatus, electronic device, and storage medium. The image processing method comprises: obtaining posture detection data for multiple parts of a target object in a current image frame; determining posture drive data for each target part associated with other parts of the multiple parts based on the posture detection data for the target part and its associated parts; and driving the corresponding part of a virtual character based on the determined posture drive data for each target part, wherein the virtual character and the target object have a one-to-one corresponding part. According to the present disclosure, the naturalness and accuracy of the driving effect of the virtual character can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of computer technology, and more particularly, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, avatars are widely used in various scenarios, including social networking, online education, and gaming. To animate avatars, actuation technology is required to transfer the user's posture from captured images to the avatar. However, this actuation technology still suffers from the problem of not providing a natural and reasonable actuation effect for the avatar. Summary of the Invention

[0003] The exemplary embodiments of the present disclosure provide an image processing method, apparatus, electronic device, and storage medium, which can improve the naturalness of the driving effect of a virtual image. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, comprising: acquiring posture detection data of multiple parts of a target object in a current image frame; for each target part associated with other parts among the multiple parts, determining posture driving data of the target part based on the posture detection data of the target part and its associated parts; and driving corresponding parts of a virtual image based on the determined posture driving data of each target part, wherein the virtual image has a one-to-one corresponding part with the target object.

[0005] Optionally, the multiple parts include hands, head and body parts, the posture detection data of the multiple parts include head posture detection data, body posture detection data and hand posture detection data, and the various target parts include body parts and hands, wherein, for each target part associated with other parts among the multiple parts, based on the posture detection data of the target part and its associated parts, the step of determining the posture driving data of the target part includes: determining the posture driving data of the neck among the body parts based on the body posture detection data and the head posture detection data, and determining the posture driving data of the body parts other than the neck based on the body posture detection data; determining the posture driving data of the wrist in the hand based on the hand posture detection data and the body posture detection data, and determining the posture driving data of the fingers in the hand based on the hand posture detection data.

[0006] Optionally, the step of determining the posture driving data of the neck in the body part based on the body posture detection data and the head posture detection data includes: determining the posture driving data of the neck based on the neck posture detection data in the body posture detection data, the head posture detection data and the neck posture association data, wherein the neck posture association data includes at least one of the following items: the posture detection data of the parent joint of the neck joint in the body posture detection data, the posture driving data of the neck determined for the previous image frame, and the posture data of the neck in a preset idle state, wherein the neck idle state is the state that the neck needs to enter when the body part and head of the target object in the current image frame are not detected.

[0007] Optionally, the step of determining the neck posture driving data includes: determining the neck posture driving data by minimizing the weighted sum of the two norms of the distances between the neck posture driving data and each item of neck fusion data, wherein the each item of neck fusion data includes: the neck posture detection data, the head posture detection data, and the neck posture association data.

[0008] Optionally, based on the hand posture detection data and the body posture detection data, the step of determining the posture driving data of the wrist in the hand includes: determining the posture driving data of the wrist based on the wrist posture detection data in the hand posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data, wherein the wrist posture association data includes at least one of the following items: the posture driving data of the wrist determined for the previous image frame, and the posture data of the preset wrist idle state, wherein the wrist idle state is the state that the wrist needs to enter when the hand of the target object in the current image frame is not detected.

[0009] Optionally, the step of determining the posture driving data of the wrist includes: determining the posture driving data of the wrist by minimizing the weighted sum of the bi-norms of the distances between the posture driving data of the wrist and each wrist fusion data, wherein the each wrist fusion data includes: the wrist posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data.

[0010] Optionally, the step of obtaining posture detection data of multiple parts of the target object in the current image frame includes: inputting the current image frame into a body detection model to obtain body posture detection data; inputting the head area image in the current image frame into a head detection model to obtain head posture detection data; inputting the hand area image in the current image frame into a hand detection model to obtain hand posture detection data.

[0011] Optionally, the step of inputting the current image frame into the body detection model to obtain body posture detection data includes: inputting the current image frame into the body detection model to obtain the body posture detection data, neck position data, and wrist position data; wherein the image processing method also includes: according to the neck position data, identifying and cutting out the area where the head is located from the current image frame as the head area image; according to the wrist position data, identifying and cutting out the area where the hand is located from the current image frame as the hand area image.

[0012] Optionally, the step of determining the posture driving data of the body parts other than the neck based on the body posture detection data includes: determining the posture driving data of the body parts other than the neck based on the posture detection data of the body parts other than the neck and the body posture association data in the body posture detection data, wherein the body posture association data includes at least one of the following items: the posture driving data of the body parts other than the neck determined for the previous image frame, and the posture data of the preset body parts in an idle state, wherein the body part idle state is the state that the body parts other than the neck need to enter when the body parts of the target object in the current image frame are not detected.

[0013] Optionally, the step of determining the posture driving data of the body part other than the neck includes: determining the posture driving data of the body part other than the neck by minimizing the weighted sum of the bi-norms of the distances between the posture driving data of the body part other than the neck and each body fusion data, wherein the each body fusion data includes: the posture detection data of the body part other than the neck and the body posture association data.

[0014] Optionally, the step of determining the posture driving data of the fingers in the hand based on the hand posture detection data includes: determining the posture driving data of the fingers based on the finger posture detection data and the finger posture association data in the hand posture detection data, wherein the finger posture association data includes at least one of the following items: the posture driving data of the fingers determined for the previous image frame, and the posture data of the preset finger idle state, wherein the finger idle state is the state that the finger needs to enter when the finger of the target object in the current image frame is not detected.

[0015] Optionally, the step of determining the posture driving data of the finger includes: determining the posture driving data of the finger by minimizing the weighted sum of the bi-norms of the distances between the posture driving data of the finger and each finger fusion data, wherein the each finger fusion data includes: the finger posture detection data and the finger posture association data.

[0016] Optionally, the step of driving the corresponding parts of the virtual image based on the determined posture driving data of each target part includes: based on the posture driving data of the wrist, correcting the posture driving data of the parent joint of the wrist joint in the posture driving data of the body parts other than the neck to obtain the corrected posture driving data of the body parts other than the neck; and driving the corresponding parts of the virtual image based on the corrected posture driving data of the body parts other than the neck, the posture driving data of the fingers, the posture driving data of the neck and the posture driving data of the wrist.

[0017] Optionally, the current image frame is an image frame captured by a monocular camera.

[0018] According to a second aspect of an embodiment of the present disclosure, an image processing device is provided, comprising: a detection data acquisition unit configured to acquire posture detection data of multiple parts of a target object in a current image frame; a drive data acquisition unit configured to determine, for each target part associated with other parts among the multiple parts, posture driving data of the target part and its associated parts; and a drive unit configured to drive corresponding parts of a virtual image based on the determined posture driving data of each target part, wherein the virtual image has a one-to-one corresponding part with the target object.

[0019] Optionally, the multiple parts include hands, head and body parts, the posture detection data of the multiple parts include head posture detection data, body posture detection data and hand posture detection data, and the various target parts include body parts and hands, wherein the driving data acquisition unit is configured to: determine the posture driving data of the neck among the body parts based on the body posture detection data and the head posture detection data, and determine the posture driving data of the body parts other than the neck based on the body posture detection data; determine the posture driving data of the wrist in the hand based on the hand posture detection data and the body posture detection data, and determine the posture driving data of the fingers in the hand based on the hand posture detection data.

[0020] Optionally, the driving data acquisition unit is configured to: determine the neck posture driving data based on the neck posture detection data in the body posture detection data, the head posture detection data and the neck posture association data, wherein the neck posture association data includes at least one of the following items: the posture detection data of the parent joint of the neck joint in the body posture detection data, the posture driving data of the neck determined for the previous image frame, and the posture data of the preset neck idle state, wherein the neck idle state is the state that the neck needs to enter when the body part and head of the target object in the current image frame are not detected.

[0021] Optionally, the driving data acquisition unit is configured to determine the neck posture driving data by minimizing the weighted sum of the two norms of the distances between the neck posture driving data and each item of neck fusion data, wherein the each item of neck fusion data includes: the neck posture detection data, the head posture detection data, and the neck posture association data.

[0022] Optionally, the driving data acquisition unit is configured to determine the wrist posture driving data based on the wrist posture detection data in the hand posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data, wherein the wrist posture association data includes at least one of the following items: the wrist posture driving data determined for the previous image frame, and the posture data of the preset wrist idle state, wherein the wrist idle state is the state that the wrist needs to enter when the hand of the target object in the current image frame is not detected.

[0023] Optionally, the drive data acquisition unit is configured to determine the posture drive data of the wrist by minimizing the weighted sum of the bi-norms of the distances between the posture drive data of the wrist and each wrist fusion data, wherein the each wrist fusion data includes: the wrist posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data.

[0024] Optionally, the detection data acquisition unit is configured to: input the current image frame into the body detection model to obtain body posture detection data; input the head area image in the current image frame into the head detection model to obtain head posture detection data; input the hand area image in the current image frame into the hand detection model to obtain hand posture detection data.

[0025] Optionally, the detection data acquisition unit is configured to: input the current image frame into the body detection model to obtain the body posture detection data, neck position data, and wrist position data; wherein, the detection data acquisition unit is further configured to: identify and cut out the area where the head is located from the current image frame according to the neck position data, as the head area image; identify and cut out the area where the hand is located from the current image frame according to the wrist position data, as the hand area image.

[0026] Optionally, the driving data acquisition unit is configured to: determine the posture driving data of the body parts other than the neck based on the posture detection data of the body parts other than the neck in the body posture detection data and the body posture association data, wherein the body posture association data includes at least one of the following items: the posture driving data of the body parts other than the neck determined for the previous image frame, and the posture data of the preset body parts in an idle state, wherein the body part idle state is the state that the body parts other than the neck need to enter when the body parts of the target object in the current image frame are not detected.

[0027] Optionally, the driving data acquisition unit is configured to determine the posture driving data of the body parts other than the neck by minimizing the weighted sum of the bi-norms of the distances between the posture driving data of the body parts other than the neck and the various body fusion data, wherein the various body fusion data include: the posture detection data of the body parts other than the neck and the body posture association data.

[0028] Optionally, the driving data acquisition unit is configured to: determine the finger posture driving data based on the finger posture detection data and the finger posture association data in the hand posture detection data, wherein the finger posture association data includes at least one of the following items: the finger posture driving data determined for the previous image frame, and the posture data in a preset finger idle state, wherein the finger idle state is the state that the finger needs to enter when the finger of the target object in the current image frame is not detected.

[0029] Optionally, the driving data acquisition unit is configured to determine the posture driving data of the finger by minimizing the weighted sum of the bi-norms of the distances between the posture driving data of the finger and each finger fusion data, wherein the each finger fusion data includes: the finger posture detection data and the finger posture association data.

[0030] Optionally, the driving unit is configured to: correct the posture driving data of the parent joint of the wrist joint in the posture driving data of the body parts other than the neck based on the posture driving data of the wrist to obtain the corrected posture driving data of the body parts other than the neck; and drive the corresponding parts of the virtual image based on the corrected posture driving data of the body parts other than the neck, the posture driving data of the fingers, the posture driving data of the neck and the posture driving data of the wrist.

[0031] Optionally, the current image frame is an image frame captured by a monocular camera.

[0032] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, prompt the at least one processor to execute the image processing method described above.

[0033] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which, when instructions in the computer-readable storage medium are executed by at least one processor, prompts the at least one processor to execute the image processing method as described above.

[0034] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which implement the image processing method described above when executed by at least one processor.

[0035] According to the image processing method, device, electronic device, and storage medium of the exemplary embodiments of the present disclosure, the naturalness and accuracy of the driving effect of the virtual image can be improved. In addition, the problem of jitter in the driving effect can be avoided, and the robustness and stability are high. In addition, according to the image processing method of the exemplary embodiments of the present disclosure, the driving of the virtual image can also be achieved based on the depth-free image captured by the monocular camera. Since no complex external equipment is required, the requirements for the capture device are low, and the implementation cost of the virtual image driving is reduced, thereby achieving a strong applicability and low cost effect.

[0036] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0038] Figure 1 A flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure;

[0039] Figure 2 A flowchart illustrating a method for acquiring posture detection data of multiple parts of a target object in a current image frame according to an exemplary embodiment of the present disclosure is shown;

[0040] Figure 3 A flowchart illustrating a method for determining posture driving data of a target part according to an exemplary embodiment of the present disclosure is shown;

[0041] Figure 4 A block diagram showing a structure of an image processing apparatus according to an exemplary embodiment of the present disclosure is shown;

[0042] Figure 5 A structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0043] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0044] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0045] It should be noted that the phrase "at least one of the items" in this disclosure includes three types of parallel situations: "any one of the items", "a combination of any multiple items of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. For another example, "performing at least one of step 1 and step 2" includes the following three parallel situations: (1) performing step 1; (2) performing step 2; and (3) performing steps 1 and 2.

[0046] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0047] Figure 1 A flowchart illustrating an image processing method according to an exemplary embodiment of the present disclosure is shown.

[0048] Reference Figure 1 In step S101, posture detection data of multiple parts of the target object in the current image frame is obtained.

[0049] As an example, the multiple parts may include: hands, head, and body parts (here, the body parts refer to the parts of the human body other than the head and hands). Accordingly, the posture detection data of the multiple parts may include: hand posture detection data, head posture detection data, and body posture detection data.

[0050] The head posture detection data is the head posture data of the target object obtained by the initial detection, the body posture detection data is the body part posture data of the target object obtained by the initial detection, and the hand posture detection data is the hand posture data of the target object obtained by the initial detection.

[0051] As an example, head posture data may include Euler angles of the head. For example, Euler angles may include at least one of the following: pitch angle α, yaw angle β, and roll angle γ. It should be understood that head posture may also be described using other appropriate methods, such as rotation quaternions. This disclosure does not limit the specific method for describing head posture, and in other words, does not limit the specific form of head posture data.

[0052] As an example, body posture data may include: Euler angles of various joints of a body part, and these joints may be pre-specified. It should be understood that the body part posture may also be described by other appropriate means, such as rotation quaternions. This disclosure does not limit the specific manner of describing the body part posture, that is, the specific form of the body posture data is not limited.

[0053] As an example, hand posture data may include: Euler angles of various joints of the hands (including the left and right hands), and these joints may be pre-specified. It should be understood that hand postures may also be described by other appropriate means, such as rotation quaternions. This disclosure does not limit the specific method of describing hand postures, that is, the specific form of hand posture data is not limited.

[0054] As an example, the current image frame may be an image frame captured by a monocular camera. It should be understood that the current image frame may also be an image frame captured by other types of camera devices, and the present disclosure does not limit this.

[0055] As an example, the current image frame may be an image obtained by capturing the entire body of the target object.

[0056] Figure 2 A flowchart illustrating a method for acquiring posture detection data of multiple parts of a target object in a current image frame according to an exemplary embodiment of the present disclosure is shown.

[0057] Reference Figure 2 In step S201, the current image frame is input into a body detection model to obtain body posture detection data. In addition, for example, neck position data and wrist position data can also be obtained.

[0058] As an example, the body posture detection data may include but is not limited to the following items: posture detection data of human body joints other than hands and head. For example, the human body other than hands and head may include: neck, torso, lower limbs, and arms.

[0059] In step S202, the head region image in the current image frame is input into a head detection model to obtain head posture detection data.

[0060] For example, based on the neck position data, the area where the head is located can be identified and cut out from the current image frame as the head area image.

[0061] As an example, the head detection model can obtain head posture detection data based on the facial information of the target object in the head area image.

[0062] In step S203, the hand region image in the current image frame is input into the hand detection model to obtain hand posture detection data.

[0063] For example, based on the wrist position data, the area where the hand is located can be identified and cut out from the current image frame as the hand area image.

[0064] As an example, the hand gesture detection data may include, but is not limited to, the following items: finger gesture detection data and wrist gesture detection data.

[0065] As an example, the body detection model, hand detection model, and head detection model can be constructed based on convolutional neural networks respectively. It should be understood that they can also be constructed based on other appropriate algorithms, and the present disclosure does not limit this.

[0066] According to the embodiments of the present disclosure, the detection rates of the head and hand detection models are improved by using body capture results as pre-processing information. In other words, when the body is detected, the head and hands are more easily found based on the body detection results, thereby improving the detection rate of the head and hands, and further improving the robustness and accuracy of driving the virtual avatar.

[0067] return Figure 1 In step S102, for each target part associated with other parts among the multiple parts, posture driving data of the target part is determined based on the posture detection data of the target part and its associated parts.

[0068] As examples, various target locations may include body parts and hands.

[0069] Figure 3 A flowchart illustrating a method for determining gesture driving data of a target part according to an exemplary embodiment of the present disclosure is shown.

[0070] Reference Figure 3 In step S301, the posture driving data of the body part is determined based on the body posture detection data and the head posture detection data.

[0071] A body part is associated with the head via its neck. Therefore, posture drive data for the body part can be determined based on the posture detection data for the body part and the head. Specifically, for example, posture drive data for the neck of the body part can be determined based on the body posture detection data and the head posture detection data, and posture drive data for body parts other than the neck can be determined based on the body posture detection data.

[0072] In step S302, hand gesture driving data is determined based on the hand gesture detection data and the body gesture detection data.

[0073] The hand is associated with a body part via its wrist. Therefore, hand gesture drive data can be determined based on the gesture detection data of the hand and the body part. Specifically, for example, gesture drive data for the wrist of the hand can be determined based on the hand gesture detection data and the body gesture detection data, and gesture drive data for the fingers of the hand can be determined based on the hand gesture detection data.

[0074] In a first embodiment, the neck posture driving data can be determined based on the neck posture detection data in the body posture detection data, the head posture detection data and the neck posture association data, wherein the neck posture association data includes at least one of the following items: the posture detection data of the parent joint of the neck joint in the body posture detection data, the posture driving data of the neck determined for the previous image frame, and the posture data of the preset neck idle state, wherein the neck idle state is the state that the neck needs to enter when the body part and head of the target object in the current image frame are not detected.

[0075] It should be immediately understood that the neck posture drive data can be determined based on the above data using an appropriate method. For example, the neck posture drive data can be determined by minimizing the weighted sum of the bi-norms of the distances between the neck posture drive data and various neck fusion data items, where the various neck fusion data items include the neck posture detection data, the head posture detection data, and the neck posture association data. In other words, the neck posture drive data that minimizes this weighted sum can be determined through minimization optimization.

[0076] As an example, the neck posture data r1 that minimizes the value of the following formula (ie, formula (1)) may be determined as the posture driving data of the neck:

[0077] w1||r1-r head ||2+w2||r1-r bodyneck ||2+w3||r1-r preneck ||2+w4||r1-r neckfather ||2+w5||r1-r idle1||2 (1)

[0078] Among them, r head Represents the head posture detection data, r bodyneck Represents the neck posture detection data, r neckfather Represents the posture detection data of the parent joint of the neck joint, r preneck Represents the posture driving data of the neck determined for the previous image frame, r idle1 Represents the posture data of the neck in the idle state, w1 represents the confidence of the head posture detection data, w2 represents the confidence of the neck posture detection data, w3 represents the smoothing weight related to the frame rate, w4 represents the muscle distortion penalty weight, and w5 represents the loss processing weight.

[0079] As an example, the value of w1 is related to the face angle of the target object in the current image frame. For example, when the face of the target object is not detected, it can be set to 0, and when the face is detected at a frontal angle, it can be set to 1. As the face angle changes from a frontal angle to an angle where no face is detected, w1 gradually transitions from 1 to 0.

[0080] As an example, the value of w2 may be related to the value of w1, for example, may be 1-w1.

[0081] As an example, by presetting r idle1 , when the body part and head of the target object in the current image frame are not detected, the neck of the avatar can be slowly regressed to the default neck state. When the body part and / or head of the target object in the current image frame are detected, w5 can be set to 0, and when the body part and head of the target object in the current image frame are not detected, w5 can be set to a value greater than 0 and less than or equal to 1. For example, when the body part and head of the target object in the current image frame (for example, the i-th frame) are not initially detected, w5 can be set to a smaller value a1 greater than 0; when the body part and head of the target object in the current image frame (for example, the i+1-th frame) are continuously not detected, w5 can be set to a value a2 greater than a1, and so on, until w5 is set to 1, so that the neck of the avatar slowly regresses to the default neck state, thereby achieving inter-frame smoothing.

[0082] It should be understood that w1-w5 may vary with the current image frame, that is, their values ​​may be adaptively adjusted. For example, for different image frames, the corresponding w1-w5 may be different or the same.

[0083] According to an embodiment of the present disclosure, in the process of calculating the neck posture driving data using the neck posture detection data in the body posture detection data, the head posture detection data is integrated and the energy formula is used for optimization solution, thereby enhancing the smoothness of the neck posture driving data corresponding to the previous image frame (i.e., taking into account the constraints of the previous frame), taking into account the confidence of the posture detection data, muscle distortion penalty, and inter-frame smoothing, and improving the stability, robustness, naturalness, and accuracy of the driving in multi-frame video scenarios.

[0084] In a second embodiment, the wrist posture driving data can be determined based on the wrist posture detection data in the hand posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data, wherein the wrist posture association data includes at least one of the following items: the wrist posture driving data determined for the previous image frame, and the posture data of the preset wrist idle state, wherein the wrist idle state is the state that the wrist needs to enter when the hand of the target object in the current image frame is not detected.

[0085] It should be immediately understood that wrist posture drive data can be determined based on the above data using an appropriate method. As an example, the wrist posture drive data can be determined by minimizing the weighted sum of the bi-norms of the distances between the wrist posture drive data and various wrist fusion data, where the various wrist fusion data include: the wrist posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data.

[0086] As an example, the wrist posture data r2 that minimizes the value of formula (2) can be determined as the posture driving data of the wrist.

[0087] w6||r2-r wrist ||2+w7||r2-r prewrist ||2+w8||r2-r wristfather ||2+w9||r2-r idle2 ||2 (2)

[0088] Among them, r wrist represents the wrist posture detection data, r prewrist represents the wrist posture driving data determined for the previous image frame, r wristfather represents the posture detection data of the parent joint of the wrist joint (e.g., the forearm joint), r idle2 It represents the posture data of the wrist in the idle state, w6 represents the confidence of the wrist posture detection data, w7 represents the smoothing weight related to the frame rate, w8 represents the muscle distortion penalty weight, and w9 represents the loss processing weight.

[0089] As an example, by presetting r idle2 , when the hand of the target object in the current image frame is not detected, the wrist of the avatar can be slowly retracted to the default wrist state. When the hand of the target object in the current image frame is detected, w9 can be set to 0, and when the hand of the target object in the current image frame is not detected, w9 can be set to a value greater than 0 and less than or equal to 1. For example, when the hand of the target object in the current image frame (for example, the i-th frame) is not initially detected, w9 can be set to a smaller value b1 greater than 0; when the hand of the target object in the current image frame (for example, the i+1-th frame) is continuously not detected, w9 can be set to a value b2 greater than b1, and so on, until w9 is set to 1, so that the wrist of the avatar slowly retracts to the default wrist state, thereby achieving inter-frame smoothing.

[0090] It should be understood that w6-w9 may vary depending on the current image frame, that is, their values ​​may be adaptively adjusted. For example, for different image frames, the corresponding w6-w9 may be different or the same.

[0091] According to an embodiment of the present disclosure, in the process of calculating the wrist posture driving data using the wrist posture detection data in the hand posture detection data, the body posture detection data is integrated and the energy formula is used for optimization solution, thereby enhancing the smoothness of the wrist posture driving data corresponding to the previous image frame (i.e., taking into account the constraints of the previous frame), taking into account the confidence of the posture detection data, muscle distortion penalty, and inter-frame smoothing, and improving the stability, robustness, naturalness, and accuracy of the driving in multi-frame video scenarios.

[0092] In a third embodiment, the posture driving data of the body parts other than the neck can be determined based on the posture detection data of the body parts other than the neck in the body posture detection data (that is, the posture detection data other than the neck posture detection data in the body posture detection data) and the body posture association data, wherein the body posture association data includes at least one of the following items: the posture driving data of the body parts other than the neck determined for the previous image frame, and the posture data of the preset body parts in an idle state, wherein the body part idle state is the state that the body parts other than the neck need to enter when the body parts of the target object in the current image frame are not detected.

[0093] It should be immediately understood that the posture drive data for body parts other than the neck can be determined based on the above data using an appropriate method. As an example, the posture drive data for body parts other than the neck can be determined by minimizing the weighted sum of the two norms of the distances between the posture drive data for the body parts other than the neck and the various body fusion data, wherein the various body fusion data include: the posture detection data for the body parts other than the neck and the body posture association data.

[0094] As an example, the posture data r3 of the body part other than the neck that minimizes the value of formula (3) can be determined as the posture driving data of the body part other than the neck.

[0095] w 10 ||r3-r body ||2+w 11 ||r3-r prebody ||2+w 12 ||r3-r idle3 ||2 (3)

[0096] Among them, r body represents the posture detection data of the body parts other than the neck, r prebody represents the posture driving data of the body parts other than the neck determined for the previous image frame, r idle3 Indicates the posture data of the body parts when they are idle, w 10 represents the confidence of the posture detection data of body parts other than the neck, w 11 represents the smoothing weight related to the frame rate, w 12 Represents the loss processing weight.

[0097] As an example, by presetting r idle3 , when the target object's body part is not detected in the current image frame, the body parts of the avatar except the neck will slowly fall back to the default body part state. 12 Can be set to 0. When no body part of the target object in the current image frame is detected, w 12 can be set to a value greater than 0 and less than or equal to 1. For example, when the body part of the target object in the current image frame (eg, the i-th frame) is not initially detected, w 12 is set to a smaller value c1 greater than 0; when the body part of the target object in the current image frame (for example, the i+1th frame) is not detected continuously, w 12 Set to a value c2 greater than c1, and so on, until w 12Set to 1 to slowly roll back the avatar's body parts, except the neck, to their default state, thus achieving smoothing between frames.

[0098] It should be understood that 10 -w 12 It can vary with the current image frame, that is, its value can be adaptively adjusted. For example, for different image frames, the corresponding w 10 -w 12 Can be different or the same.

[0099] According to an embodiment of the present disclosure, in the process of calculating the posture driving data of body parts other than the neck using the posture detection data of body parts other than the neck, the smoothness of the posture driving data corresponding to the previous image frame is enhanced (that is, the constraints of the previous frame are taken into account), the confidence of the posture detection data and the inter-frame smoothing are taken into account, and the stability, robustness, naturalness and accuracy of the driving in multi-frame video scenes are improved.

[0100] In a fourth embodiment, the finger posture driving data can be determined based on the finger posture detection data and the finger posture association data in the hand posture detection data, wherein the finger posture association data includes at least one of the following items: the finger posture driving data determined for the previous image frame, and the posture data of the preset finger idle state, wherein the finger idle state is the state that the finger needs to enter when the finger of the target object in the current image frame is not detected.

[0101] It should be immediately understood that the finger gesture driving data can be determined based on the above data using an appropriate method. As an example, the finger gesture driving data can be determined by minimizing the weighted sum of the bi-norms of the distances between the finger gesture driving data and each finger fusion data item, wherein the each finger fusion data item includes: the finger gesture detection data and the finger gesture association data.

[0102] As an example, the finger posture data r4 that minimizes the value of formula (4) can be determined as the finger posture driving data,

[0103] w 13 ||r4-r hand ||2+w 14 ||r4-r prehand ||2+w 15 ||r4-r idle4 ||2 (4)

[0104] Among them, r hand Represents the finger posture detection data, r prehand Represents the finger posture driving data determined for the previous image frame, r idle4Indicates the posture data of the finger when it is idle, w 13 Indicates the confidence of finger posture detection data, w 14 represents the smoothing weight related to the frame rate, w 15 Represents the loss processing weight.

[0105] As an example, by presetting r idle4 , which can make the virtual image's fingers slowly return to the default finger state when the target object's hand is not detected in the current image frame. When the target object's hand is detected in the current image frame, w 15 Can be set to 0. When the hand of the target object in the current image frame is not detected, w 15 can be set to a value greater than 0 and less than or equal to 1. For example, when the hand of the target object in the current image frame (eg, the i-th frame) is not initially detected, w 15 is set to a smaller value d1 greater than 0; when the hand of the target object in the current image frame (for example, the i+1th frame) is not detected continuously, w 15 Set to a value d2 that is greater than d1, and so on, until w 15 Set to 1 to make the avatar's fingers slowly fall back to the default finger state, thus achieving smoothness between frames.

[0106] It should be understood that 13 -w 15 It can vary with the current image frame, that is, its value can be adaptively adjusted. For example, for different image frames, the corresponding w 13 -w 15 Can be different or the same.

[0107] According to an embodiment of the present disclosure, in the process of calculating the finger posture driving data using finger posture detection data, the smoothness of the finger posture driving data corresponding to the previous image frame is enhanced (that is, the constraints of the previous frame are taken into account), the confidence of the posture detection data and inter-frame smoothing are taken into account, and the stability, robustness, naturalness and accuracy of the driving in multi-frame video scenes are improved.

[0108] In step S103, based on the determined posture driving data of each target part, the corresponding part of the virtual image is driven. The virtual image and the target object have one-to-one corresponding parts.

[0109] As an example, the neck of the virtual image can be driven according to the neck posture driving data, the body parts of the virtual image other than the neck can be driven according to the posture driving data of the body parts other than the neck, the wrist of the virtual image can be driven according to the wrist posture driving data, and the fingers of the virtual image can be driven according to the finger posture driving data.

[0110] As an example, the posture data of each part of the virtual image may be set as the posture driving data of the part determined in step S102 , thereby driving the corresponding part of the virtual image based on the determined posture driving data of each part.

[0111] As an example, step S103 may include: based on the posture driving data of the wrist, correcting the posture driving data of the parent joint of the wrist joint in the posture driving data of the body parts other than the neck to obtain the corrected posture driving data of the body parts other than the neck; then, based on the corrected posture driving data of the body parts other than the neck, the posture driving data of the fingers, the posture driving data of the neck and the posture driving data of the wrist, driving the corresponding parts of the virtual image.

[0112] For example, based on the posture drive data of the wrist and the posture drive data of the arm, the x-axis rotation component (i.e., the x-axis angle) of the relative rotation of the wrist and forearm can be determined. Then, the posture drive data of the forearm can be corrected based on the product of the determined rotation component and a preset proportional coefficient (for example, the posture drive data of the forearm is multiplied by the rotation quaternion corresponding to the product as the corrected data) to transmit a part of the determined rotation component to the forearm.

[0113] According to an embodiment of the present disclosure, the virtual image can be driven to be in a posture consistent with the target object in the current image frame based on the current image frame, that is, the virtual image can be driven to make an action consistent with the target object in the current image frame, thereby forming an animation frame of the virtual image.

[0114] As an example, the image processing method according to an exemplary embodiment of the present disclosure may further include: forming a continuous animation frame sequence based on the generated animation frames corresponding to the respective image frames and the timing of the respective image frames. Thus, the client can see a vivid virtual image.

[0115] As an example, the current image frame can be a video frame captured in real time, or a video frame from a received video. For example, after executing steps S101-S103 to drive the virtual character based on the current image frame (e.g., the image frame captured at the current moment), the image frame captured at the next moment can be used as the current image frame to continue executing steps S101-S103 to drive the virtual character, thereby enabling the virtual character to be driven continuously, forming an animation frame sequence of the virtual character. Alternatively, steps S101-S103 can be executed in sequence using each video frame from the received video as the current image frame, and based on the generated animation frames corresponding to each image frame and the timing of each image frame, a continuous animation frame sequence of the virtual character is formed.

[0116] According to exemplary embodiments of the present disclosure, a low-cost, robust, and natural-looking method for full-body capture and driving of avatars is provided. This disclosure considers that a monocular camera can independently capture, detect, and output the postures of the head, body parts, and hands. However, to fully drive an avatar, these outputs must be integrated and the constraints between the various parts must be considered. This approach addresses issues such as jitter, unnaturalness, irrationality, and missed or false detection of hands and head, which are common with full-body driving solutions.

[0117] Figure 4 A block diagram illustrating the structure of an image processing apparatus according to an exemplary embodiment of the present disclosure is shown.

[0118] like Figure 4 As shown, the image processing apparatus 10 according to an exemplary embodiment of the present disclosure includes: a detection data acquisition unit 101 , a driving data acquisition unit 102 , and a driving unit 103 .

[0119] Specifically, the detection data acquisition unit 101 is configured to acquire posture detection data of multiple parts of the target object in the current image frame.

[0120] The driving data acquisition unit 102 is configured to determine, for each target part associated with other parts among the multiple parts, posture driving data of the target part based on the posture detection data of the target part and its associated parts.

[0121] The driving unit 103 is configured to drive corresponding parts of the virtual image based on the determined posture driving data of each target part, and the virtual image and the target object have one-to-one corresponding parts.

[0122] As an example, the multiple parts may include hands, head and body parts, the posture detection data of the multiple parts may include head posture detection data, body posture detection data and hand posture detection data, and the various target parts may include body parts and hands.

[0123] As an example, the drive data acquisition unit 102 can be configured to: determine the posture driving data of the neck among the body parts based on the body posture detection data and the head posture detection data, and determine the posture driving data of the body parts other than the neck based on the body posture detection data; determine the posture driving data of the wrist in the hand based on the hand posture detection data and the body posture detection data, and determine the posture driving data of the fingers in the hand based on the hand posture detection data.

[0124] As an example, the driving data acquisition unit 102 can be configured to: determine the neck posture driving data based on the neck posture detection data in the body posture detection data, the head posture detection data and the neck posture association data, wherein the neck posture association data includes at least one of the following items: the posture detection data of the parent joint of the neck joint in the body posture detection data, the posture driving data of the neck determined for the previous image frame, and the posture data of the preset neck idle state, wherein the neck idle state is the state that the neck needs to enter when the body part and head of the target object in the current image frame are not detected.

[0125] As an example, the drive data acquisition unit 102 can be configured to determine the neck posture drive data by minimizing the weighted sum of the two norms of the distances between the neck posture drive data and each item of neck fusion data, wherein the each item of neck fusion data includes: the neck posture detection data, the head posture detection data, and the neck posture association data.

[0126] As an example, the drive data acquisition unit 102 can be configured to: determine the wrist posture driving data based on the wrist posture detection data in the hand posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data, wherein the wrist posture association data includes at least one of the following items: the wrist posture driving data determined for the previous image frame, and the posture data of the preset wrist idle state, wherein the wrist idle state is the state that the wrist needs to enter when the hand of the target object in the current image frame is not detected.

[0127] As an example, the drive data acquisition unit 102 can be configured to determine the posture drive data of the wrist by minimizing the weighted sum of the two norms of the distances between the posture drive data of the wrist and each wrist fusion data, wherein the each wrist fusion data includes: the wrist posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data.

[0128] As an example, the detection data acquisition unit 101 can be configured to: input the current image frame into the body detection model to obtain body posture detection data; input the head area image in the current image frame into the head detection model to obtain head posture detection data; input the hand area image in the current image frame into the hand detection model to obtain hand posture detection data.

[0129] As an example, the detection data acquisition unit 101 can be configured to: input the current image frame into the body detection model to obtain the body posture detection data, neck position data, and wrist position data; wherein, the detection data acquisition unit 101 can also be configured to: identify and cut out the area where the head is located from the current image frame according to the neck position data, as the head area image; identify and cut out the area where the hand is located from the current image frame according to the wrist position data, as the hand area image.

[0130] As an example, the driving data acquisition unit 102 can be configured to: determine the posture driving data of the body parts other than the neck based on the posture detection data and body posture association data of the body parts other than the neck in the body posture detection data, wherein the body posture association data includes at least one of the following items: the posture driving data of the body parts other than the neck determined for the previous image frame, and the posture data of the preset body parts in an idle state, wherein the body part idle state is the state that the body parts other than the neck need to enter when the body parts of the target object in the current image frame are not detected.

[0131] As an example, the drive data acquisition unit 102 can be configured to determine the posture drive data of the body part other than the neck by minimizing the weighted sum of the bi-norm of the distance between the posture drive data of the body part other than the neck and the various body fusion data, wherein the various body fusion data include: the posture detection data of the body part other than the neck and the body posture association data.

[0132] As an example, the driving data acquisition unit 102 can be configured to: determine the finger posture driving data based on the finger posture detection data and the finger posture association data in the hand posture detection data, wherein the finger posture association data includes at least one of the following items: the finger posture driving data determined for the previous image frame, and the posture data of the preset finger idle state, wherein the finger idle state is the state that the finger needs to enter when the finger of the target object in the current image frame is not detected.

[0133] As an example, the driving data acquisition unit 102 can be configured to determine the posture driving data of the finger by minimizing the weighted sum of the bi-norms of the distances between the posture driving data of the finger and each finger fusion data, wherein the each finger fusion data includes: the finger posture detection data and the finger posture association data.

[0134] As an example, the driving unit 103 can be configured to: based on the posture driving data of the wrist, correct the posture driving data of the parent joint of the wrist joint in the posture driving data of the body parts other than the neck to obtain the corrected posture driving data of the body parts other than the neck; drive the corresponding parts of the virtual image based on the corrected posture driving data of the body parts other than the neck, the posture driving data of the fingers, the posture driving data of the neck and the posture driving data of the wrist.

[0135] As an example, the current image frame may be an image frame captured by a monocular camera.

[0136] Regarding the image processing device 10 in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0137] In addition, it should be understood that the various units in the image processing apparatus 10 according to the exemplary embodiment of the present disclosure may be implemented as hardware components and / or software components. Those skilled in the art may implement the various units using, for example, a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), depending on the processing performed by the defined various units.

[0138] Figure 5 FIG. 1 shows a structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Figure 5 The electronic device 20 includes: at least one memory 201 and at least one processor 202, wherein the at least one memory 201 stores a set of computer-executable instructions. When the computer-executable instruction set is executed by the at least one processor 202, the image processing method described in the above exemplary embodiment is executed.

[0139] As an example, the electronic device 20 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above-mentioned instruction set. Here, the electronic device 20 is not necessarily a single electronic device, but may also be any collection of devices or circuits capable of executing the above-mentioned instructions (or instruction sets) individually or in combination. The electronic device 20 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0140] In the electronic device 20, the processor 202 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 202 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0141] The processor 202 can execute instructions or codes stored in the memory 201, wherein the memory 201 can also store data. Instructions and data can also be sent and received over a network via a network interface device, wherein the network interface device can use any known transmission protocol.

[0142] The memory 201 may be integrated with the processor 202, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, the memory 201 may comprise a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. The memory 201 and the processor 202 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, or the like, such that the processor 202 can access files stored in the memory.

[0143] In addition, the electronic device 20 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 20 may be connected to each other via a bus and / or a network.

[0144] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, the at least one processor is prompted to perform the image processing method as described in the above exemplary embodiment. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or ultra-fast digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device configured to store the computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0145] According to an exemplary embodiment of the present disclosure, a computer program product may be provided. Instructions in the computer program product may be executed by at least one processor to implement the image processing method as described in the exemplary embodiment above.

[0146] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0147] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image processing method, characterized in that: include: Obtaining posture detection data of multiple parts of the target object in the current image frame; For each target part associated with other parts among the multiple parts, determining posture driving data of the target part based on posture detection data of the target part and its associated parts; Based on the determined posture driving data of each target part, driving the corresponding part of the virtual image, the virtual image and the target object having a one-to-one corresponding part; The multiple parts include hands, head, and body parts, the posture detection data of the multiple parts include head posture detection data, body posture detection data, and hand posture detection data; the target part includes a body part, and for each target part associated with other parts among the multiple parts, the step of determining the posture driving data of the target part based on the posture detection data of the target part and its associated parts includes: Determine the posture driving data of the neck by minimizing optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the neck and each item of neck fusion data is minimized; determining posture driving data of body parts other than the neck based on the body posture detection data; Among them, the various items of neck fusion data include: neck posture detection data in the body posture detection data, the head posture detection data, and neck posture association data; the neck posture association data includes at least one of the following items: posture detection data of the parent joint of the neck joint in the body posture detection data, neck posture driving data determined for the previous image frame, and preset neck idle state posture data.

2. The image processing method according to claim 1, wherein: The target part also includes a hand. For each target part among the multiple parts that is associated with other parts, the step of determining the posture driving data of the target part based on the posture detection data of the target part and its associated parts further includes: Based on the hand posture detection data and the body posture detection data, posture driving data of a wrist in a hand is determined, and based on the hand posture detection data, posture driving data of fingers in a hand is determined.

3. The image processing method according to claim 1, wherein: The neck idle state is a state that the neck needs to enter when the body part and the head of the target object in the current image frame are not detected.

4. The image processing method according to claim 2, wherein: The step of determining gesture driving data of a wrist in a hand based on the hand gesture detection data and the body gesture detection data comprises: Determine the posture driving data of the wrist based on the wrist posture detection data in the hand posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data, The wrist posture associated data includes at least one of the following items: wrist posture driving data determined for the previous image frame, preset wrist posture data in an idle state, The wrist idle state is a state that the wrist needs to enter when the hand of the target object in the current image frame is not detected.

5. The image processing method according to claim 4, characterized in that The step of determining the posture driving data of the wrist includes: The posture driving data of the wrist is determined by minimization optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the wrist and each wrist fusion data is minimized. The various wrist fusion data include: the wrist posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data.

6. The image processing method according to claim 1, wherein: The steps of obtaining posture detection data of multiple parts of the target object in the current image frame include: Input the current image frame into the body detection model to obtain body posture detection data; Input the head area image in the current image frame into the head detection model to obtain head posture detection data; The hand area image in the current image frame is input into the hand detection model to obtain hand posture detection data.

7. The image processing method according to claim 6, characterized in that: The step of inputting the current image frame into a body detection model to obtain body posture detection data includes: inputting the current image frame into the body detection model to obtain the body posture detection data, neck position data, and wrist position data; Wherein, the image processing method further includes: identifying and cutting out the region where the head is located from the current image frame according to the neck position data as the head region image; According to the wrist position data, the area where the hand is located is identified and cut out from the current image frame as the hand area image.

8. The image processing method according to claim 1, wherein: The step of determining posture driving data of body parts other than the neck based on the body posture detection data comprises: determining the posture driving data of the body parts other than the neck based on the posture detection data of the body parts other than the neck and the body posture association data in the body posture detection data, The body posture associated data includes at least one of the following items: posture driving data of body parts other than the neck determined for the previous image frame, and posture data of preset body parts in an idle state. The body part idle state is a state that body parts other than the neck need to enter when the body part of the target object in the current image frame is not detected.

9. The image processing method according to claim 8, characterized in that: The step of determining the posture driving data of the body part other than the neck comprises: The posture driving data of the body parts other than the neck are determined by minimization optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the body parts other than the neck and the various body fusion data is minimized. The various items of body fusion data include: posture detection data of the body parts other than the neck and the body posture association data.

10. The image processing method according to claim 2, wherein: The step of determining gesture driving data of fingers in the hand based on the hand gesture detection data comprises: determining the finger posture driving data based on the finger posture detection data and the finger posture association data in the hand posture detection data, The finger posture associated data includes at least one of the following items: finger posture driving data determined for the previous image frame, preset finger posture data in an idle state, The finger idle state is a state that the finger needs to enter when the finger of the target object in the current image frame is not detected.

11. The image processing method according to claim 10, wherein: The step of determining the gesture driving data of the finger comprises: The posture driving data of the finger is determined by minimization optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the finger and the fusion data of each finger is minimized. The various finger fusion data include: the finger posture detection data and the finger posture association data.

12. The image processing method according to claim 2, wherein: The steps of driving the corresponding parts of the virtual image based on the determined posture driving data of each target part include: Based on the posture drive data of the wrist, correcting the posture drive data of the parent joint of the wrist joint in the posture drive data of the body parts other than the neck to obtain corrected posture drive data of the body parts other than the neck; Based on the corrected posture driving data of the body parts other than the neck, the posture driving data of the fingers, the posture driving data of the neck, and the posture driving data of the wrist, the corresponding parts of the avatar are driven.

13. The image processing method according to claim 1, wherein: The current image frame is an image frame captured by a monocular camera.

14. An image processing device, characterized in that: include: a detection data acquisition unit, configured to acquire posture detection data of multiple parts of the target object in the current image frame; a driving data acquisition unit configured to determine, for each target part associated with other parts among the plurality of parts, posture driving data of the target part based on posture detection data of the target part and its associated parts; a driving unit configured to drive corresponding parts of the virtual image based on the determined posture driving data of each target part, wherein the virtual image has a one-to-one corresponding part with the target object; The multiple parts include hands, head, and body parts, the posture detection data of the multiple parts include head posture detection data, body posture detection data, and hand posture detection data; the target part includes a body part, and the driving data acquisition unit is configured to: Determine the posture driving data of the neck by minimizing optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the neck and each item of neck fusion data is minimized; determining posture driving data of body parts other than the neck based on the body posture detection data; Among them, the various items of neck fusion data include: neck posture detection data in the body posture detection data, the head posture detection data, and neck posture association data; the neck posture association data includes at least one of the following items: posture detection data of the parent joint of the neck joint in the body posture detection data, neck posture driving data determined for the previous image frame, and preset neck idle state posture data.

15. The image processing device according to claim 14, wherein: The target part also includes a hand, and the driving data acquisition unit is further configured to: Based on the hand posture detection data and the body posture detection data, posture driving data of the wrist in the hand is determined, and based on the hand posture detection data, posture driving data of the fingers in the hand is determined.

16. The image processing device according to claim 14, wherein: The neck idle state is a state that the neck needs to enter when the body part and the head of the target object in the current image frame are not detected.

17. The image processing device according to claim 15, wherein: The driving data acquisition unit is configured to determine the posture driving data of the wrist based on the wrist posture detection data in the hand posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data. The wrist posture associated data includes at least one of the following items: wrist posture driving data determined for the previous image frame, preset wrist posture data in an idle state, The wrist idle state is a state that the wrist needs to enter when the hand of the target object in the current image frame is not detected.

18. The image processing device according to claim 17, wherein The drive data acquisition unit is configured to determine the posture drive data of the wrist by minimizing optimization, wherein the weighted sum of the two norms of the distances between the posture drive data of the wrist and each wrist fusion data is minimized. The various wrist fusion data include: the wrist posture detection data, the posture detection data of the parent joint of the wrist joint in the body posture detection data, and the wrist posture association data.

19. The image processing device according to claim 14, wherein The detection data acquisition unit is configured to: Input the current image frame into the body detection model to obtain body posture detection data; Input the head area image in the current image frame into the head detection model to obtain head posture detection data; The hand area image in the current image frame is input into the hand detection model to obtain hand posture detection data.

20. The image processing device according to claim 19, wherein The detection data acquisition unit is configured to: input the current image frame into the body detection model to obtain the body posture detection data, neck position data, and wrist position data; In which, the detection data acquisition unit is further configured to: identify and cut out the area where the head is located from the current image frame according to the neck position data as the head area image; identify and cut out the area where the hand is located from the current image frame according to the wrist position data as the hand area image.

21. The image processing device according to claim 14, wherein The driving data acquisition unit is configured to: determine the posture driving data of the body parts other than the neck based on the posture detection data of the body parts other than the neck in the body posture detection data and the body posture association data, The body posture associated data includes at least one of the following items: posture driving data of body parts other than the neck determined for the previous image frame, and posture data of preset body parts in an idle state. The body part idle state is a state that body parts other than the neck need to enter when the body part of the target object in the current image frame is not detected.

22. The image processing device according to claim 21, wherein The driving data acquisition unit is configured to determine the posture driving data of the body parts other than the neck by minimizing optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the body parts other than the neck and the various body fusion data is minimized. The various items of body fusion data include: posture detection data of the body parts other than the neck and the body posture association data.

23. The image processing device according to claim 15, wherein The driving data acquisition unit is configured to: determine the finger posture driving data based on the finger posture detection data and the finger posture association data in the hand posture detection data, The finger posture associated data includes at least one of the following items: finger posture driving data determined for the previous image frame, preset finger posture data in an idle state, The finger idle state is a state that the finger needs to enter when the finger of the target object in the current image frame is not detected.

24. The image processing device according to claim 23, wherein The driving data acquisition unit is configured to: determine the posture driving data of the finger by minimizing optimization, wherein the weighted sum of the two norms of the distances between the posture driving data of the finger and the fusion data of each finger is minimized, The various finger fusion data include: the finger posture detection data and the finger posture association data.

25. The image processing device according to claim 15, wherein The drive unit is configured as: Based on the posture drive data of the wrist, correcting the posture drive data of the parent joint of the wrist joint in the posture drive data of the body parts other than the neck to obtain corrected posture drive data of the body parts other than the neck; Based on the corrected posture driving data of the body parts other than the neck, the posture driving data of the fingers, the posture driving data of the neck, and the posture driving data of the wrist, the corresponding parts of the avatar are driven.

26. The image processing device according to claim 14, wherein The current image frame is an image frame captured by a monocular camera.

27. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer-executable instructions, When the computer-executable instructions are executed by the at least one processor, the computer-executable instructions prompt the at least one processor to execute the image processing method according to any one of claims 1 to 13.

28. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the image processing method according to any one of claims 1 to 13.

29. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by at least one processor, the image processing method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Virtual image driving method and device and server

    CN114519758A