Gesture data completion method and apparatus for three-dimensional object, device, storage medium, and product

By acquiring incomplete 3D posture data and posture description text, and using the text graph model to generate 2D posture images and identify joint points, the problem of high consumption of manpower and material resources in the existing technology is solved, and efficient 3D object posture data completion is achieved.

WO2025156888A9PCT designated stage Publication Date: 2025-09-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140519
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2024-12-19
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

The existing technology of building a hand motion library through motion capture gloves consumes a lot of manpower and material resources, and the collected hand motions are poorly consistent with the limb movements of the three-dimensional object.

Method used

By obtaining the incomplete three-dimensional posture data and posture description text of the three-dimensional object, a two-dimensional posture image is generated using the text graph model, and the missing joint point data is obtained through joint point recognition to complete the posture data of the three-dimensional object.

Benefits of technology

It reduces the consumption of manpower and material resources, improves the efficiency of completing 3D object posture data and the utilization rate of data, and the generated posture data perfectly matches the original data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140519_12092025_PF_FP_ABST
    Figure CN2024140519_12092025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and discloses a gesture data completion method and apparatus for a three-dimensional object, a device, a storage medium, and a product. The method comprises: acquiring three-dimensional incomplete gesture data of a three-dimensional object in a preset gesture, wherein the three-dimensional incomplete gesture data comprises first three-dimensional joint point data of some joint points of the three-dimensional object; calling a text-to-image model to generate a two-dimensional gesture image of the three-dimensional object in the preset gesture on the basis of the three-dimensional incomplete gesture data and a gesture description text, wherein the gesture description text is used for describing the preset gesture; performing joint point recognition on the two-dimensional gesture image to obtain second three-dimensional joint point data of missing joint points; and adding the second three-dimensional joint point data to the three-dimensional incomplete gesture data to obtain three-dimensional complete gesture data of the three-dimensional object. The method can quickly complete the missing joint points in gesture data of the three-dimensional object.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment, storage medium and product for completing posture data of three-dimensional objects

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 26, 2024, application number 202410113465.1, and application name “Three-dimensional object posture completion method, device, equipment, storage medium and product”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a posture data completion technology for three-dimensional objects. Background Art

[0003] With the development of computer technology, three-dimensional objects can be created through computer programs or artificial intelligence technology, and the three-dimensional objects can simulate the behavior and dialogue of people or animals.

[0004] For three-dimensional objects that only have partial limb posture data but no hand posture, related technologies use motion capture gloves to separately collect multiple hand movements to build a hand movement library. When it is necessary to drive the hand movements of the three-dimensional object, the match is directly retrieved in the hand movement library, thereby achieving fine driving of the three-dimensional object.

[0005] However, the related art method of using motion capture gloves to build a hand motion library consumes a lot of manpower and material resources, and the collected hand motions are poorly consistent with the limb motions of the three-dimensional object. Summary of the Invention

[0006] The present application provides a method, device, equipment, storage medium and product for completing the posture data of a three-dimensional object. The technical solution is as follows.

[0007] According to one aspect of the present application, a method for completing posture data of a three-dimensional object is provided, the method being executed by a computer device, the method comprising:

[0008] Acquiring incomplete three-dimensional posture data of the three-dimensional object in a preset posture, the incomplete three-dimensional posture data including first three-dimensional joint point data of some joint points of the three-dimensional object;

[0009] Invoking a Wensheng graph model to generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and a posture description text, wherein the posture description text is used to describe the preset posture;

[0010] Performing joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of missing joint points of the three-dimensional object except for the partial joint points;

[0011] The second three-dimensional joint point data is added to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

[0012] According to one aspect of the present application, a device for completing posture data of a three-dimensional object is provided. The device is deployed on a computer device and includes:

[0013] a data module, configured to obtain three-dimensional incomplete posture data of the three-dimensional object in a preset posture, wherein the three-dimensional incomplete posture data includes first three-dimensional joint point data of some joint points of the three-dimensional object;

[0014] a generating module, configured to call a Wensheng graph model and generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and a posture description text; the posture description text is used to describe the preset posture;

[0015] a recognition module, configured to perform joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of missing joint points of the three-dimensional object except for the partial joint points;

[0016] The completion module is used to add the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

[0017] According to another aspect of the present application, a computer device is provided, comprising: a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method for completing the posture data of a three-dimensional object as described above.

[0018] According to another aspect of the present application, a computer storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the method for completing the posture data of a three-dimensional object as described above.

[0019] According to another aspect of the present application, a computer program product is provided, which includes a computer program stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes the method for completing the posture data of a three-dimensional object as described above.

[0020] The beneficial effects brought about by the technical solution provided by this application include at least:

[0021] Three-dimensional incomplete pose data and pose description text of a three-dimensional object in a preset pose are obtained. The three-dimensional incomplete pose data includes first three-dimensional joint point data of some joints of the three-dimensional object, and the pose description text is used to describe the preset pose. The three-dimensional incomplete pose data and pose description text are then input into a Wensheng graph model. Because the Wensheng graph model has the characteristic of accurately controlling image generation based on text, the Wensheng graph model can generate a two-dimensional pose image of the three-dimensional object in the preset pose based on the three-dimensional incomplete pose data and pose description text. The generated object in the two-dimensional pose image has the complete preset pose and can reflect the pose of the missing joint points. In this way, by identifying and extracting the joint points of the generated object in the two-dimensional pose image, second three-dimensional joint point data of the missing joint points can be obtained. The second three-dimensional joint point data can be added to the three-dimensional incomplete pose data to complete the three-dimensional incomplete pose data, thereby obtaining the three-dimensional completed pose data of the three-dimensional object, i.e., the complete pose data. This application targets three-dimensional objects with postures that do not have missing joints. Based on the posture data and posture description text of some joints in the three-dimensional incomplete posture data, a two-dimensional posture image with a complete preset posture can be directly generated through a text graph model. The posture of the three-dimensional object can then be completed using the second three-dimensional joint data extracted from the two-dimensional posture image. This eliminates the need to collect additional postures of missing joints to construct an action library, significantly reducing the consumption of manpower and material resources, improving the efficiency of completing the posture data of the three-dimensional object, and also increasing the utilization rate of open-source three-dimensional posture data for limbless postures. Furthermore, the posture data completed by this method can perfectly match the original incomplete posture data of the three-dimensional object, improving the effectiveness of completing the posture data of the three-dimensional object. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] FIG1 is a schematic diagram of a method for completing posture data of a three-dimensional object provided by an exemplary embodiment of the present application;

[0024] FIG2 is a schematic diagram of the architecture of a computer system provided by an exemplary embodiment of the present application;

[0025] FIG3 is a flow chart of a method for completing posture data of a three-dimensional object provided by an exemplary embodiment of the present application;

[0026] FIG4 is a schematic diagram of a method for completing posture data of a three-dimensional object provided by another exemplary embodiment of the present application;

[0027] FIG5 is a flowchart of a method for completing posture data of a three-dimensional object provided by another exemplary embodiment of the present application;

[0028] FIG6 is a schematic diagram of a posture skeleton graph provided by an exemplary embodiment of the present application;

[0029] FIG7 is a schematic diagram of a Wensheng graph model and a gesture control plug-in provided by an exemplary embodiment of the present application;

[0030] FIG8 is a schematic diagram of a Wensheng graph model and a gesture control plug-in provided by another exemplary embodiment of the present application;

[0031] FIG9 is a schematic diagram of a Wensheng graph model and a gesture control plug-in provided by another exemplary embodiment of the present application;

[0032] FIG10 is a flowchart of a method for completing posture data of a three-dimensional object provided by another exemplary embodiment of the present application;

[0033] FIG11 is a schematic diagram of similarity matching provided by an exemplary embodiment of the present application;

[0034] FIG12 is a flowchart of a method for completing posture data of a three-dimensional object provided by another exemplary embodiment of the present application;

[0035] FIG13 is a block diagram of a device for completing posture data of a three-dimensional object provided by an exemplary embodiment of the present application;

[0036] FIG14 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0038] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence computer vision, which are specifically illustrated by the following embodiments.

[0039] An embodiment of the present application provides a schematic diagram of a method for completing posture data of a three-dimensional object, as shown in FIG1 . The method can be executed by a computer device, which can be a terminal or a server.

[0040] Exemplarily, taking the three-dimensional object as a human body, the computer device obtains the three-dimensional incomplete posture data 10 and posture description text 30 corresponding to the three-dimensional object; the computer device maps the three-dimensional incomplete posture data 10 to a two-dimensional plane to obtain a posture skeleton graph 20; the computer device generates a human body posture graph 60 with a human body torso posture and hand posture based on the posture description text 30 and the posture skeleton graph 20, and the human body posture graph 60 is a two-dimensional posture image.

[0041] The incomplete 3D pose data 10 is data used to describe the pose of a portion of a 3D object (including the torso and limbs, but excluding the hands); or, the incomplete 3D pose data 10 is a parameter matrix used to describe the pose of a human torso of a 3D object; or, the incomplete 3D pose data 10 is human pose data lacking a limb pose. The limb pose includes at least one of, but is not limited to, a hand pose, a foot pose, a finger pose, and a toe pose.

[0042] Optionally, the three-dimensional incomplete posture data 10 includes at least one of data describing chest posture, data describing arm posture, and data describing leg posture, but is not limited thereto.

[0043] The posture description text 30 includes descriptive words describing a preset posture of the three-dimensional object; or, the posture description text 30 is a descriptive word used to describe the posture of the torso and limbs. For example, the posture description text 30 is white shirt, black shorts, boy, black shoes, waving hands, and open hands.

[0044] Exemplarily, for a three-dimensional object without limb posture, by obtaining the three-dimensional incomplete posture data 10 corresponding to the three-dimensional object and mapping the three-dimensional incomplete posture data 10 to a two-dimensional plane, a posture skeleton diagram 20 is obtained, and the posture skeleton diagram 20 includes a posture skeleton diagram. The computer device encodes the posture skeleton diagram 20 through the posture control plug-in 40, and inputs the encoded skeleton feature vector as an intermediate vector and the description word feature vector corresponding to the posture description text 30 into the Wensheng graph model 50 to generate an image, thereby obtaining a human body posture diagram 60. In this process, the Wensheng graph model 50 generates a human body posture diagram based on the posture description text 30. In the process of generating the human body posture diagram, the Wensheng graph model 50 uses the posture skeleton diagram 20 as a constraint, that is, the torso posture in the generated human body posture diagram is the same as the torso posture corresponding to the three-dimensional incomplete posture data 10.

[0045] For example, the torso posture of the three-dimensional object described in the three-dimensional incomplete posture data 10 is a kicking posture, and the posture description text 30 is a white top, black shorts, boy, black shoes, waving hands, and open hands. In order to make the missing joint points in the human body posture diagram more matched with the partial joint points, and to avoid the uncontrolled posture of the generated object in the generated human body posture diagram, the embodiment of the present application uses the torso posture (i.e., the posture skeleton diagram) corresponding to the three-dimensional incomplete posture data 10 as a constraint, and jointly generates a human body posture based on the torso posture corresponding to the three-dimensional incomplete posture data 10 according to the posture description text 30 and the three-dimensional incomplete posture data 10, thereby completing the limb posture.

[0046] In summary, the method provided in this embodiment obtains three-dimensional incomplete posture data and posture description text corresponding to a three-dimensional object; maps the three-dimensional incomplete posture data to a two-dimensional plane to obtain a posture skeleton diagram; and generates a human posture diagram having a human torso posture and limb posture based on the posture description text and the posture skeleton diagram. When faced with a three-dimensional object without a limb posture, this application completes the limb posture of the three-dimensional object using the posture description text based on the torso posture corresponding to the three-dimensional incomplete posture data, thereby improving the efficiency of completing the posture data of the three-dimensional object and also improving the utilization rate of open-source three-dimensional incomplete posture data without limb posture.

[0047] 2 shows a schematic diagram of the architecture of a computer system provided by an embodiment of the present application. The computer system may include: a terminal 100 and a server 200.

[0048] The terminal 100 can be an electronic device such as a mobile phone, a tablet computer, a vehicle-mounted terminal (vehicle computer), a wearable device, a personal computer (PC), an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, an unmanned vending terminal, etc. The terminal 100 can be installed with a client that runs a target application. The target application can be an application that supports three-dimensional object display, or other applications that support three-dimensional object modeling, three-dimensional object rendering, and three-dimensional object storage. This application does not limit this. In addition, this application does not limit the form of the target application, including but not limited to applications (Application, App) installed in the terminal 100, mini-programs, etc., and can also be in the form of a web page.

[0049] Server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data. Server 200 can be the backend server of the target application described above, used to provide backend services to the client of the target application.

[0050] The terminal 100 and the server 200 may communicate with each other via a network, such as a wired or wireless network.

[0051] In the method for completing the posture data of a three-dimensional object provided in an embodiment of the present application, the execution subject of each step can be a computer device, which refers to an electronic device with data calculation, processing and storage capabilities. Taking the solution implementation environment shown in Figure 2 as an example, the method for completing the posture data of a three-dimensional object can be executed by the terminal 100 (such as the client of the target application installed and running in the terminal 100 executing the method for completing the posture data of the three-dimensional object), or the method for completing the posture data of the three-dimensional object can be executed by the server 200, or the terminal 100 and the server 200 can interact and cooperate to execute the method, and this application does not limit this.

[0052] FIG3 is a flowchart of a method for completing the posture data of a three-dimensional object provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be the terminal 100 or the server 200 in FIG2. The method includes the following steps.

[0053] Step 220: Acquire three-dimensional incomplete posture data of the three-dimensional object in a preset posture, where the three-dimensional incomplete posture data includes first three-dimensional joint point data of some joint points of the three-dimensional object.

[0054] A 3D object is a 3D model of a virtual object. A virtual object may include at least one of the following: a human body, an anthropomorphic creature, an animal, a plant, or a fictional creature. Alternatively, a virtual object may include at least one of the following: a building, a vehicle, an object, or a terrain.

[0055] A 3D object is composed of at least two joints (or key points). For example, if the 3D object is a human body, its 3D model includes the joints of the human body; for another example, if the 3D object is a building, its 3D model includes the key points of the building.

[0056] Joints of a 3D object are nodes within a 3D model that allow for operations such as rotation and translation. These nodes are typically located at joints within the model, such as the shoulders, elbows, and hips of the human body. These joints enable flexible transformations of 3D objects, facilitating applications such as animation and game design. In 3D modeling software, joints are typically calculated and determined through a series of geometric operations and algorithms, ensuring that 3D objects maintain a natural, smooth, and continuous trajectory during motion.

[0057] When the three-dimensional object is a human body, the three-dimensional object may include at least one of the following types of joints: head joints, neck joints, left and right shoulder joints, spine joints, waist joints, left and right elbow joints, left and right wrist joints, left and right finger joints, left and right hip joints, left and right knee joints, left and right ankle joints, and left and right foot joints. Each type of joint may include at least one joint.

[0058] The pose data of a 3D object in a preset pose includes joint point data of the 3D object. A frame of pose data for a 3D object includes at least one of the following joint point data: the 3D position coordinates of each joint point, the joint rotation angles of each joint point, and the joint connection relationships of each joint point. The preset pose can be a known quantity, allowing for input of a pose description text that accurately indicates the preset pose, thereby ensuring accurate generation of a complete 2D pose image in the preset pose.

[0059] Alternatively, the pose data of the 3D object can be directly obtained from the model data of the 3D model. For example, the model data of the 3D model includes: pose data (joint point data), vertex information (coordinates, normal vectors, etc.), patch information, topology information (connectivity between patches, relationships between bones and meshes, etc.), texture data, material data, skeletal animation data, model annotations, attribute information, etc.

[0060] Optionally, if the model data of the three-dimensional model does not include pose data (joint point data), the computer device may obtain the pose data based on the model data of the three-dimensional model. For example, the three-dimensional model may be rendered and the joint point data of each joint may be identified and extracted from the rendering result to obtain the pose data.

[0061] Three-dimensional incomplete posture data is posture data that lacks the joint point data of at least one joint point, that is, posture data that includes the joint point data of some joint points. The three-dimensional joint point data of a joint point can include multiple data. The lack of three-dimensional joint point data of a certain joint point can be that there is no three-dimensional joint point data about the joint point, or it can be that there is three-dimensional node data of the joint point, but the three-dimensional node data of the joint point is incomplete, and the embodiments of the present application do not limit this. For example, the complete posture data should include the three-dimensional joint point data of 24 joint points, and the three-dimensional incomplete posture data may only include the joint point data of 20 joint points. The joint points missing in the three-dimensional incomplete posture data can be the same type of joint points, or can include at least two types of joint points. For example, the three-dimensional incomplete posture data lacks hand (wrist and / or finger) joint points.

[0062] Optionally, the joint points that are already included in the three-dimensional incomplete posture data and whose joint point data are complete can be called partial joint points; the joint points that are missing in the three-dimensional incomplete posture data or whose joint point data are incomplete can be called missing joint points, that is, the missing joint points are the joint points of the three-dimensional object other than the partial joint points.

[0063] For example, the missing joints include at least one hand joint, or at least one foot joint, or at least one elbow joint.

[0064] Optionally, the partial joint points include at least two joint points that can indicate a rough outline of the posture of the three-dimensional object. The missing joint points include at least one joint point used to refine the posture details of the three-dimensional object.

[0065] In an optional embodiment, the partial joints include trunk joints, and the missing joints include limb joints. The trunk joints include at least one of the head joints, neck joints, chest joints, waist joints, and limb joints. The limb joints include at least one of the hand joints, wrist joints, finger joints, foot joints, toe joints, and so on.

[0066] Alternatively, the incomplete 3D pose data may include all joints, but the joint data for one or more joints may be missing. For example, the incomplete 3D pose data may include the joint rotation angles of the hand joints, but lack the 3D position coordinates of the hand joints.

[0067] Among them, the joint point data can also be called three-dimensional joint point data, and the three-dimensional joint point data includes at least one of the following data: three-dimensional coordinates of the joint point, joint point rotation angle, joint point connection relationship, joint point name, and joint point identification.

[0068] Step 240: calling the text-based image model to generate a two-dimensional posture image of the three-dimensional object in a preset posture according to the three-dimensional incomplete posture data and the posture description text, where the posture description text is used to describe the preset posture.

[0069] The text-to-image model can also be called a text-to-image model, which is a neural network model that can generate two-dimensional images based on input text.

[0070] The text-based graph model is a multimodal deep learning model that can generate a matching 2D image from a description. Its core principle is to transform natural language text into an image space and simultaneously link visual features with linguistic information to achieve a mapping between natural language text and images. The specific operation of the text-based graph model is as follows: the text description is encoded into a feature vector. A generator network is then used to synthesize an image corresponding to this feature vector. Text-based graph models can be used in a variety of applications, such as generating realistic product images for e-commerce websites, creating visual aids for people with disabilities, generating images for virtual and augmented reality applications, and creating CAPTCHA images. There are also various types of text-based graph models, such as those based on GANs (Generative Adversarial Networks) and SD (Stable Diffusion) text-based graph models. These models are trained using large datasets of paired descriptions and corresponding 2D images. After training, they are able to generate 2D images based on new descriptions.

[0071] For example, the SD model can be a Wensheng graph model. The SD model is an image generation model based on the diffusion process, which can generate high-quality, high-resolution images. It simulates the diffusion process and gradually removes the noise image (random noise matrix) to obtain the target image. This model has strong stability and controllability, and can generate images with diverse effects and good visual effects.

[0072] Optionally, the text graph model can generate a two-dimensional posture image based on the posture description text and the three-dimensional incomplete posture data. The two-dimensional posture image is generated based on the posture description text. The two-dimensional posture image can include a generated object. The generated object presents the preset posture indicated by the state description text, and the posture presented by the generated object also matches the preset posture in the three-dimensional incomplete posture data. The generated object can be an object of the same type as the three-dimensional object, or the generated object can be the same as the three-dimensional object. The generated object is generated based on the posture description text. The closer the description in the posture description text is to the three-dimensional object, the more similar the generated object is to the three-dimensional object.

[0073] The text-based graph model generates a two-dimensional posture image based on the posture description text. During the generation process, the posture indicated in the three-dimensional incomplete posture data is used as a constraint, so that the posture of the generated object in the two-dimensional posture image is close to the preset posture in the three-dimensional incomplete posture data.

[0074] The posture description text includes at least one descriptive word for describing the posture of the three-dimensional object. The posture description text is used to describe the preset posture of the three-dimensional object in the incomplete three-dimensional posture data. It should be noted that the incomplete three-dimensional posture data in step 220 is a frame of posture data, or the incomplete three-dimensional posture data is a frame of posture data in a posture sequence of a continuous movement of the three-dimensional object. The posture description text may also include at least one descriptive word for describing the continuous movement.

[0075] For example, the descriptive words in the gesture description text may include at least one of the following: action descriptive words, movement descriptive words, limb state descriptive words, movement speed descriptive words, movement style descriptive words, movement purpose descriptive words, and the like.

[0076] In addition, to make the generated object closely resemble the three-dimensional object, the posture description text may also include at least one descriptive word for the three-dimensional object. For example, the posture description text may include a descriptive word for the three-dimensional object's appearance, clothing, gender, age, or personality.

[0077] It is important to note that the order of the descriptors in the pose description text is also very important, because their order affects the weight of the generated image. Generally speaking, the more descriptors are placed at the beginning, the greater the weight, and the more descriptors are placed at the end, the smaller the weight.

[0078] Exemplarily, the posture description text includes positive descriptive words and negative descriptive words; the positive descriptive words include the positive requirement text of the two-dimensional posture image, and the positive descriptive words include the preset posture; the negative descriptive words include at least one descriptive word for describing image defects, and the negative descriptive words are used to guide the text-based image model to avoid generating defective images with image defects.

[0079] For example, negative descriptors (exclusion words) are used to describe unwanted content in a two-dimensional pose image, such as low quality, watermarks, skin blemishes, and the like.

[0080] For example, positive descriptors could be: masterpiece, best quality, ultra detailed, best shadows, HD, high resolution, best details, 1 person, perfect hands, white t-shirt, black shorts, black shoes, short black hair, simple background, white background, a man kicking something or someone with his left leg.

[0081] Negative descriptors could be: worst quality: 2, low quality: 2, normal quality: 2, low quality, normal quality, monochrome: 1.2, grayscale: 1.2, skin spots, acne, skin blemishes, unusual number of fingers, unusual number of limbs, twisted joints, unusual number of organs, skin damage, unusual body fat status, unsuitable for office hours, hair accessories, selfie, bad anatomy, text, error, extra digits, fewer digits, cropped, worst quality, low quality, normal quality, jpeg artifacts, signature, watermark, username, blurry.

[0082] The embodiment of the present application includes positive descriptive words and negative descriptive words in the posture description text, so that the text-generated graph model can generate a two-dimensional posture image reflecting the preset posture according to the positive descriptive words, and at the same time avoid the appearance of unwanted content in the two-dimensional posture image according to the negative descriptive words, thereby ensuring that the generated two-dimensional posture image better meets the requirements.

[0083] Step 260: Perform joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of the missing joint points of the three-dimensional object except for some joint points.

[0084] Alternatively, a joint point recognition algorithm for a three-dimensional object can be used to identify all joint points of the three-dimensional object from the two-dimensional pose image, and missing joint points can be found from all the identified joint points.

[0085] The joint recognition algorithm can obtain 3D joint data for each joint based on 2D pose images. A joint recognition algorithm is a neural network model specifically trained to identify joints from 2D images. During training, the algorithm learns the positional and connection relationships between each joint. Based on the inherent connections between joints, it identifies each joint from the input 2D image and outputs 3D joint data for each joint.

[0086] For the joint point recognition of a single part, a dedicated joint point recognition algorithm can be trained. For example, for the task of recognizing the hand joint points, a hand joint point recognition algorithm (hand posture estimation algorithm) can be trained.

[0087] Optionally, when the missing joint points of a three-dimensional object are concentrated in a certain part, the target area where the target part is located in the two-dimensional posture image can be identified first, the target area in the two-dimensional posture image is intercepted as the target image, and the missing joint points are identified using the joint point recognition algorithm of the target part.

[0088] For example, when the missing joints are hand joints, the ACR (Attention Collaboration-based Regressor) hand posture estimation algorithm can be used to identify the hand joints in the two-dimensional posture image to obtain the joint data of the missing hand joints.

[0089] Since the text-based graph model can generate an image with the posture (two-dimensional posture image) based on the posture description text, and can use three-dimensional incomplete posture data to constrain the posture in the two-dimensional posture image, the text-based graph model can generate a two-dimensional posture image that is closer to the preset posture.

[0090] If the object is generated in a complete posture in the generated two-dimensional posture image, then by performing joint point extraction on the two-dimensional posture image, the complete joint point distribution under the preset posture of the three-dimensional object can be extracted. Therefore, based on the two-dimensional posture image, the second three-dimensional joint point data can be identified and extracted, and then the identified second three-dimensional joint point data can be used to complete the three-dimensional incomplete posture data.

[0091] Taking the three-dimensional object as a human body and the missing joint points as hand joint points as an example, as shown in FIG4 , the two-dimensional posture image 301 contains the complete posture of the human body. By estimating the hand posture of the human hand 302 in the two-dimensional posture image 301 , the three-dimensional joint point data of the hand joint points can be obtained.

[0092] Step 280: Add the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

[0093] Exemplarily, the identified second three-dimensional joint point data is filled into the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data. It should be noted that if the missing joint point is a joint point that is completely not included in the three-dimensional incomplete posture data, then when adding the second three-dimensional joint point data to the three-dimensional incomplete posture data, the missing joint point can be directly matched with some joint points, and the second three-dimensional joint point data is added to the appropriate position in the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data. If some of the missing joint points are joint points included in the three-dimensional incomplete posture data, but their joint point data is incomplete, then when adding the second three-dimensional joint point data to the three-dimensional incomplete posture data, it is necessary to match the missing joint point with some joint points, and add the second three-dimensional joint point data to the appropriate position in the three-dimensional incomplete posture data. Moreover, since the three-dimensional incomplete posture data contains some joint point data of the missing joint point, it is also necessary to perform deduplication when adding to avoid duplicate joint point data.

[0094] The 3D completed posture data includes 3D joint point data of all joint points of the 3D object, that is, the 3D completed posture data includes: the first 3D joint point data of the 3D object, and some joint point data of missing joint points.

[0095] In summary, the method provided in this embodiment obtains three-dimensional incomplete pose data and pose description text for a three-dimensional object in a preset pose. The three-dimensional incomplete pose data includes first three-dimensional joint point data for some of the joints of the three-dimensional object, and the pose description text is used to describe the preset pose. The three-dimensional incomplete pose data and pose description text are then input into a Vincent graph model. Because the Vincent graph model has the characteristic of accurately controlling image generation based on text, the Vincent graph model can generate a two-dimensional pose image of the three-dimensional object in the preset pose based on the three-dimensional incomplete pose data and pose description text. The generated object in the two-dimensional pose image has a complete preset pose and can reflect the pose of the missing joint points. In this way, by identifying and extracting the joint points of the generated object in the two-dimensional pose image, second three-dimensional joint point data for the missing joint points can be obtained. By adding the second three-dimensional joint point data to the three-dimensional incomplete pose data, the three-dimensional incomplete pose data can be completed, resulting in three-dimensional completed pose data of the three-dimensional object, i.e., complete pose data. This application targets three-dimensional objects with postures that do not have missing joints. Based on the posture data and posture description text of some joints in the three-dimensional incomplete posture data, a two-dimensional posture image with a complete preset posture can be directly generated through a text graph model. The posture of the three-dimensional object can then be completed using the second three-dimensional joint data extracted from the two-dimensional posture image. This eliminates the need to collect additional postures of missing joints to construct an action library, significantly reducing the consumption of manpower and material resources, improving the efficiency of completing the posture data of the three-dimensional object, and also increasing the utilization rate of open-source three-dimensional posture data for limbless postures. Furthermore, the posture data completed by this method can perfectly match the original incomplete posture data of the three-dimensional object, improving the effectiveness of completing the posture data of the three-dimensional object.

[0096] An exemplary embodiment of generating a two-dimensional gesture image using a text graph model with a gesture control plug-in is given below.

[0097] FIG5 is a flowchart of a method for completing the pose data of a three-dimensional object provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be the terminal 100 or the server 200 in FIG2 . Based on the embodiment shown in FIG3 , step 240 includes steps 241 and 242.

[0098] Step 220: Acquire three-dimensional incomplete posture data of the three-dimensional object in a preset posture, where the three-dimensional incomplete posture data includes first three-dimensional joint point data of some joint points of the three-dimensional object.

[0099] For example, if the three-dimensional object is a human body, the three-dimensional incomplete posture data of the human body can be obtained. The three-dimensional incomplete posture data includes the first three-dimensional joint point data of the human body, and some joint points include the joint points of the head, torso and limbs of the human body; the three-dimensional incomplete posture data lacks the second three-dimensional joint point data of the three-dimensional human body, and the missing joint points include the hand joint points of the human body.

[0100] Step 241: Map the three-dimensional incomplete posture data to a two-dimensional plane to obtain a posture skeleton graph.

[0101] The posture skeleton diagram is a two-dimensional image that marks the positions of some joints and the connection relationships between some joints.

[0102] Optionally, a pose skeleton diagram can be drawn according to the bone and joint colors specified in the OpenPose algorithm. As shown in (1) in Figure 6, the pose skeleton diagram uses dots of different colors to mark the locations of the joints, and dots of different colors represent different joints. Line segments of different colors are used to connect the joints, and line segments of different colors are used to indicate different parts of the human body. More clearly, as shown in (2) in Figure 6, the pose skeleton diagram can intuitively indicate the location and connection relationship of each joint.

[0103] Optionally, the two-dimensional plane can be a two-dimensional imaging plane of a virtual camera. The computer device maps the three-dimensional incomplete posture data to the two-dimensional imaging plane of the virtual camera according to the virtual camera parameters to obtain a joint point graph, which includes two-dimensional joint points of at least two partial joint points; according to the joint point connection relationship of at least two partial joint points, the two-dimensional joint point coordinates of at least two partial joint points in the joint point graph are connected to obtain a posture skeleton graph.

[0104] The virtual camera parameters include at least one of the coordinates of the virtual camera, the position of the virtual camera relative to the three-dimensional object, and the built-in parameters of the virtual camera.

[0105] For example, the following virtual camera parameters can be used: resolution 512*512, focal length 50mm, sensor size 36mm, distance from the root node of the three-dimensional object is 8m, flush with the waist Root node, and shooting direction perpendicular to the body plane of the three-dimensional object (the body plane can be a plane determined based on the two shoulder joints and the waist root node).

[0106] In the embodiment of the present application, the three-dimensional incomplete posture data is projected onto the two-dimensional imaging plane of the virtual camera to obtain a posture skeleton diagram. Since the setting of the virtual camera can simulate the shooting process of a real camera, the three-dimensional incomplete posture data is projected onto the two-dimensional imaging plane of the virtual camera, so that an accurate joint point diagram can be obtained, and then a more accurate posture skeleton diagram can be obtained, thereby improving the generation accuracy of the two-dimensional posture image.

[0107] Step 242: Input the posture description text into the text graph model, call the posture control plug-in to constrain the image generation process of the text graph model according to the posture skeleton graph, and obtain a two-dimensional posture image.

[0108] Among them, the postures of some joint points in the two-dimensional posture image are consistent with the posture skeleton graph.

[0109] Exemplarily, the Wensheng graph model includes a posture control plug-in, and the posture control plug-in is used to constrain the two-dimensional posture image generated by the Wensheng graph model according to the posture skeleton graph.

[0110] The posture control plug-in uses the principle shown in FIG7 to constrain the image generation process of the Wensheng graph model. As shown in FIG7 (1), there is a neural network block 303 in the Wensheng graph model, and the input of the neural network block 303 is x and the output is y. When the posture control plug-in is used to constrain the generation process, as shown in FIG7 (2), the network of the posture control plug-in includes a trainable copy 304 of the neural network block and two zero convolution layers. Among them, the trainable copy 304 of the neural network block is obtained by directly copying the neural network block 303 of the Wensheng graph model. During the application process, the constraint condition c (posture skeleton graph) is input into the first zero convolution layer, and then the output of the first zero convolution layer is added to the input x, and the addition result is input into the trainable copy 304, and then the output of the trainable copy 304 is input into the second zero convolution layer, and the output of the second zero convolution layer is added to the original output y of the neural network block 303 to obtain the final output y'. In this way, the posture control plug-in can constrain the image generation process of the Wensheng graph model, so that the constraint condition c can constrain the output data of the Wensheng graph model.

[0111] It should be noted that when training the posture control plug-in, the network parameters within neural network block 303 in the Vincent graph model are locked (fixed parameters do not participate in parameter adjustment during training). The training samples are then used to adjust the network parameters in the two zero convolution layers and the trainable replica in the posture control plug-in, allowing the Vincent graph model to output the target output of the training samples within the constraints of the posture control plug-in. When initializing the posture control plug-in, the network parameters within the two zero convolution layers are set to 0, and the initial parameters of the trainable replica 304 are the same as those of neural network block 303.

[0112] In an optional embodiment, as shown in FIG8 , the Wensheng graph model includes a first network 401 and a second network 402; the posture control plug-in includes a first zero-convolution layer 404, a network copy 403 of the first network, and a second zero-convolution layer 405; the network copy 403 is a network obtained by initialization and training using the network structure and network parameters of the first network 401, that is, the network copy 403 has the same network structure as the first network 401, but the network parameters are not necessarily the same; step 242 may include the following steps:

[0113] 1) Input the gesture description text into the first network 401 to obtain text features. Optionally, input the gesture description text into the text encoder 406 to obtain a text encoding result, and then input the text encoding result into the first network 401 to obtain text features.

[0114] 2) Input the posture skeleton image into the first zero convolution layer 404 to obtain the posture convolution result.

[0115] 3) The pose convolution result is added to the random noise matrix to obtain the constrained noise matrix. The random noise matrix is ​​a random matrix that conforms to a Gaussian distribution. The image-text generation model denoises the random noise matrix based on the input pose description text to obtain the final 2D pose image.

[0116] 4) Input the constraint noise matrix and the posture skeleton graph into the network copy 403 to obtain the first constraint feature.

[0117] 5) Input the first constraint feature into the second zero convolution layer 405 to obtain the second constraint feature.

[0118] 6) Add the second constraint feature and the text feature to obtain the text constraint feature.

[0119] 7) Input the text constraint features and the gesture description text into the second network 402 to obtain a two-dimensional gesture image. Optionally, input the text constraint features and the text encoding result into the second network to obtain a two-dimensional gesture image.

[0120] In an embodiment of the present application, a multi-layer structure of a posture control plug-in is used to generate a second constraint feature based on a posture skeleton graph, and then the text features of the posture description text are combined as text constraint features to control the posture of the object generated in the two-dimensional posture image to be consistent with the posture skeleton graph through the text constraint features, thereby ensuring the constraint effect, making the missing joint points identified more closely match the partial joint points, and improving the efficiency and effect of posture data completion of three-dimensional objects.

[0121] Exemplarily, the first network includes at least one encoder; and the second network includes at least one decoder.

[0122] In an optional embodiment, as shown in FIG9 , the Wensheng graph model can adopt the SD model, and the posture control plug-in can adopt the OpenPose mode in the Control Net plug-in in the SD model. That is, the Wensheng graph model includes a text encoder, encoder 1, encoder 2, encoder 3, encoder 4, an intermediate network, decoder 4, decoder 3, decoder 2, and decoder 1. The posture control plug-in includes a first zero convolution layer, a copy of encoder 1, a copy of encoder 2, a copy of encoder 3, a copy of encoder 4, a copy of the intermediate network, a second zero convolution layer 1, a second zero convolution layer 2, a second zero convolution layer 3, a second zero convolution layer 4, and a second zero convolution layer 5.

[0123] The process of using the Wensheng graph model and the posture control plug-in to obtain a 2D posture image is as follows:

[0124] (1) Input the posture description text into the text encoder to obtain the text encoding result. Then input the text encoding result and the random noise matrix into encoder 1, input the text encoding result and the output of encoder 1 into encoder 2, input the text encoding result and the output of encoder 2 into encoder 3, and input the text encoding result and the output of encoder 3 into encoder 4. The random noise matrix can be 64*64 dimensional data, the output of encoder 1 can be 32*32 dimensional data, the output of encoder 2 can be 16*16 dimensional data, the output of encoder 3 can be 8*8 dimensional data, and the output of encoder 4 can be 8*8 dimensional data.

[0125] (2) Input the posture skeleton image into the first zero convolution layer. Add the output of the first zero convolution layer to the random noise matrix. Input the addition result and the text encoding result into the encoder 1 replica, input the text encoding result and the output of the encoder 1 replica into the encoder 2 replica, input the text encoding result and the output of the encoder 2 replica into the encoder 3 replica, input the text encoding result and the output of the encoder 3 replica into the encoder 4 replica, and input the text encoding result and the output of the encoder 4 replica into the intermediate network replica.

[0126] (3) The output of the intermediate network copy is input into the second zero convolution layer 1. The text encoding result and the output of the encoder 4 are input into the intermediate network. The output of the second zero convolution layer 1 is added to the output of at least one layer of network blocks in the intermediate network (see the process illustrated in Figure 7), so that the intermediate network obtains the output of the intermediate network based on the result of the addition.

[0127] (4) The output of the encoder 4 copy is input to the second zero convolution layer 2. The text encoding result and the output of the intermediate network are input to the decoder 4. The output of the second zero convolution layer 2 is added to the output of at least one layer of network blocks in the decoder 4 (see the process illustrated in Figure 7), so that the decoder 4 obtains the output of the decoder 4 according to the result of the addition.

[0128] (5) The output of the copy of encoder 3 is input to the second zero convolution layer 3. The text encoding result and the output of decoder 4 are input to decoder 3. The output of the second zero convolution layer 3 is added to the output of at least one layer of network blocks in decoder 3 (see the process illustrated in Figure 7), so that decoder 3 obtains the output of decoder 3 based on the result of the addition.

[0129] (6) The output of the copy of encoder 2 is input to the second zero convolution layer 4. The text encoding result and the output of decoder 3 are input to decoder 2. The output of the second zero convolution layer 4 is added to the output of at least one layer of network blocks in decoder 2 (see the process illustrated in Figure 7), so that decoder 2 obtains the output of decoder 2 based on the result of the addition.

[0130] (7) The output of the copy of encoder 1 is input to the second zero convolution layer 5. The text encoding result and the output of decoder 2 are input to decoder 1. The output of the second zero convolution layer 5 is added to the output of at least one layer of network blocks in decoder 1 (see the process illustrated in Figure 7), so that decoder 1 obtains the output of decoder 1 based on the result of the addition.

[0131] Optionally, the output result of decoder 1 can be used as input 407 again (replacing the random noise matrix in the above process), and the above processes (1) to (7) can be iteratively repeated until the number of iterations meets the number threshold, and the output of the last decoder 1 is used as the two-dimensional posture image finally output by the Wensheng graph model.

[0132] In the embodiment of the present application, the first network is configured to include at least one encoder, and the second network is configured to include at least one decoder. Thus, in the image generation process of the Wensheng graph model constrained by the posture control plug-in, the encoder can reduce the dimension of the data. The encoder can remove redundant information and retain key features, thereby improving the speed and accuracy of subsequent processing. The decoder can accurately capture and reconstruct the structure and details of the input data, thereby ensuring the accuracy of the output data, thereby improving the data processing efficiency and accuracy throughout the process and enhancing the generalization ability of the network.

[0133] Step 260: Perform joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of the missing joint points of the three-dimensional object except for some joint points.

[0134] Step 280: Add the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

[0135] In summary, the method provided in this embodiment uses a pose control plug-in to constrain the image generation process of the Vincent graph model, enabling the Vincent graph model to generate a 2D pose image based on the constraints of the pose skeleton graph. This ensures that the pose of the object generated in the 2D pose image is consistent with the pose skeleton graph. This allows the identified missing joint points to be more closely aligned with the partial joint points, improving the efficiency and effectiveness of pose data completion for 3D objects. It also increases the utilization of open-source 3D pose data for limbless poses.

[0136] For example, the posture of the three-dimensional object can be a frame posture in a continuous action. In order to make the previous and next postures in the continuous action of the three-dimensional object smoother, the similarity of the completed posture will be determined based on the previous and next frame postures. If the similarity is poor, the completed posture can be regenerated.

[0137] For example, after completing each posture in a continuous action using the above method, each posture in the continuous action may be smoothed to further improve the action continuity.

[0138] Figure 10 is a flowchart of a method for completing the pose data of a three-dimensional object provided by an exemplary embodiment of the present application. This method can be executed by a computer device, which may be terminal 100 or server 200 in Figure 2. Based on the embodiment shown in Figure 3, step 260 may be followed by step 270, and / or step 280 may be followed by step 290.

[0139] Step 220: Acquire three-dimensional incomplete posture data of the three-dimensional object in a preset posture, where the three-dimensional incomplete posture data includes first three-dimensional joint point data of some joint points of the three-dimensional object.

[0140] Exemplarily, the incomplete 3D pose data is a frame of pose data in a motion sequence of a 3D object, where the motion sequence includes at least two frames of pose data. That is, step 220 may involve obtaining the i-th frame of incomplete 3D pose data in the motion sequence of the 3D object. The motion sequence includes n frames of pose data, where i is a positive integer not greater than n, and n is a positive integer.

[0141] For example, a set of motion sequences of a three-dimensional object is obtained, where the set of motion sequences includes at least two frames of pose data. At least one frame of pose data lacks a joint point. Alternatively, each frame of pose data in the motion sequence lacks a joint point. Optionally, the joint points missing from each frame of pose data may be the same or different.

[0142] The present application embodiment uses the example of a situation where each frame of posture data in an action sequence lacks hand joint points. To complete the hand joint points in each frame of posture data in the action sequence, the method described in the aforementioned embodiment can be used to complete the hand joint points in each frame of posture data in the action sequence.

[0143] In order to make the hand movements smoother after the action sequence is completed, step 270 and / or step 290 provided in this embodiment may be used for smoothing.

[0144] Step 240: calling the text-based image model to generate a two-dimensional posture image of the three-dimensional object in a preset posture according to the three-dimensional incomplete posture data and the posture description text, where the posture description text is used to describe the preset posture.

[0145] For the i-th frame of incomplete 3D pose data, a text-based model is invoked to generate the i-th frame of a 2D pose image based on the i-th frame of incomplete 3D pose data and the pose description text. The pose description text is used to describe the actions in the action sequence. The pose of the object generated in the i-th frame of the 2D pose image is the same as the preset pose in the i-th frame of incomplete 3D pose data.

[0146] Step 260: Perform joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of the missing joint points of the three-dimensional object except for some joint points.

[0147] Perform joint point recognition on the i-th frame of the two-dimensional posture image to obtain second three-dimensional joint point data in the i-th frame of three-dimensional incomplete posture data.

[0148] Step 270: Calculate the posture similarity between the first joint point data and the second joint point data, and when the posture similarity between the first joint point data and the second joint point data is less than the similarity threshold, re-execute the following steps until the posture similarity is not less than the similarity threshold.

[0149] Among them, re-executing the following steps may be re-executing steps 240, 260 and 270, and executing step 280 when the posture similarity between the first joint point data and the second joint point data is not less than the similarity threshold.

[0150] Among them, the first joint point data includes the second three-dimensional joint point data in the two-dimensional posture image, and the second joint point data includes the second three-dimensional joint point data in the historical posture data; the historical posture data includes at least one frame of posture data in the action sequence that is located before the three-dimensional incomplete posture data.

[0151] Specifically, the similarity between the missing joint points in the i-th frame and the missing joint points in the i-1th frame is calculated. If the similarity is high, the missing joint points for the next frame (i+1th frame) are generated. If the similarity is low, the missing joint points for the i-th frame are regenerated. The missing joint points in the i-th frame are the second 3D joint point data derived from the incomplete 3D pose data in the i-th frame. The missing joint points in the i-1th frame are the second 3D joint point data derived from the incomplete 3D pose data in the i-1th frame of the action sequence.

[0152] It should be noted that when i is 1, step 270 of calculating the posture similarity between two adjacent frames may not be performed.

[0153] In summary, when the three-dimensional incomplete posture data is a frame of posture data in the action sequence of a three-dimensional object, the method provided in this embodiment can, after generating the missing joint points of each frame, also perform posture similarity matching on the missing joint points of the current frame based on the missing joint points of the previous frame. If the posture difference between the missing joint points of the current frame and the missing joint points of the previous frame is too large, the missing joint points of the current frame are regenerated until the posture difference between the missing joint points of the current frame and the missing joint points of the previous frame is less than the threshold. In this way, the coherence and smoothness of the missing joint points in an action sequence can be ensured, thereby improving the joint point completion effect. Exemplarily, the three-dimensional joint point data includes three-dimensional position coordinates and joint rotation angles. A method for calculating posture similarity is given below:

[0154] A completion matrix is ​​obtained based on the three-dimensional position coordinates of the first joint point data; and a history matrix is ​​obtained based on the three-dimensional position coordinates of the second joint point data; the cosine similarity of the completion matrix and the history matrix is ​​calculated to obtain a first similarity; and the difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data is calculated to obtain a second similarity; the first similarity and the second similarity are weightedly summed to obtain posture similarity.

[0155] The embodiment of the present application performs calculations from two dimensions: joint point coordinates and joint point rotation angles. The joint point coordinates directly provide the position information of the joints in space, and the joint point rotation angles reflect the relative motion relationship between the joints. Therefore, when calculating the posture similarity, the position information and relative motion relationship of the joints are comprehensively considered, thereby improving the accuracy of the posture similarity calculation.

[0156] In one possible implementation, the number of missing joint points is at least two, and the second similarity can be calculated according to the following method: calculate the difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data to obtain the joint rotation angle difference of each missing joint point; and perform weighted summation on at least two joint rotation angle differences according to the weight of each missing joint point in the missing joint points to obtain the second similarity.

[0157] In the embodiment of the present application, when there are multiple missing joint points, since each missing joint point has its own joint rotation angle difference, the weighted sum of at least two joint rotation angle differences can be performed based on the weight of each missing joint point, and the second similarity can be obtained more accurately.

[0158] It is understandable that the weight of each missing joint point can be pre-set, and the weight of each missing joint point can be related to the importance of the missing joint point. In a possible implementation, the importance of the missing joint point is related to the type of the missing joint point. The type of the missing joint point can include a parent node and a child node. In some cases, the importance of the parent node is higher than the importance of the child node. Therefore, when setting the weight, the weight of the parent node in the missing joint point can be set higher than the weight of the child node; the number of joint points between the parent node and the root node in the three-dimensional object is the first number, and the amount of joint point data between the child node and the root node is the second number, and the first number is less than the second number. That is, of the two connected joint points, the joint point that is closer to the root node level is the parent node, and the joint point that is farther from the root node level is the child node. For example, the wrist joint is the parent node of the hand joint.

[0159] In the embodiment of the present application, the weight of the parent node is set to be higher than that of the child node, thereby setting a higher weight for the important nodes in the missing joint points, further improving the calculation accuracy of the second similarity.

[0160] By using the method of step 270, it is possible to ensure that the posture similarity of the missing joint points in two adjacent frames is greater than the similarity threshold, thereby ensuring the posture continuity of the two adjacent frames. For example, as shown in (1) in Figure 11, if the hand posture in the current frame 408 is too different from the hand posture in the previous frame 409, the hand posture of the current frame is regenerated. As shown in (2) in Figure 11, the regenerated hand posture 410 is less different from the hand posture in the previous frame 409, and the movement is smoother.

[0161] Step 280: Add the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

[0162] Optionally, the i-th frame of three-dimensional incomplete posture data is completed according to the second three-dimensional joint point data in the i-th frame of three-dimensional incomplete posture data to obtain the i-th frame of three-dimensional completed posture data.

[0163] Step 290: Smoothing the three-dimensional completed posture data according to the adjacent posture data in the action sequence to obtain three-dimensional smoothed posture data.

[0164] Among them, the adjacent posture data includes: at least one frame of posture data in the action sequence that is located before the three-dimensional incomplete posture data, and at least one frame of posture data in the action sequence that is located after the three-dimensional incomplete posture data; or, the adjacent posture data includes: at least one frame of posture data in the action sequence that is located before the three-dimensional incomplete posture data; or, the adjacent posture data includes: at least one frame of posture data in the action sequence that is located after the three-dimensional incomplete posture data.

[0165] Optionally, a smoothing algorithm may be used to smooth the second three-dimensional joint point data in the motion sequence to obtain smoothed three-dimensional posture data frame by frame. For example, the smoothing algorithm may be a moving average smoothing algorithm, an exponential smoothing algorithm, a median filter smoothing algorithm, a local polynomial smoothing algorithm, or the like.

[0166] The method provided in this embodiment can also use a smoothing algorithm to perform data smoothing on the missing joint points in the action sequence after completing the missing joint points in each frame in the action sequence, so that the posture transition of the missing joint points in the entire action sequence is smoother, ensuring that the action sequence obtained after completion has a better visual effect.

[0167] The following is an exemplary embodiment of using the method provided in the embodiment of the present application to complete the hand posture of a three-dimensional human body.

[0168] FIG12 is a flowchart of a method for completing the posture data of a three-dimensional object provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be the terminal 100 or the server 200 in FIG2. The method includes the following steps.

[0169] This embodiment mainly uses the Stable Diffusion graph framework and the OpenPose mode of the Control Net plug-in to complete the human hand posture.

[0170] Stable Diffusion (SD) is a text graph framework that uses a differentiable diffusion equation to simulate the process of gradually reducing noise to generate high-quality samples. At the same time, text description features are introduced in the process of diffusion denoising to control the probability distribution of denoising, thereby generating images related to the input text description. However, simply using text to control human body images is not constrained enough, and it is difficult to obtain a specified human body action image. Therefore, this embodiment also uses the OpenPose mode of the Control Net plug-in to further constrain the human body posture of the generated image. The principle of the Control Net plug-in is to insert a conditional control branch into the SD diffusion model to affect the generated image. The branch can input various forms of conditional control images, such as depth maps, edge maps, human body posture maps, semantic segmentation maps, etc., to achieve precise control of the generated image. The OpenPose mode is used here, that is, the human body posture map is input to control the human body posture in the SD text graph.

[0171] Step 901: 3D human joint reprojection and planar skeleton graph production: reproject the 3D human joint coordinates to a 2D plane, and draw a planar skeleton graph sequence according to the OpenPose skeleton graph standard.

[0172] Since Control Net can only input 2D planar skeleton images, preprocessing is required after obtaining 3D human joint coordinates (i.e., incomplete 3D pose data). Here, a 3D human pose frame is imported into Blender, and a virtual camera is set up with reasonable virtual camera parameters (including intrinsic and extrinsic parameters) so that the entire human skeleton is centered in the captured image. The intrinsic and extrinsic parameters of the virtual camera are also recorded. The reference values ​​for the virtual camera parameters used are as follows: 512*512 resolution, 50mm focal length, 36mm sensor size, 8m distance from the human skeleton, and alignment of the root node at the waist. The intrinsic and extrinsic parameter matrix of the virtual camera is then calculated based on these virtual camera parameters. The 3D human joint coordinates can then be projected onto the 2D imaging plane to obtain the corresponding 2D coordinates of the joints on the image. Following the bone and joint colors specified in the OpenPose algorithm, the joint points are drawn with the corresponding colors based on the previously obtained 2D joint coordinates on a 512*512 black image. The joints are then connected using the corresponding colors to form a planar skeleton image, known as the pose skeleton image.

[0173] Step 902: Generate description words according to the action sequence description and the text-based graph description word template.

[0174] Before generating images, a set of positive and negative descriptor templates for the textual image are first developed. This ensures that the images are generated as close to the positive descriptors as possible while minimizing the negative descriptors. Furthermore, when generating images of the human body for each action sequence, the positive descriptors are combined with the textual description of the action to ensure that the generated image is more consistent with the action, whether it is a single action like standing, walking, or kicking, or an interactive action like raising a glass.

[0175] The text image description word template includes a positive description word template and a negative description word template. The positive description word template includes at least one positive description word and a description text of a preset posture. The positive description word template can be a fixed template, while the description text of the preset posture can be changed according to different requirements. The negative description word template includes at least one negative description word, and the negative description word is a fixed template.

[0176] Optionally, the positive description word template used is:

[0177] "(masterpiece,best quality,ultra-detailed,best shadow),HD,high resolution,best details,1boy,perfect hand,white T-shirt,black short pants,black shoes,short black hair,simple background,white background,{action prompt}".

[0178] The {action prompt} along with the curly brackets needs to be replaced with a text description of the corresponding action, such as "a man kicks something or someone with his left leg". The final positive description word template is:

[0179] "(masterpiece,best quality,ultra-detailed,best shadow),HD,high resolution,best details,1boy,perfect hand,white T-shirt,black short pants,black shoes,short black hair,simple background,white background,a man kicks something or someone with his left leg."

[0180] Optionally, the negative descriptor template used is:

[0181] "(worst quality:2),(low quality:2),(normal quality:2),lowers,normal quality,(monochrome:1.2),(grayscale:1.2),skin spots,acnes,skin blemishes,jpeg artifacts,cropped,bad anatomy,nsfw,hair ornaments,selfie,lowres,text,error,worst quality,low quality,normal quality, signature, watermark, username, blurry, the number of fingers does not conform to common sense, the number of limbs does not conform to common sense, joint distortion, the number of organs does not conform to common, the skin damage, body fat percentage does not conform to common sense."

[0182] Step 903: Generate a two-dimensional posture image based on the stable diffusion Vincent graph model and the posture control plug-in.

[0183] Enter the positive and negative descriptors into the description word text box of the SD Wensheng graph framework. Extract a frame of planar skeleton image from the planar skeleton image sequence and enter it into the image selector of the Control Net plug-in in the SD framework. Select the OpenPose mode and None for the processor. The other parameters for the Wensheng graph are as follows: set the sampler to DPM (Data Processing Module)++2M a Karras, the sampling step to 20, the CFG (Configuration File) scale to 7, and the size to 512*512. Click Generate to obtain the human pose image (i.e., 2D pose image) for the specified action (i.e., preset pose), as shown in Figure 4.

[0184] Step 904: Extract hand motions from the two-dimensional posture image using the three-dimensional hand posture estimation.

[0185] After obtaining a human body pose image with a reasonable hand, 3D hand pose estimation is used to extract the hand movements in the 2D pose image, namely, the joint point data of each joint of the hands. The 3D hand pose estimation algorithm can be an ACR hand pose estimation algorithm, thereby obtaining the joint point data of each joint of the hands in the image.

[0186] Step 905: Calculate the similarity with the hand movement of the previous frame and determine whether the similarity is high enough.

[0187] In order to ensure that the hand movements are coherent and smooth throughout the entire action sequence, a temporal stability judgment mechanism is introduced in the process of generating each frame of image. The principle of this mechanism is to obtain the hand posture matrix in the image generated by the previous frame and the current frame, and then calculate the cosine similarity of the two matrices and the difference of each joint, and then combine all the values ​​​​weighted to obtain the hand movement similarity of the two frames of image. The weight of each joint is related to the importance of the joint. It is generally believed that the parent joint has a greater weight and the child joint has a smaller weight. When the similarity is too low, step 903 is re-executed to generate the current frame image until the difference between the hand movements of the previous and next frames is not too large. The effect is shown in Figure 11. It can be seen that due to the introduction of the temporal stability judgment mechanism, the hand movements in the previous and next frames are more coherent and reasonable. At this time, the current hand movement will be merged into the body posture for storage.

[0188] Step 906: After all frames are generated, the entire hand motion sequence is smoothed to obtain a complete full-body posture.

[0189] After completing the hand movement completion of the entire action, the movement sequence of both hands will be smoothed again to complete the hand movement completion process.

[0190] In summary, the method provided in this embodiment proposes a method based on Stable Diffusion graphs. It uses Control Net and a planar image of the human skeleton to impose strong constraints on the human torso pose in the generated image. Leveraging the rich prior information in the Stable Diffusion diffusion model, the method generates reasonable human images for this torso pose based on descriptors. Even actions interacting with objects can be constrained using descriptors, resulting in more reasonable and diverse hand poses. Furthermore, a mechanism for determining the temporal stability of hand poses is introduced, resulting in a smoother sequence of generated hand poses.

[0191] The method provided in this embodiment utilizes Stable Diffusion and ControlNet Graph methods to complete reasonable hand poses for 3D human pose data lacking hand poses, thereby achieving more refined driving of virtual humans. Furthermore, in the field of deep learning for human pose data generation, it can more effectively utilize open-source human pose datasets and reduce data acquisition costs. After obtaining a 3D human pose sequence without hand poses, it is first reprojected onto a 2D plane, and a planar skeleton graph for each frame is drawn according to the OpenPose human pose skeleton graph standard. Next, descriptors for the Graph are generated based on the action sequence's behavioral labels. The descriptors and all planar skeleton graphs are then input into the Stable Diffusion Graph framework equipped with the ControlNet plug-in to generate a planar image of the human body in the corresponding pose. At this point, a 3D hand pose estimation algorithm is used to obtain the pose of the human hand in the generated image, completing the pose data. Furthermore, to ensure smoother hand movements throughout the pose sequence, the similarity between the hand pose of the current frame and the previous frame is calculated while the current frame image is being generated, thereby eliminating the problem of drastic changes in hand pose.

[0192] The method provided in this embodiment does not require additional collection of hand movements or training of deep learning models when faced with body movement data without hand movements. Hand movement completion can be achieved by simply using the SD Wenshengtu framework. The generated movements are richer and more appropriate, which can effectively reduce manpower and material costs and improve the utilization rate of open source human posture data without hands.

[0193] The posture data completion method provided in the embodiments of the present application can be applied to applications with three-dimensional virtual objects, for example, to complete the posture of a three-dimensional virtual character in a game application, or to complete the joint points of a three-dimensional object or the key points of a three-dimensional terrain in a virtual reality (VR) / augmented reality (AR) application, or to complete the posture of a three-dimensional virtual anchor in a virtual live broadcast application, or to complete the posture of an AI virtual image in an artificial intelligence (AI) question-and-answer application, or to complete the posture of a three-dimensional animation character in an animation production application, and so on.

[0194] The following is an example of completing the hand gestures of a three-dimensional virtual character (three-dimensional object) in a game application. This method can be executed by the client of the game application or by the server of the game application.

[0195] The game application stores an action library. The action library is used to store a complete pose sequence of at least one action of a 3D virtual character. When rendering the 3D virtual character performing an action, the complete pose sequence of the target action can be directly read from the action library. The 3D virtual character is then rendered according to the complete pose sequence of the target action, thereby displaying the 3D virtual character performing the target action.

[0196] The incomplete posture sequence for each action can be manually drawn by the developer, who can manually determine the positions of the various body joints of the three-dimensional virtual character when performing an action, as well as the motion trajectory of each body joint. However, due to the large number of hand joints and the flexible motion trajectory of the hand joints, manually determining the positions and motion trajectories of the hand joints in an action is a huge workload, and the efficiency of manual execution is too low. Therefore, the method provided in the embodiment of the present application can be used to perform posture data completion based on the manually determined incomplete posture, and complete the hand posture to obtain a complete posture.

[0197] 1. Read a fragmented pose sequence of the target action (i.e., preset pose) from the action library. The fragmented pose sequence includes at least two frames of fragmented poses of the target action. Each frame of the fragmented pose sequence includes three-dimensional joint point data (i.e., first three-dimensional joint point data) of the body joints. The body joints include: head joints, neck joints, torso joints, and limb joints. The hand joints (wrist joints and finger joints) are missing from the body joints.

[0198] 2. Obtain a frame of incomplete pose from at least two frames of incomplete pose. Invoke the text-based image model to generate a two-dimensional pose image of the three-dimensional virtual character based on the incomplete pose data of the frame and the pose description text of the target action. The pose description text is used to describe the target action.

[0199] 3. Perform joint recognition on the two-dimensional posture image to obtain three-dimensional joint point data (second three-dimensional joint point data) of the hand joints of the three-dimensional virtual character.

[0200] 4. Complete the incomplete posture of the frame based on the 3D joint data of the hand joints to obtain a complete posture.

[0201] 5. Then continue to obtain the next frame of incomplete posture in the incomplete posture sequence, execute steps 1, 2, 3, 4, and 5 to complete each frame of the incomplete posture sequence in the target action and obtain the complete posture sequence of the target action.

[0202] The method of using the text-based graph model to complete the incomplete posture can refer to the method provided in any of the above embodiments, and will not be repeated here.

[0203] After completing the pose data for each action in the action library, the 3D virtual character can be controlled to perform each action according to the pose sequence stored in the action library. For example, when a trigger operation is received to control the 3D virtual character to perform a target action, the complete pose sequence corresponding to the target action is extracted from the action library. Each frame of the 3D virtual character is rendered in sequence according to the complete pose sequence. In other words, a frame-by-frame image of the 3D virtual character performing the target action is obtained. Playing these frames in sequence can then show the 3D virtual character performing the target action.

[0204] In summary, the above method can be used to complete the pose data for each action in the action library. The hand pose is then completed based on the manually drawn body pose data, resulting in a complete pose for the action. This improves the efficiency of action development and also increases the utilization of open-source 3D pose data without hand poses. Furthermore, the hand poses completed by this method can perfectly match the body pose of the 3D virtual character, improving the effectiveness of pose data completion for 3D virtual characters.

[0205] FIG13 shows a schematic diagram of a device for completing the posture data of a three-dimensional object provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:

[0206] A data module 1001 is configured to obtain incomplete 3D posture data of the 3D object in a preset posture, wherein the incomplete 3D posture data includes first 3D joint point data of some joint points of the 3D object;

[0207] A generating module 1002 is configured to call a Wensheng graph model to generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and a posture description text, wherein the posture description text is used to describe the preset posture;

[0208] A recognition module 1003 is configured to perform joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of the missing joint points of the three-dimensional object except for the partial joint points;

[0209] The completion module 1004 is configured to add the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

[0210] In an optional embodiment, the Wensheng graph model includes a posture control plug-in, and the posture control plug-in is used to constrain the two-dimensional posture image generated by the Wensheng graph model according to the posture skeleton graph;

[0211] The generating module 1002 is configured to map the three-dimensional incomplete posture data to a two-dimensional plane to obtain the posture skeleton graph;

[0212] The generation module 1002 is configured to input the posture description text into the Wensheng graph model, call the posture control plug-in to constrain the image generation process of the Wensheng graph model according to the posture skeleton graph, and obtain the two-dimensional posture image;

[0213] The postures of the part of the joint points in the two-dimensional posture image are consistent with the posture skeleton graph.

[0214] In an optional embodiment, the culture graph model includes a first network and a second network;

[0215] The posture control plug-in includes a first zero convolution layer, a network copy of the first network, and a second zero convolution layer; the network copy is a network obtained by initializing and training using the network structure and network parameters of the first network;

[0216] The generating module 1002 is configured to input the posture description text into the first network to obtain text features;

[0217] The generating module 1002 is configured to input the posture skeleton graph into the first zero convolution layer to obtain a posture convolution result;

[0218] The generating module 1002 is configured to add the posture convolution result to a random noise matrix to obtain a constrained noise matrix; the random noise matrix is ​​a random matrix that conforms to a Gaussian distribution;

[0219] The generating module 1002 is configured to input the constraint noise matrix and the posture skeleton graph into the network replica to obtain a first constraint feature;

[0220] The generating module 1002 is configured to input the first constraint feature into the second zero convolution layer to obtain a second constraint feature;

[0221] The generating module 1002 is configured to add the second constraint feature and the text feature to obtain a text constraint feature;

[0222] The generating module 1002 is configured to input the text constraint features and the posture description text into a second network to obtain the two-dimensional posture image.

[0223] In an optional embodiment, the first network includes at least one encoder; and the second network includes at least one decoder.

[0224] In an optional embodiment, the gesture description text includes positive descriptive words and negative descriptive words;

[0225] The positive description words include the positive requirement text of the two-dimensional posture image, and the positive description words include the preset posture;

[0226] The negative descriptor includes at least one descriptor for describing an image defect, and the negative descriptor is used to guide the text graph model to avoid generating a defective image having the image defect.

[0227] In an optional embodiment, the incomplete three-dimensional posture data is a frame of posture data in a motion sequence of the three-dimensional object, and the motion sequence includes at least two frames of posture data; the apparatus further includes:

[0228] The similarity matching module 1005 is configured to calculate the posture similarity between the first joint point data and the second joint point data after performing joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of the missing joint points of the three-dimensional object except for the partial joint points, wherein the first joint point data includes the second three-dimensional joint point data in the two-dimensional posture image, and the second joint point data includes the second three-dimensional joint point data in the historical posture data; the historical posture data includes at least one frame of posture data in the action sequence that is located before the three-dimensional incomplete posture data; and if the posture similarity between the first joint point data and the second joint point data is less than a similarity threshold, re-execute the following steps until the posture similarity is not less than the similarity threshold:

[0229] Calling the Wensheng graph model to generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and the posture description text;

[0230] Perform joint point recognition on the two-dimensional posture image to obtain the second three-dimensional joint point data.

[0231] In an optional embodiment, the three-dimensional joint point data includes three-dimensional position coordinates and joint rotation angles;

[0232] The similarity matching module 1005 is configured to obtain a completion matrix based on the three-dimensional position coordinates of the first joint point data; and obtain a history matrix based on the three-dimensional position coordinates of the second joint point data;

[0233] The similarity matching module 1005 is configured to calculate the cosine similarity between the completion matrix and the history matrix to obtain a first similarity; and calculate the difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data to obtain a second similarity;

[0234] The similarity matching module 1005 is configured to perform weighted summation of the first similarity and the second similarity to obtain the posture similarity.

[0235] In an optional embodiment, the number of the missing joint points is at least two;

[0236] The similarity matching module 1005 is configured to calculate the difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data, to obtain the joint rotation angle difference of each missing joint point;

[0237] The similarity matching module 1005 is configured to perform weighted summation on at least two joint rotation angle differences according to the weight of each missing joint point in the missing joint points to obtain the second similarity.

[0238] In an optional embodiment, the weight of the parent node in the missing joint point is higher than the weight of the child node; the number of joint points between the parent node and the root node in the three-dimensional object is a first number, and the amount of joint point data between the child node and the root node is a second number, and the first number is less than the second number.

[0239] In an optional embodiment, the incomplete three-dimensional posture data is a frame of posture data in a motion sequence of the three-dimensional object, and the motion sequence includes at least two frames of posture data; the apparatus further includes:

[0240] a smoothing module 1006 for, after adding the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object, smoothing the three-dimensional completed posture data according to adjacent posture data in the action sequence to obtain three-dimensional smoothed posture data;

[0241] The adjacent posture data includes: at least one frame of posture data in the motion sequence that is located before the three-dimensional incomplete posture data, and at least one frame of posture data in the motion sequence that is located after the three-dimensional incomplete posture data;

[0242] Alternatively, the adjacent posture data includes: at least one frame of posture data preceding the three-dimensional incomplete posture data in the motion sequence;

[0243] Alternatively, the adjacent posture data includes: at least one frame of posture data following the three-dimensional incomplete posture data in the action sequence.

[0244] In an optional embodiment, the two-dimensional plane is a two-dimensional imaging plane of a virtual camera, and mapping the three-dimensional incomplete posture data to the two-dimensional plane to obtain the posture skeleton graph includes:

[0245] The generating module 1002 is configured to map the three-dimensional incomplete posture data to a two-dimensional imaging plane of the virtual camera according to virtual camera parameters to obtain a joint point graph, wherein the joint point graph includes two-dimensional joint point coordinates of at least two of the partial joint points;

[0246] The generating module 1002 is configured to connect the two-dimensional joint point coordinates of at least two of the partial joint points in the joint point graph according to the joint point connection relationship of at least two of the partial joint points to obtain the posture skeleton graph.

[0247] In an optional embodiment, the virtual camera parameters include at least one of the coordinates of the virtual camera, the position of the virtual camera relative to the three-dimensional object, and built-in parameters of the virtual camera.

[0248] FIG14 shows a block diagram of a computer device 1400 according to an exemplary embodiment of the present application. The computer device can be implemented as the server in the above-mentioned solution of the present application. The computer device 1400 includes a central processing unit (CPU) 1401, a system memory 1404 including a random access memory (RAM) 1402 and a read-only memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the central processing unit 1401. The computer device 1400 also includes a mass storage device 1406 for storing an operating system 1409, application programs 1410, and other program modules 1411.

[0249] The mass storage device 1406 is connected to the central processing unit 1401 via a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1406 and its associated computer-readable media provide non-volatile storage for the computer device 1400. In other words, the mass storage device 1406 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0250] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, Erasable Programmable Read Only Memory (EPROM), Electronically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state storage technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media are not limited to the above-mentioned ones. The above-mentioned system memory 1404 and mass storage device 1406 can be collectively referred to as memory.

[0251] According to various embodiments of the present application, the computer device 1400 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1400 may be connected to a network 1408 via a network interface unit 1407 connected to the system bus 1405, or the network interface unit 1407 may be used to connect to other types of networks or remote computer systems (not shown).

[0252] The memory also includes at least one computer program, which is stored in the memory. The central processing unit 1401 implements all or part of the steps in the three-dimensional object posture data completion method shown in the above embodiments by executing the at least one program.

[0253] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the three-dimensional object posture data completion method provided by the above-mentioned method embodiments.

[0254] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to implement the three-dimensional object posture data completion method provided by the above-mentioned method embodiments.

[0255] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes to implement the three-dimensional object posture data completion method provided by the above-mentioned method embodiments.

[0256] It is understandable that in the specific implementation methods of this application, the data involved, historical data, and portraits and other data related to user data processing related to user identity or characteristics, when the above embodiments of this application are applied to specific products or technologies, need to obtain user permission or consent, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0257] It should be noted that, unless otherwise expressly defined herein, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field. Unless otherwise expressly stated, all references to "an element, device, component, device, step, etc." are to be interpreted openly as referring to at least one instance of an element, device, component, device, step, etc. Unless expressly stated otherwise, the steps of any method disclosed herein do not have to be performed in the exact order disclosed.

[0258] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0259] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0260] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent switches, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for completing posture data of a three-dimensional object, the method being executed by a computer device, the method comprising: Acquiring incomplete three-dimensional posture data of the three-dimensional object in a preset posture, the incomplete three-dimensional posture data including first three-dimensional joint point data of some joint points of the three-dimensional object; Invoking a Wensheng graph model to generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and a posture description text, wherein the posture description text is used to describe the preset posture; Performing joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of missing joint points of the three-dimensional object except for the partial joint points; The second three-dimensional joint point data is added to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

2. The method according to claim 1, wherein the Wensheng graph model includes a posture control plug-in, and the posture control plug-in is used to constrain the two-dimensional posture image generated by the Wensheng graph model according to the posture skeleton graph; The calling of the text-generated image model to generate the two-dimensional posture image of the preset posture according to the three-dimensional incomplete posture data and the posture description text includes: Mapping the three-dimensional incomplete posture data to a two-dimensional plane to obtain the posture skeleton graph; Inputting the posture description text into the Wensheng graph model, calling the posture control plug-in to constrain the image generation process of the Wensheng graph model according to the posture skeleton graph to obtain the two-dimensional posture image; The postures of the part of the joint points in the two-dimensional posture image are consistent with the posture skeleton graph.

3. The method according to claim 2, wherein the cultural graph model comprises a first network and a second network; The posture control plug-in includes a first zero convolution layer, a network copy of the first network, and a second zero convolution layer; the network copy is a network obtained by initializing and training using the network structure and network parameters of the first network; The step of inputting the posture description text into the Wensheng graph model and calling the posture control plug-in to constrain the image generation process of the Wensheng graph model according to the posture skeleton graph to obtain the two-dimensional posture image of the three-dimensional object includes: Inputting the posture description text into the first network to obtain text features; Inputting the posture skeleton graph into the first zero convolution layer to obtain a posture convolution result; Adding the posture convolution result to a random noise matrix to obtain a constrained noise matrix; the random noise matrix is ​​a random matrix that conforms to a Gaussian distribution; Inputting the constrained noise matrix and the posture skeleton graph into the network copy to obtain a first constrained feature; inputting the first constrained feature into the second zero convolution layer to obtain a second constrained feature; Adding the second constraint feature and the text feature to obtain a text constraint feature; The text constraint features and the posture description text are input into a second network to obtain the two-dimensional posture image.

4. The method according to claim 3, wherein the first network comprises at least one encoder; and the second network comprises at least one decoder.

5. The method according to any one of claims 1 to 4, wherein the posture description text includes positive description words and negative description words; The positive description words include the positive requirement text of the two-dimensional posture image, and the positive description words include the preset posture; The negative descriptor includes at least one descriptor for describing an image defect, and the negative descriptor is used to guide the text graph model to avoid generating a defective image having the image defect.

6. The method according to any one of claims 1 to 5, wherein the incomplete three-dimensional pose data is a frame of pose data in a motion sequence of the three-dimensional object, the motion sequence including at least two frames of pose data; after performing joint point recognition on the two-dimensional pose image to obtain second three-dimensional joint point data of the missing joint points of the three-dimensional object except for the partial joint points, the method further comprises: Calculating posture similarity between first joint point data and second joint point data, wherein the first joint point data includes the second three-dimensional joint point data in the two-dimensional posture image, and the second joint point data includes the second three-dimensional joint point data in historical posture data; the historical posture data includes at least one frame of posture data in the action sequence that is located before the three-dimensional incomplete posture data; When the posture similarity between the first joint point data and the second joint point data is less than the similarity threshold, the following steps are re-executed until the posture similarity is not less than the similarity threshold: Calling the Wensheng graph model to generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and the posture description text; Perform joint point recognition on the two-dimensional posture image to obtain the second three-dimensional joint point data.

7. The method according to claim 6, wherein the three-dimensional joint point data includes three-dimensional position coordinates and joint rotation angles; and the step of calculating the posture similarity between the first joint point data and the second joint point data comprises: Obtaining a completion matrix according to the three-dimensional position coordinates of the first joint point data; and obtaining a history matrix according to the three-dimensional position coordinates of the second joint point data; Calculating the cosine similarity between the completion matrix and the history matrix to obtain a first similarity; and calculating a difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data to obtain a second similarity; The first similarity and the second similarity are weightedly summed to obtain the posture similarity.

8. The method according to claim 7, wherein the number of missing joint points is at least two; The calculating a difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data to obtain a second similarity includes: Calculating the difference between the joint rotation angle of the first joint point data and the joint rotation angle of the second joint point data to obtain the joint rotation angle difference of each missing joint point; According to the weight of each missing joint point in the missing joint points, a weighted sum is performed on at least two joint rotation angle differences to obtain the second similarity.

9. According to the method according to claim 8, the weight of the parent node in the missing joint point is higher than the weight of the child node; the number of joint points between the parent node and the root node in the three-dimensional object is a first number, and the amount of joint point data between the child node and the root node is a second number, and the first number is less than the second number.

10. The method according to any one of claims 1 to 9, wherein the incomplete 3D pose data is a frame of pose data in a motion sequence of the 3D object, the motion sequence comprising at least two frames of pose data; and after adding the second 3D joint point data to the incomplete 3D pose data to complete the incomplete 3D pose data and obtain the completed 3D pose data of the 3D object, the method further comprises: Smoothing the three-dimensional completed posture data according to adjacent posture data in the action sequence to obtain three-dimensional smoothed posture data; The adjacent posture data includes: at least one frame of posture data in the motion sequence that is located before the three-dimensional incomplete posture data, and at least one frame of posture data in the motion sequence that is located after the three-dimensional incomplete posture data; Alternatively, the adjacent posture data includes: at least one frame of posture data preceding the three-dimensional incomplete posture data in the motion sequence; Alternatively, the adjacent posture data includes: at least one frame of posture data following the three-dimensional incomplete posture data in the action sequence.

11. The method according to any one of claims 2 to 10, wherein the two-dimensional plane is a two-dimensional imaging plane of a virtual camera, and mapping the three-dimensional incomplete posture data to the two-dimensional plane to obtain the posture skeleton graph comprises: Mapping the three-dimensional incomplete posture data to a two-dimensional imaging plane of the virtual camera according to virtual camera parameters to obtain a joint point graph, wherein the joint point graph includes two-dimensional joint point coordinates of at least two of the partial joint points; According to the joint point connection relationship of at least two of the partial joint points, the two-dimensional joint point coordinates of at least two of the partial joint points in the joint point graph are connected to obtain the posture skeleton graph.

12. The method according to claim 11, wherein the virtual camera parameters include: At least one of the coordinates of the virtual camera, the position of the virtual camera relative to the three-dimensional object, and the built-in parameters of the virtual camera.

13. A device for completing posture data of a three-dimensional object, the device being deployed on a computer device, the device comprising: a data module, configured to obtain three-dimensional incomplete posture data of the three-dimensional object in a preset posture, wherein the three-dimensional incomplete posture data includes first three-dimensional joint point data of some joint points of the three-dimensional object; a generating module, configured to call a Wensheng graph model to generate a two-dimensional posture image of the three-dimensional object in the preset posture according to the three-dimensional incomplete posture data and a posture description text, wherein the posture description text is used to describe the preset posture; a recognition module, configured to perform joint point recognition on the two-dimensional posture image to obtain second three-dimensional joint point data of missing joint points of the three-dimensional object except for the partial joint points; The completion module is used to add the second three-dimensional joint point data to the three-dimensional incomplete posture data to complete the three-dimensional incomplete posture data and obtain the three-dimensional completed posture data of the three-dimensional object.

14. A computer device, comprising: A processor and a memory, wherein at least one computer program is stored in the memory, and at least one computer program is loaded and executed by the processor to implement the method for completing the posture data of a three-dimensional object as described in any one of claims 1 to 12.

15. A computer storage medium, wherein the computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the method for completing the posture data of a three-dimensional object as described in any one of claims 1 to 12.

16. A computer program product, comprising a computer program stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes the three-dimensional object posture data completion method as described in any one of claims 1 to 12.