A cross-domain human body unified modeling method, device, equipment and storage medium

By constructing a relative pose similarity space and filtering target pose update anchor prompt sets, the problem of cross-task deep learning model is solved, and cross-task unified modeling is realized, and the adaptability and flexibility of the model is improved.

CN120107495BActive Publication Date: 2025-08-15PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510595696.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing technology is difficult to implement deep learning model reasoning across tasks, and different models are required for different tasks.

Method used

By constructing a relative posture similarity space, filter out the target posture update anchor prompt set, and combine soft prompts to achieve unified cross-domain human body modeling.

Benefits of technology

It realizes unified modeling under different tasks, improves the generalization ability and task adaptability of the model, and reduces the need for model replacement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107495B_ABST
    Figure CN120107495B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of human body modeling, and in particular to a cross-domain unified human body modeling method, apparatus, equipment and storage medium. The present invention uses a relative posture similarity space to represent the relative similarity between any two skeleton postures of the human body, and then extracts the relative similarity between the unsampled skeleton posture and the anchor prompt from the relative posture similarity space. Based on the relative similarity, the target posture is screened out from the unsampled skeleton posture, and the anchor prompt set is updated with the target posture to obtain a final set of anchor prompts. Finally, the anchor prompts and soft prompts in the final set of anchor prompts cooperate with each other to complete the reasoning task of the model. Since the soft prompts are based on the parameters generated when the model is trained for each reasoning task, when processing different reasoning tasks, it is only necessary to replace the soft prompts corresponding to the reasoning task to cooperate with the anchor prompts to complete the reasoning task without replacing the model. Therefore, the present invention can achieve unified modeling across tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human body modeling, and in particular to a cross-domain human body unified modeling method, apparatus, device and storage medium. Background Art

[0002] Human body modeling involves using deep learning to accomplish a variety of human-centric tasks, including human pose estimation, motion prediction, action recognition, mesh reconstruction, and joint completion. Existing technologies require different deep learning models for different tasks to perform model inference. For example, human pose estimation requires a deep learning model that matches human pose estimation to infer human poses. Similarly, motion prediction requires a deep learning model that matches motion prediction to infer human motion trends.

[0003] In summary, it is difficult to achieve cross-task reasoning using the same model using existing technologies.

[0004] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a cross-domain human body unified modeling method, device, equipment and storage medium, which solves the problem that the existing technology is difficult to achieve cross-task reasoning using the same model.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a cross-domain human body unified modeling method, which includes:

[0008] Acquiring a human skeleton posture, determining a relative similarity between any two of the skeleton postures, and constructing a relative posture similarity space based on the relative similarity between any two of the skeleton postures;

[0009] Setting an anchor hint set, and initializing the anchor hint set and defining an unsampled pose set using the skeleton pose;

[0010] Extracting a relative similarity between an unsampled skeleton pose and an anchor hint from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set and the anchor hint is a skeleton pose in the anchor hint set;

[0011] Based on the relative similarity between the unsampled skeleton pose and the anchor prompt, a target pose is screened from the unsampled pose set, and the anchor prompt set is updated with the target pose to obtain a final set of anchor prompts. The final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.

[0012] In one implementation, determining the relative similarity of any two skeleton poses includes:

[0013] Obtaining joint positions of the human body when it is in any two of the skeleton postures;

[0014] The relative similarity between any two skeleton postures is determined based on the joint positions corresponding to the any two skeleton postures.

[0015] In one implementation, a relative posture similarity space is constructed based on the similarity of any two skeleton postures, including:

[0016] Determining a reference pose and a non-reference pose included in any two of the skeleton poses;

[0017] Setting a blank preset space, and determining a position of the non-reference posture in the preset space based on a relative similarity between the non-reference posture and the reference posture;

[0018] The relative similarities corresponding to any two of the skeleton postures including the non-reference posture are filled in the position to obtain the relative posture similarity space.

[0019] In one implementation, the skeleton pose in the unsampled pose set is the skeleton pose after removing the anchor hint set.

[0020] In one implementation, selecting a target pose from the set of unsampled poses based on a relative similarity between the unsampled skeleton pose and the anchor cue includes:

[0021] Determining a maximum relative similarity between each of the unsampled skeleton poses and each of the anchor hints based on the relative similarities between the unsampled skeleton poses and the anchor hints;

[0022] The smallest maximum relative similarity is screened out from the maximum relative similarities corresponding to each of the unsampled skeleton postures, and the unsampled skeleton posture corresponding to the smallest maximum relative similarity is used as the target posture.

[0023] In one implementation, updating the anchor prompt set with the target pose to obtain a final set of anchor prompts includes:

[0024] The target pose is removed from the unsampled pose set, and the target pose is added to the anchor prompt set to update the anchor prompt set until the number of anchor prompts in the anchor prompt set reaches a set number, thereby obtaining a final set of anchor prompts.

[0025] In one implementation, the modality of the skeleton posture includes a two-dimensional modality, a three-dimensional modality and a mesh modality; the two-dimensional modality contains the serial number of the skeleton posture and the two-dimensional coordinates of the joints of the skeleton posture; the three-dimensional modality contains the serial number of the skeleton posture and the three-dimensional coordinates of the joints of the skeleton posture; the mesh modality includes the joint rotation information of the skeleton posture and the shape of the skeleton posture.

[0026] In a second aspect, an embodiment of the present invention further provides a cross-domain human body unified modeling device, wherein the device includes the following components:

[0027] A relative similarity calculation module is used to obtain the skeleton posture of the human body, determine the relative similarity between any two skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two skeleton postures;

[0028] A set setting module, configured to set an anchor prompt set, initialize the anchor prompt set using the skeleton pose, and define an unsampled pose set;

[0029] an information extraction module for extracting relative similarities between unsampled skeleton poses and anchor hints from the relative pose similarity space, wherein the unsampled skeleton poses are skeleton poses in the unsampled pose set and the anchor hints are skeleton poses in the anchor hint set;

[0030] An updating module is configured to filter out a target pose from the unsampled pose set based on the relative similarity between the unsampled skeleton pose and the anchor prompt, and to update the anchor prompt set with the target pose to obtain a final set of anchor prompts. The final set of anchor prompts is used in conjunction with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.

[0031] In the third aspect, an embodiment of the present invention also provides a terminal device, wherein the terminal device includes a memory, a processor, and a cross-domain human body unified modeling program stored in the memory and runnable on the processor, and when the processor executes the cross-domain human body unified modeling program, the steps of the above-mentioned cross-domain human body unified modeling method are implemented.

[0032] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a cross-domain human body unified modeling program is stored. When the cross-domain human body unified modeling program is executed by a processor, the steps of the above-mentioned cross-domain human body unified modeling method are implemented.

[0033] Beneficial effects: The present invention uses a relative posture similarity space to represent the relative similarity between any two skeleton postures of the human body, and then extracts the relative similarity between the unsampled skeleton posture and the anchor prompt from the relative posture similarity space. Based on the relative similarity, the target posture is screened out from the unsampled skeleton posture, and the anchor prompt set is updated with the target posture to obtain a final set of anchor prompts. Finally, the anchor prompts and soft prompts in the final set of anchor prompts are used to cooperate with each other to complete the reasoning task of the model. Since the soft prompts are based on the parameters generated when the model is trained for each reasoning task, when processing different reasoning tasks, it is only necessary to replace the soft prompts corresponding to the reasoning task to cooperate with the anchor prompts to complete the reasoning task without replacing the model. Therefore, the present invention can achieve unified modeling across tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is the overall flow chart of the present invention;

[0035] Figure 2 is a schematic diagram of a skeleton posture mode in an embodiment of the present invention;

[0036] Figure 3 A schematic diagram of context learning in an embodiment of the present invention;

[0037] Figure 4 A schematic diagram of query and prompt in an embodiment of the present invention;

[0038] Figure 5 A structural diagram of the cross-domain unified human body modeling device provided by the present invention;

[0039] Figure 6 This is a block diagram of the internal structure of a terminal device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The following is a clear and complete description of the technical solutions of the present invention in conjunction with the embodiments and the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0041] Research has shown that human body modeling involves using deep learning techniques to accomplish a variety of human-centric tasks, including human pose estimation, motion prediction, action recognition, mesh reconstruction, and joint completion. Existing technologies require different deep learning models for different tasks to perform model inference. For example, human pose estimation requires a deep learning model that matches human pose estimation to infer human poses. Similarly, motion prediction requires a deep learning model that matches motion prediction to infer human motion trends.

[0042] To solve the above technical problems, the present invention provides a cross-domain human body unified modeling method, device, equipment and storage medium, which solves the problem that the existing technology is difficult to achieve cross-task reasoning using the same model.

[0043] The cross-domain human body unified modeling method of this embodiment can be applied to a terminal device, which can be a terminal product with data processing function, such as a computer. Figure 1 As shown in , the cross-domain human body unified modeling method specifically includes the following steps:

[0044] S100, obtaining a human skeleton posture, determining a relative similarity between any two skeleton postures, and constructing a relative posture similarity space based on the relative similarity between any two skeleton postures;

[0045] S200, setting an anchor prompt set, and initializing the anchor prompt set and defining an unsampled pose set using the skeleton pose;

[0046] S300, extracting relative similarities between unsampled skeleton poses and anchor hints from the relative pose similarity space, wherein the unsampled skeleton poses are skeleton poses in the unsampled pose set, and the anchor hints are skeleton poses in the anchor hint set;

[0047] S400 , based on the relative similarity between the unsampled skeleton pose and the anchor hint, a target pose is selected from the unsampled pose set, and the anchor hint set is updated with the target pose to obtain a final anchor hint set.

[0048] Among them, the final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. Soft prompts are learning parameters generated when training the model based on reasoning tasks.

[0049] The model is as follows Figure 3 The model after context learning shown in , that is, the model after training, as shown in Figure 3 As shown in the figure, context learning is to input the skeleton posture of the human body into the model, and the model outputs the next skeleton posture through context learning.

[0050] Example 1: In this example, Figure 2 As shown, the modes of skeleton posture include two-dimensional modes (two-dimensional modes are Figure 2 In the 2D representation, the key points are the joints), the three-dimensional mode (the three-dimensional mode is Figure 2 In the 3D representation of key points) and meshes, the two-dimensional modality is to represent the skeleton pose with two-dimensional coordinates, the three-dimensional modality is to represent the skeleton pose with three-dimensional coordinates, and the mesh modality is to represent the skeleton pose with rotation information and shape.

[0051] When the mode is 2D or 3D, the sequence of setting the skeleton pose is: , among which 1,2, are the frame numbers. is the serial number of the largest frame, C is the element (the element is also the coordinate dimension), is the total number of joints, represents a data set of real numbers, Indicates the The skeleton pose of the frame, Represents a set consisting of the frame number, the total number of elements, and the total number of joints. When the skeleton pose is 2D, C=2, which means that the coordinates of each joint on the skeleton are represented by 2D coordinates (x, y); when the skeleton pose is 3D, C=3, which means that the coordinates of each joint on the skeleton are represented by 3D coordinates (x, y, z).

[0052] When using meshes to represent skeleton poses, mesh-based human motion models the human skeleton as a 3D mesh consisting of vertices and faces. This captures the skeleton pose and joint configuration of the human body as well as the finer geometric details of the human body surface. In other words, a set of compact parameters is used to describe the shape and pose of the human body. In mesh-based human motion, the 3D mesh is composed of pose parameters. Drive, posture parameters describe the relative rotation of the joints; shape parameters It captures individual differences in height, weight, muscle mass, etc. The mesh consists of vertices Definition, representing the position of the human body surface in three-dimensional space, with vertices connected by triangles to form a complete structure. By applying linear blend skinning, the mesh is deformed according to the joint transformation, ensuring smooth movement during joint movement and making the joints have anatomically consistent movement.

[0053] Example 2, based on Example 1, in this example, the relative similarity of any two skeleton postures in step S100 is calculated, including: obtaining the joint positions of the human body when it is in any two of the skeleton postures; determining the relative similarity of any two of the skeleton postures based on the joint positions corresponding to the any two skeleton postures :

[0054] ;

[0055] Represents similarity, N is the total number of joints, C is the coordinate dimension, is the skeleton pose among any two skeleton poses , is the skeleton pose among any two skeleton poses , for The The first joint Elements ( The first joint The first element joint positions), when When it is 2, that is, two-dimensional, the The element is the horizontal coordinate or vertical coordinate ; In three dimensions, The element is the horizontal coordinate or vertical coordinate or vertical coordinate . for The The first joint elements.

[0056] The minus sign in the and When the difference is large, it means the similarity is low, that is, when the similarity is low, it means and The two postures are quite different. The minus sign is also used to indicate that and When the difference is small, it means the similarity is high, that is, when the similarity is high, it means and The two poses are relatively small in difference. In summary, the negative sign is used to make two poses with large differences correspond to low similarity, and two poses with small differences correspond to high similarity.

[0057] After calculating the relative similarity of any two skeleton postures, this embodiment uses the relative similarity of any two skeleton postures to construct a relative posture similarity space, including the following specific steps: determining the reference posture and non-reference posture contained in any two of the skeleton postures; setting a blank preset space, and determining the position of the non-reference posture in the preset space based on the relative similarity between the non-reference posture and the reference posture; filling the relative similarity corresponding to any two of the skeleton postures containing the non-reference posture in the position to obtain the relative posture similarity space.

[0058] The reference pose is the T-shaped posture exhibited by a naturally stretched human body. The non-reference pose is the pose after removing the reference pose from all skeletal poses. The relative similarity between a non-reference pose and the reference pose is calculated using the aforementioned formula (denoted as SIM1). The size of SIM1 determines the position of the non-reference pose in the blank preset space. The relative similarity between this non-reference pose and any other non-reference pose is then calculated (denoted as SIM2), and SIM2 is placed in that position. This same process is repeated for all non-reference poses to obtain the relative pose similarity space, which records the relative similarity between any two non-reference poses.

[0059] For example, if the blank preset space is a blank matrix, then SIM1 determines the number of rows in the space matrix corresponding to the relative similarity of one of the non-reference postures, and SIM2 is placed on the row corresponding to SIM1. The size of SIM2 determines the number of columns it is in.

[0060] That is, each pose (i.e., a non-reference pose) is represented based on its relative similarity to a canonical pose (i.e., a reference pose) and its relationship to other poses in the dataset. The canonical pose serves as a reference point in this space, and the positions of other poses are determined by their similarity or difference to the canonical pose and their relative ranking in the pose distribution. This relative pose similarity space helps identify diverse and representative poses in a dataset (e.g., anchor cues).

[0061] Example 3, based on Example 1 or Example 2, in this example, the anchor prompt set after initialization Contains reference poses and non-reference poses, anchor prompt set The reference poses and non-reference poses in are called anchor cues. Remove anchor hint sets for all skeleton poses The skeleton pose afterward.

[0062] The step S400 of this embodiment of the present invention includes the following specific steps: determining the maximum relative similarity between each unsampled skeleton pose and each anchor prompt based on the relative similarity between the unsampled skeleton pose and the anchor prompt; selecting the minimum maximum relative similarity from the maximum relative similarities corresponding to each unsampled skeleton pose, and using the unsampled skeleton pose corresponding to the minimum maximum relative similarity as the target pose; .

[0063] That is, each cycle never samples the posture set Unsampled skeleton poses in Filter out target poses as anchor prompt sets The next anchor prompt in the . The specific process is as follows:

[0064] Corresponding to the unsampled pose set Each unsampled skeleton pose in , that is, calculation ( That is unsampled skeleton poses) and a set of anchor hints The similarity of all anchor prompts in and select the largest one:

[0065] ;

[0066] Where, Anchor prompt collection Anchor hint in .

[0067] ;

[0068] In the formula Represents the solution satisfaction of , that is, the unsampled pose with the least similarity to the existing anchor prompt is used as the next sample to ensure that the newly selected pose is the most different from the existing anchor prompt set, thereby increasing the diversity of the sampled data.

[0069] For example, the unsampled pose set Including unsampled skeleton pose 1, unsampled skeleton pose 2, unsampled skeleton pose 3, anchor prompt set Including anchor prompt one and anchor prompt two. Calculate the similarity sim11 between the unsampled skeleton pose one and the anchor prompt one, calculate the similarity sim12 between the unsampled skeleton pose one and the anchor prompt two, and select the maximum similarity from the similarity sim11 and the similarity sim12, and record it as Maxsim1; calculate the similarity sim21 between the unsampled skeleton pose two and the anchor prompt one, calculate the similarity sim22 between the unsampled skeleton pose two and the anchor prompt two, and select the maximum similarity from the similarity sim21 and the similarity sim22, and record it as Maxsim2; calculate the similarity sim31 between the unsampled skeleton pose three and the anchor prompt one, calculate the similarity sim32 between the unsampled skeleton pose three and the anchor prompt two, and select the maximum similarity from the similarity sim31 and the similarity sim32, and record it as Maxsim3. Then select the minimum similarity from Maxsim1, Maxsim2, and Maxsim3. If the minimum similarity is Maxsim3, then the unsampled skeleton pose three is the target pose of this cycle. .

[0070] Get the target posture in each cycle Afterwards, the target pose is removed from the unsampled pose set and added to the anchor prompt set to update the anchor prompt set until the number of anchor prompts in the anchor prompt set reaches a set number, thereby obtaining a final set of anchor prompts.

[0071] Example 4, based on Example 1 or Example 2 or Example 3, as Figure 4 As shown, this embodiment matches each anchor prompt in the final set of anchor prompts with a corresponding soft prompt, so as to guide the reasoning process of the model by combining with the corresponding soft prompt. This enables the model to not only identify key patterns similar to the query sample, but also to make appropriate adjustments between different tasks based on the guidance of the soft prompt. That is, the representative anchor prompts in the anchor prompt set are used to dynamically adapt to diverse task requirements, thereby significantly improving the generalization ability and task adaptability of the model. By selecting the most similar anchor prompt, the model can model the query sample (the query sample is the human skeleton posture as the sample) more accurately. At the same time, the introduction of soft prompts provides the model with more contextual information, helping it to better reason and make decisions in complex tasks. This embodiment is based on a joint strategy of anchor prompts and soft prompts, which not only enhances the expressive power of the model, but also improves its flexibility and adaptability when facing different tasks.

[0072] This embodiment also provides a cross-domain human body unified modeling device, such as Figure 5 As shown, the device includes the following components:

[0073] The relative similarity calculation module 01 is used to obtain the skeleton posture of the human body, determine the relative similarity between any two skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two skeleton postures;

[0074] A set setting module 02 is used to set an anchor prompt set, and initialize the anchor prompt set using the skeleton pose and define an unsampled pose set;

[0075] An information extraction module 03 is configured to extract relative similarities between unsampled skeleton poses and anchor hints from the relative pose similarity space, wherein the unsampled skeleton poses are skeleton poses in the unsampled pose set, and the anchor hints are skeleton poses in the anchor hint set;

[0076] An updating module 04 is configured to filter out a target pose from the unsampled pose set based on the relative similarity between the unsampled skeleton pose and the anchor prompt, and to update the anchor prompt set with the target pose to obtain a final set of anchor prompts. The final set of anchor prompts is used in conjunction with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when the model is trained based on the reasoning task.

[0077] Based on the above embodiment, the present invention further provides a terminal device, whose principle block diagram can be shown as follows: Figure 6 As shown. The terminal device includes a processor, a memory, a network interface, and a display screen connected via a system bus. The processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a cross-domain unified human body modeling method is implemented. The display screen of the terminal device can be a liquid crystal display or an electronic ink display.

[0078] Those skilled in the art will understand that Figure 6 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0079] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a cross-domain human body unified modeling program stored in the memory and executable on the processor. When the processor executes the cross-domain human body unified modeling program, the following operating instructions are implemented:

[0080] Acquiring a human skeleton posture, determining a relative similarity between any two of the skeleton postures, and constructing a relative posture similarity space based on the relative similarity between any two of the skeleton postures;

[0081] Setting an anchor hint set, and initializing the anchor hint set and defining an unsampled pose set using the skeleton pose;

[0082] Extracting a relative similarity between an unsampled skeleton pose and an anchor hint from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set and the anchor hint is a skeleton pose in the anchor hint set;

[0083] Based on the relative similarity between the unsampled skeleton pose and the anchor prompt, a target pose is screened from the unsampled pose set, and the anchor prompt set is updated with the target pose to obtain a final set of anchor prompts. The final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.

[0084] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A cross-domain human body unified modeling method, characterized in that: include: Acquiring a human skeleton posture, determining a relative similarity between any two of the skeleton postures, and constructing a relative posture similarity space based on the relative similarity between any two of the skeleton postures; Setting an anchor hint set, and initializing the anchor hint set and defining an unsampled pose set using the skeleton pose; Extracting a relative similarity between an unsampled skeleton pose and an anchor hint from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set and the anchor hint is a skeleton pose in the anchor hint set; Filtering a target pose from the set of unsampled poses based on the relative similarity between the unsampled skeleton pose and the anchor hints, and updating the set of anchor hints with the target pose to obtain a final set of anchor hints, wherein the final set of anchor hints is used in conjunction with soft hints to complete the model's reasoning task on the human body, wherein the soft hints are learning parameters generated when training the model based on the reasoning task; Determining the relative similarity of any two of the skeleton poses includes: Obtaining joint positions of the human body when it is in any two of the skeleton postures; Determining the relative similarity of any two skeleton postures based on the joint positions corresponding to the any two skeleton postures; Filtering a target pose from the set of unsampled poses based on a relative similarity between the unsampled skeleton pose and the anchor hint, comprising: Determining a maximum relative similarity between each of the unsampled skeleton poses and each of the anchor hints based on the relative similarities between the unsampled skeleton poses and the anchor hints; The smallest maximum relative similarity is screened out from the maximum relative similarities corresponding to each of the unsampled skeleton postures, and the unsampled skeleton posture corresponding to the smallest maximum relative similarity is used as the target posture.

2. The cross-domain human body unified modeling method according to claim 1, characterized in that: Based on the relative similarity of any two skeleton postures, a relative posture similarity space is constructed, including: Determining a reference pose and a non-reference pose included in any two of the skeleton poses; Setting a blank preset space, and determining a position of the non-reference posture in the preset space based on a relative similarity between the non-reference posture and the reference posture; The relative similarities corresponding to any two of the skeleton postures including the non-reference posture are filled in the position to obtain the relative posture similarity space.

3. The cross-domain human body unified modeling method according to claim 1, characterized in that: The skeleton poses in the unsampled pose set are the skeleton poses after removing the anchor hint set.

4. The cross-domain human body unified modeling method according to claim 1, characterized in that: Updating the anchor prompt set with the target pose to obtain a final set of anchor prompts, including: The target pose is removed from the unsampled pose set, and the target pose is added to the anchor prompt set to update the anchor prompt set until the number of anchor prompts in the anchor prompt set reaches a set number, thereby obtaining a final set of anchor prompts.

5. The cross-domain unified human body modeling method according to any one of claims 1 to 4, characterized in that: The modality of the skeleton posture includes a two-dimensional modality, a three-dimensional modality and a mesh modality; the two-dimensional modality contains the serial number of the skeleton posture and the two-dimensional coordinates of the joints of the skeleton posture; the three-dimensional modality contains the serial number of the skeleton posture and the three-dimensional coordinates of the joints of the skeleton posture; the mesh modality includes the joint rotation information of the skeleton posture and the shape of the skeleton posture.

6. A cross-domain human body unified modeling device, characterized in that: The device comprises the following components: A relative similarity calculation module is used to obtain the skeleton posture of the human body, determine the relative similarity between any two skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two skeleton postures; A set setting module, configured to set an anchor prompt set, initialize the anchor prompt set using the skeleton pose, and define an unsampled pose set; an information extraction module for extracting relative similarities between unsampled skeleton poses and anchor hints from the relative pose similarity space, wherein the unsampled skeleton poses are skeleton poses in the unsampled pose set and the anchor hints are skeleton poses in the anchor hint set; an updating module for filtering a target pose from the set of unsampled poses based on a relative similarity between the unsampled skeleton pose and the anchor hints, and updating the set of anchor hints with the target pose to obtain a final set of anchor hints, wherein the final set of anchor hints is used in conjunction with soft hints to complete the model's reasoning task on the human body, wherein the soft hints are learning parameters generated when training the model based on the reasoning task; Determining the relative similarity of any two of the skeleton poses includes: Obtaining joint positions of the human body when it is in any two of the skeleton postures; Determining the relative similarity of any two skeleton postures based on the joint positions corresponding to the any two skeleton postures; Filtering a target pose from the set of unsampled poses based on a relative similarity between the unsampled skeleton pose and the anchor hint, comprising: Determining a maximum relative similarity between each of the unsampled skeleton poses and each of the anchor hints based on the relative similarities between the unsampled skeleton poses and the anchor hints; The smallest maximum relative similarity is screened out from the maximum relative similarities corresponding to each of the unsampled skeleton postures, and the unsampled skeleton posture corresponding to the smallest maximum relative similarity is used as the target posture.

7. A terminal device, characterized in that: The terminal device includes a memory, a processor, and a cross-domain human body unified modeling program stored in the memory and runnable on the processor. When the processor executes the cross-domain human body unified modeling program, the steps of the cross-domain human body unified modeling method as described in any one of claims 1-5 are implemented.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a cross-domain human body unified modeling program. When the cross-domain human body unified modeling program is executed by the processor, the steps of the cross-domain human body unified modeling method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method, device, equipment and medium

    CN117711066A

  • Cross-domain attitude estimation method based on skeleton graph structure constraint and application

    CN118865486A