Cross-domain human body unified modeling method and device, equipment and storage medium
By constructing relative pose similarity space and updating anchor prompt sets, the problem that models in the existing technology are difficult to achieve cross-task reasoning, and unified cross-task modeling is realized, and the generalization ability and task adaptability of the model are improved.
Patent Information
- Application Number
- CN202510595696.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The prior art uses the same model to achieve cross-task inference and cannot effectively share and utilize the knowledge of deep learning models among different tasks.
By obtaining the human skeleton pose, building a relative pose similarity space, setting up an anchor prompt set, and filtering the target pose through the relative similarity between the unsampled pose and the anchor prompt, updating the anchor prompt set, and finally using the final set of anchor prompts combined with the soft prompt to complete the cross-task reasoning of the model.
It realizes unified modeling across tasks, and can share and utilize model knowledge among different inference tasks, improving the generalization ability and task adaptability of the model.
Smart Images

Figure CN120107495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human body modeling, and in particular to a cross-domain human body unified modeling method, device, equipment and storage medium. Background Art
[0002] Human body modeling refers to the use of deep learning technology to complete a variety of human-centric tasks, including human posture estimation, motion prediction, action recognition, mesh reconstruction and joint completion. Existing technologies require different deep learning models for different tasks to complete model reasoning. For example, for human posture estimation tasks, a deep learning model that matches human posture estimation is required to reason about the human body's posture. Similarly, for motion prediction, a deep learning model that matches motion prediction is also required to reason about the human body's motion trend.
[0003] In summary, it is difficult to achieve cross-task reasoning using the same model using existing technologies.
[0004] Therefore, the prior art still needs to be improved and enhanced. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a cross-domain human body unified modeling method, device, equipment and storage medium, which solves the problem that the same model is difficult to achieve cross-task reasoning in the prior art.
[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a cross-domain human unified modeling method, which includes: Acquire a human skeleton posture, determine the relative similarity between any two of the skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two of the skeleton postures; Setting an anchor hint set, and initializing the anchor hint set and defining an unsampled pose set using the skeleton pose; Extracting a relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; According to the relative similarity between the unsampled skeleton posture and the anchor prompt, the target posture is screened out from the unsampled posture set, and the anchor prompt set is updated with the target posture to obtain a final set of anchor prompts. The final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.
[0007] In one implementation, determining the relative similarity of any two of the skeleton postures includes: Obtaining joint positions of the human body when it is in any two of the skeleton postures; The relative similarity between any two skeleton postures is determined based on the joint positions corresponding to any two skeleton postures.
[0008] In one implementation, based on the similarity of any two skeleton postures, a relative posture similarity space is constructed, including: Determine a reference pose and a non-reference pose included in any two of the skeleton poses; Setting a blank preset space, and determining a position of the non-reference gesture in the preset space according to a relative similarity between the non-reference gesture and the reference gesture; The relative similarities corresponding to any two of the skeleton postures including the non-reference posture are filled in the position to obtain the relative posture similarity space.
[0009] In one implementation, the skeleton pose in the unsampled pose set is the skeleton pose after the anchor hint set is removed.
[0010] In one implementation, selecting a target pose from the unsampled pose set based on a relative similarity between the unsampled skeleton pose and the anchor hint includes: Determining a maximum relative similarity between each of the unsampled skeleton poses and each of the anchor cues based on the relative similarities between the unsampled skeleton poses and the anchor cues; The smallest maximum relative similarity is screened out from the maximum relative similarities corresponding to each of the unsampled skeleton postures, and the unsampled skeleton posture corresponding to the smallest maximum relative similarity is used as the target posture.
[0011] In one implementation, updating the anchor prompt set with the target pose to obtain a final set of anchor prompts includes: The target posture is removed from the unsampled posture set, and the target posture is added to the anchor prompt set to update the anchor prompt set until the number of anchor prompts in the anchor prompt set reaches a set number, thereby obtaining a final set of anchor prompts.
[0012] In one implementation, the modality of the skeleton posture includes a two-dimensional modality, a three-dimensional modality and a mesh modality; the two-dimensional modality contains the serial number of the skeleton posture and the two-dimensional coordinates of the joints of the skeleton posture; the three-dimensional modality contains the serial number of the skeleton posture and the three-dimensional coordinates of the joints of the skeleton posture; the mesh modality includes the joint rotation information of the skeleton posture and the shape of the skeleton posture.
[0013] In a second aspect, an embodiment of the present invention further provides a cross-domain human unified modeling device, wherein the device includes the following components: A relative similarity calculation module is used to obtain the skeleton posture of the human body, determine the relative similarity between any two skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two skeleton postures; A set setting module, used for setting an anchor prompt set, initializing the anchor prompt set using the skeleton posture and defining an unsampled posture set; an information extraction module, configured to extract a relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; An updating module is used to filter out a target posture from the unsampled posture set based on the relative similarity between the unsampled skeleton posture and the anchor prompt, and use the target posture to update the anchor prompt set to obtain a final set of anchor prompts, wherein the final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body, and the soft prompts are learning parameters generated when the model is trained based on the reasoning task.
[0014] In the third aspect, an embodiment of the present invention further provides a terminal device, wherein the terminal device includes a memory, a processor, and a cross-domain human body unified modeling program stored in the memory and executable on the processor, and when the processor executes the cross-domain human body unified modeling program, the steps of the above-mentioned cross-domain human body unified modeling method are implemented.
[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a cross-domain human body unified modeling program is stored. When the cross-domain human body unified modeling program is executed by a processor, the steps of the above-mentioned cross-domain human body unified modeling method are implemented.
[0016] Beneficial effect: The present invention uses a relative posture similarity space to represent the relative similarity between any two skeleton postures of the human body, and then extracts the relative similarity between the unsampled skeleton posture and the anchor prompt from the relative posture similarity space. Based on the relative similarity, the target posture is screened out from the unsampled skeleton posture, and the anchor prompt set is updated with the target posture to obtain a final set of anchor prompts. Finally, the anchor prompts and soft prompts in the final set of anchor prompts are used to cooperate with each other to complete the reasoning task of the model. Since the soft prompts are based on the parameters generated when the model is trained for each reasoning task, when processing different reasoning tasks, it is only necessary to replace the soft prompts corresponding to the reasoning task to cooperate with the anchor prompts to complete the reasoning task without changing the model. Therefore, the present invention can achieve unified modeling across tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall flow chart of the present invention; Figure 2 is a schematic diagram of a skeleton posture mode in an embodiment of the present invention; Figure 3 A schematic diagram of context learning in an embodiment of the present invention; Figure 4 A schematic diagram of a query and prompt in an embodiment of the present invention; Figure 5 A structural diagram of the cross-domain human body unified modeling device provided by the present invention; Figure 6 This is a block diagram of the internal structure principle of the terminal device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following is a clear and complete description of the technical solution of the present invention in combination with the embodiments and the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] Research has found that human body modeling refers to the use of deep learning technology to complete a variety of human-centric tasks, including human posture estimation, motion prediction, action recognition, mesh reconstruction and joint completion. Existing technologies require different deep learning models for different tasks to complete model reasoning. For example, for human posture estimation tasks, a deep learning model that matches human posture estimation is required to reason about the human body's posture. Similarly, for motion prediction, a deep learning model that matches motion prediction is also required to reason about the human body's motion trend.
[0020] In order to solve the above technical problems, the present invention provides a cross-domain human body unified modeling method, device, equipment and storage medium, which solves the problem that the same model is difficult to achieve cross-task reasoning in the prior art.
[0021] The cross-domain human unified modeling method of this embodiment can be applied to a terminal device, which can be a terminal product with a data processing function, such as a computer. Figure 1 As shown in , the cross-domain human unified modeling method specifically includes the following steps: S100, acquiring a skeleton posture of a human body, determining a relative similarity between any two skeleton postures, and constructing a relative posture similarity space based on the relative similarity between any two skeleton postures; S200, setting an anchor prompt set, and using the skeleton posture to initialize the anchor prompt set and define an unsampled posture set; S300, extracting the relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; S400 , based on the relative similarity between the unsampled skeleton pose and the anchor prompt, a target pose is selected from the unsampled pose set, and the anchor prompt set is updated with the target pose to obtain a final anchor prompt set.
[0022] Among them, the final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning tasks for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.
[0023] The model is as follows Figure 3 The model after context learning, that is, the model after training, is shown in Figure 3 As shown in the figure, context learning is to input the skeleton posture of the human body into the model, and the model outputs the next skeleton posture through context learning.
[0024] Embodiment 1, in this embodiment, as Figure 2 As shown, the modes of skeleton posture include two-dimensional modes (two-dimensional modes are Figure 2 In the 2D representation, the key points are the joints), the 3D mode (the 3D mode is Figure 2 In the 3D representation of key points) and meshes, the two-dimensional modality is to represent the skeleton posture with two-dimensional coordinates, the three-dimensional modality is to represent the skeleton posture with three-dimensional coordinates, and the mesh modality is to represent the skeleton posture with rotation information and shape.
[0025] When the mode is 2D mode or 3D mode, the sequence of setting the skeleton posture is , where 1,2, are the frame numbers. is the serial number of the largest frame, C is the element (the element is also the coordinate dimension), is the total number of joints, represents a data set of real numbers, Indicates The skeleton pose of the frame, Represents a set consisting of the frame number, the total number of elements, and the total number of joints. When the skeleton posture is a two-dimensional skeleton, C=2, that is, the coordinates of each joint on the skeleton are represented by two-dimensional coordinates (x, y); when the skeleton posture is a three-dimensional skeleton, C=3, that is, the coordinates of each joint on the skeleton are represented by three-dimensional coordinates (x, y, z).
[0026] When a mesh is used to represent the skeleton pose, mesh-based human motion models the human skeleton as a 3D mesh consisting of vertices and faces. This captures the skeleton pose and joint configuration of the human body as well as the finer geometric details of the human body surface. That is, a set of compact parameters is used to describe the shape and pose of the human body. In mesh-based human motion, the 3D mesh consists of pose parameters. Drive, posture parameters describe the relative rotation of the joints; shape parameters captures individual differences in height, weight, and muscle mass. Definition, represents the position of the human body surface in three-dimensional space, and the vertices are connected by triangles to form a complete structure. By applying linear blend skinning, the mesh is deformed according to the joint transformation, thereby ensuring smooth movement during joint movement and making the joints have anatomically consistent movement.
[0027] Embodiment 2, based on embodiment 1, in this embodiment, the relative similarity of any two skeleton postures in step S100 is calculated, including: obtaining the joint positions of the human body when it is in any two of the skeleton postures; determining the relative similarity of any two of the skeleton postures according to the joint positions corresponding to the any two skeleton postures : ; represents similarity, N is the total number of joints, C is the coordinate dimension, is the skeleton pose of any two skeleton poses , is the skeleton pose of any two skeleton poses , for The The joint Elements ( The joint The first element joint positions), when When it is 2, that is, two-dimensional, the The horizontal coordinate Or vertical coordinate ; In three dimensions, The horizontal coordinate Or vertical coordinate Or vertical coordinate . for The The joint elements.
[0028] The minus sign in the and When the difference is large, it means the similarity is low, that is, when the similarity is low, it means and The two postures are quite different. The minus sign is also used to indicate when and When the difference is small, it means the similarity is high, that is, when the similarity is high, it means and The two poses are relatively small in difference. In summary, the negative sign is used to make two poses with large differences correspond to low similarity, and two poses with small differences correspond to high similarity.
[0029] After calculating the relative similarity of any two skeleton postures, this embodiment uses the relative similarity of any two skeleton postures to construct a relative posture similarity space, including the following specific steps: determining the reference posture and non-reference posture contained in any two of the skeleton postures; setting a blank preset space, and determining the position of the non-reference posture in the preset space according to the relative similarity between the non-reference posture and the reference posture; filling the relative similarity corresponding to any two of the skeleton postures containing the non-reference posture in the position to obtain the relative posture similarity space.
[0030] The reference posture is the T-shaped posture presented when the human body stretches naturally, and the non-reference posture is the posture after removing the reference posture from all skeleton postures. The relative similarity calculation formula mentioned above is used to calculate the relative similarity between one of the non-reference postures and the reference posture (this relative similarity is recorded as SIM1). The size of SIM1 determines the position of the non-reference posture in the blank preset space; then the relative similarity between the non-reference posture and any other non-reference posture is calculated (this relative similarity is recorded as SIM2), and SIM2 is filled in this position. The same operation is performed for all non-reference postures to obtain the relative posture similarity space, which is used to record the relative similarity between any two non-reference postures.
[0031] For example, if the blank preset space is a blank matrix, then SIM1 determines the number of rows in the space matrix corresponding to the relative similarity of one of the non-reference postures, and SIM2 is placed on the row corresponding to SIM1, and the size of SIM2 determines the number of columns in which it is located.
[0032] That is, each pose (i.e., non-reference pose) is represented based on its relative similarity to the canonical pose (i.e., the reference pose) and its relationship to other poses in the dataset. The canonical pose serves as a reference point in this space, and the positions of other poses are determined by their similarity or difference to the canonical pose and their relative arrangement in the pose distribution. The relative pose similarity space helps to identify diverse and representative poses from the dataset (e.g., anchor cues).
[0033] Embodiment 3, based on embodiment 1 or embodiment 2, in this embodiment, the anchor prompt set after initialization Contains reference poses and non-reference poses, anchor prompt set The reference poses and non-reference poses in are called anchor cues. Remove anchor hints set for all skeleton poses The skeleton pose afterward.
[0034] The step S400 of this embodiment includes the following specific steps of selecting the target posture: determining the maximum relative similarity between each unsampled skeleton posture and each anchor prompt according to the relative similarity between the unsampled skeleton posture and the anchor prompt; selecting the minimum maximum relative similarity from the maximum relative similarities corresponding to each unsampled skeleton posture, and taking the unsampled skeleton posture corresponding to the minimum maximum relative similarity as the target posture. .
[0035] That is, each cycle never samples the posture set Unsampled skeleton poses in Filter out target poses as anchor prompt sets The next anchor prompt in the . The specific process is as follows: Corresponding to the unsampled pose set Each unsampled skeleton pose in , that is, calculation ( That is unsampled skeleton poses) and a set of anchor hints The similarities of all anchor prompts in and select the largest one: ; In the formula, Anchor prompt collection Anchor hint in .
[0036] ; In the formula Represents the solution to satisfy of , that is, the unsampled pose with the least similarity to the existing anchor prompt is taken as the next sample to ensure that the newly selected pose has the largest difference with the existing anchor prompt set, thereby increasing the diversity of the sampled data.
[0037] For example, the unsampled pose set Including unsampled skeleton pose 1, unsampled skeleton pose 2, unsampled skeleton pose 3, anchor prompt set Includes anchor prompt one and anchor prompt two. Calculate the similarity sim11 between the unsampled skeleton pose one and the anchor prompt one, calculate the similarity sim12 between the unsampled skeleton pose one and the anchor prompt two, and select the maximum similarity from the similarity sim11 and the similarity sim12, and record it as Maxsim1; calculate the similarity sim21 between the unsampled skeleton pose two and the anchor prompt one, calculate the similarity sim22 between the unsampled skeleton pose two and the anchor prompt two, and select the maximum similarity from the similarity sim21 and the similarity sim22, and record it as Maxsim2; calculate the similarity sim31 between the unsampled skeleton pose three and the anchor prompt one, calculate the similarity sim32 between the unsampled skeleton pose three and the anchor prompt two, and select the maximum similarity from the similarity sim31 and the similarity sim32, and record it as Maxsim3. Then select the minimum similarity from Maxsim1, Maxsim2, and Maxsim3. If the minimum similarity is Maxsim3, then the unsampled skeleton pose three is the target pose of this cycle. .
[0038] Each cycle gets the target posture Afterwards, the target posture is removed from the unsampled posture set, and the target posture is added to the anchor prompt set to update the anchor prompt set, until the number of anchor prompts in the anchor prompt set reaches a set number, thereby obtaining a final set of anchor prompts.
[0039] Embodiment 4, based on embodiment 1 or embodiment 2 or embodiment 3, as Figure 4As shown, this embodiment matches the corresponding soft prompt for each anchor prompt in the final set of anchor prompts, so as to guide the reasoning process of the model by combining with the corresponding soft prompt. The model can not only identify key patterns similar to the query sample, but also make appropriate adjustments between different tasks according to the guidance of the soft prompt. That is, the representative anchor prompts in the anchor prompt set are used to dynamically adapt to the diverse task requirements, thereby significantly improving the generalization ability and task adaptability of the model. By selecting the most similar anchor prompt, the model can model the query sample (the query sample is the human skeleton posture as the sample) more accurately. At the same time, the introduction of soft prompts provides the model with more contextual information, helping it to better reason and make decisions in complex tasks. This embodiment is based on the joint strategy of anchor prompts and soft prompts, which not only enhances the expressiveness of the model, but also improves its flexibility and adaptability when facing different tasks.
[0040] This embodiment also provides a cross-domain human body unified modeling device, such as Figure 5 As shown, the device comprises the following components: The relative similarity calculation module 01 is used to obtain the skeleton posture of the human body, determine the relative similarity between any two skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two skeleton postures; A set setting module 02 is used to set an anchor prompt set, and initialize the anchor prompt set using the skeleton posture and define an unsampled posture set; An information extraction module 03, configured to extract the relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; The updating module 04 is used to filter out the target posture from the unsampled posture set according to the relative similarity between the unsampled skeleton posture and the anchor prompt, and use the target posture to update the anchor prompt set to obtain a final set of anchor prompts. The final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when the model is trained based on the reasoning task.
[0041] Based on the above embodiments, the present invention further provides a terminal device, whose principle block diagram can be shown as follows: Figure 6As shown. The terminal device includes a processor, a memory, a network interface, and a display screen connected via a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a cross-domain unified human body modeling method is implemented. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.
[0042] Those skilled in the art will understand that Figure 6 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal device to which the scheme of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0043] In one embodiment, a terminal device is provided, the terminal device comprising a memory, a processor, and a cross-domain human body unified modeling program stored in the memory and executable on the processor, and when the processor executes the cross-domain human body unified modeling program, the following operation instructions are implemented: Acquire a human skeleton posture, determine the relative similarity between any two of the skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two of the skeleton postures; Setting an anchor hint set, and initializing the anchor hint set and defining an unsampled pose set using the skeleton pose; Extracting a relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; According to the relative similarity between the unsampled skeleton posture and the anchor prompt, the target posture is screened out from the unsampled posture set, and the anchor prompt set is updated with the target posture to obtain a final set of anchor prompts. The final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.
[0044] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-domain human unified modeling method, characterized in that: include: Acquire a human skeleton posture, determine the relative similarity between any two of the skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two of the skeleton postures; Setting an anchor hint set, and initializing the anchor hint set and defining an unsampled pose set using the skeleton pose; Extracting a relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; According to the relative similarity between the unsampled skeleton posture and the anchor prompt, the target posture is screened out from the unsampled posture set, and the anchor prompt set is updated with the target posture to obtain a final set of anchor prompts. The final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body. The soft prompts are learning parameters generated when training the model based on the reasoning task.
2. The cross-domain human unified modeling method according to claim 1, characterized in that: Determining the relative similarity of any two of the skeleton poses includes: Obtaining joint positions of the human body when it is in any two of the skeleton postures; The relative similarity between any two skeleton postures is determined based on the joint positions corresponding to any two skeleton postures.
3. The cross-domain human unified modeling method according to claim 2, characterized in that: Based on the relative similarity of any two skeleton postures, a relative posture similarity space is constructed, including: Determine a reference pose and a non-reference pose included in any two of the skeleton poses; Setting a blank preset space, and determining a position of the non-reference gesture in the preset space according to a relative similarity between the non-reference gesture and the reference gesture; The relative similarities corresponding to any two of the skeleton postures including the non-reference posture are filled in the position to obtain the relative posture similarity space.
4. The cross-domain human unified modeling method according to claim 1, characterized in that: The skeleton poses in the unsampled pose set are the skeleton poses after removing the anchor prompt set.
5. The cross-domain human unified modeling method according to claim 1, characterized in that: Filtering a target pose from the unsampled pose set according to a relative similarity between the unsampled skeleton pose and the anchor hint comprises: Determining a maximum relative similarity between each of the unsampled skeleton poses and each of the anchor cues based on the relative similarities between the unsampled skeleton poses and the anchor cues; The smallest maximum relative similarity is screened out from the maximum relative similarities corresponding to each of the unsampled skeleton postures, and the unsampled skeleton posture corresponding to the smallest maximum relative similarity is used as the target posture.
6. The cross-domain human unified modeling method according to claim 1, characterized in that: Updating the anchor prompt set with the target posture to obtain a final set of anchor prompts, including: The target posture is removed from the unsampled posture set, and the target posture is added to the anchor prompt set to update the anchor prompt set until the number of anchor prompts in the anchor prompt set reaches a set number, thereby obtaining a final set of anchor prompts.
7. The cross-domain human unified modeling method according to any one of claims 1 to 6, characterized in that: The modality of the skeleton posture includes a two-dimensional modality, a three-dimensional modality and a mesh modality; the two-dimensional modality contains the serial number of the skeleton posture and the two-dimensional coordinates of the joints of the skeleton posture; the three-dimensional modality contains the serial number of the skeleton posture and the three-dimensional coordinates of the joints of the skeleton posture; the mesh modality includes the joint rotation information of the skeleton posture and the shape of the skeleton posture.
8. A cross-domain human body unified modeling device, characterized in that: The device comprises the following components: A relative similarity calculation module is used to obtain the skeleton posture of the human body, determine the relative similarity between any two skeleton postures, and construct a relative posture similarity space based on the relative similarity between any two skeleton postures; A set setting module, used for setting an anchor prompt set, initializing the anchor prompt set using the skeleton posture and defining an unsampled posture set; an information extraction module, configured to extract a relative similarity between an unsampled skeleton pose and an anchor prompt from the relative pose similarity space, wherein the unsampled skeleton pose is a skeleton pose in the unsampled pose set, and the anchor prompt is a skeleton pose in the anchor prompt set; An updating module is used to filter out a target posture from the unsampled posture set based on the relative similarity between the unsampled skeleton posture and the anchor prompt, and use the target posture to update the anchor prompt set to obtain a final set of anchor prompts, wherein the final set of anchor prompts is used to cooperate with soft prompts to complete the model's reasoning task for the human body, and the soft prompts are learning parameters generated when the model is trained based on the reasoning task.
9. A terminal device, characterized in that: The terminal device includes a memory, a processor, and a cross-domain human body unified modeling program stored in the memory and executable on the processor. When the processor executes the cross-domain human body unified modeling program, the steps of the cross-domain human body unified modeling method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a cross-domain human body unified modeling program, and when the cross-domain human body unified modeling program is executed by the processor, the steps of the cross-domain human body unified modeling method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Three-dimensional human body posture estimation method, device, equipment and medium
CN117711066A
Cross-domain attitude estimation method based on skeleton graph structure constraint and application
CN118865486A
Posture-driven multi-layer perception method and system for character interaction behaviors
CN119649464A