Data processing method and apparatus thereof

By calculating the similarity information between the skeleton nodes and weighting the sum, the problem of large errors during the action migration process is solved, and the quality of action generation and the improvement of digital human visual effects are achieved.

WO2025092554A1PCT designated stage expired Publication Date: 2025-05-08HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126997
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-10-24
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art is prone to introduce action errors during the movement migration process, resulting in unsmooth and natural movements, mold penetration, slipping, etc., which affects the visual effect of digital people.

Method used

By obtaining the feature representation of the two skeleton topology, the similarity information between the skeleton nodes is calculated and summed as weights as weights, and the action data feature representation is fused to achieve high-quality migration of the action.

Benefits of technology

It improves the quality of action generation, reduces errors during action migration, and improves the visual effect and persuasion of digital people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126997_08052025_PF_FP_ABST
    Figure CN2024126997_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method, relating to the field of artificial intelligence, and comprises: acquiring a first feature representation and a second feature representation, the first feature representation being a feature representation of a skeleton topology of a first object, and the second feature representation being a feature representation of a skeleton topology of a second object; obtaining similarity information between a second skeleton node in the skeleton topology of the second object and each first skeleton node in the skeleton topology of the first object on the basis of the first feature representation and the second feature representation; fusing feature representations of action data corresponding to the plurality of first skeleton nodes on the basis of the weight of the first skeleton nodes to obtain a feature representation of action data of the second skeleton node; and obtaining action data of the second skeleton node on the basis of the feature representation of the second skeleton node. According to the present application, the action migration is carried out on the basis of the similarity relationship between the skeleton nodes of different skeleton topologies, and the action migration of different skeleton topologies is implemented.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method and device thereof

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on October 30, 2023, with application number 202311428239.4 and application name “A data processing method and device thereof”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to a data processing method and device thereof. Background Art

[0003] With technological advancements, digital humans are increasingly being used in a variety of scenarios, including digital idols, virtual social interactions, intelligent diagnosis and treatment, hosts, and even digital astronauts. The interaction and cognition of digital humans in these scenarios require the creation of virtual humans with natural movements, rich expressions, and intelligent brains.

[0004] For example, in conversational digital humans, body language can enhance the rhythm of a speech, making it more vivid and persuasive. Research has shown that body language plays a crucial role in communication. First, body language more accurately conveys intent and emotion, complementing the message conveyed by speech. Second, body language helps users focus on the content of the digital human communication. Third, it enhances the digital human's persuasiveness, credibility, and authenticity. Finally, it reflects the speaker's intentions and personality. A lack of body language or rigid body language in communication can lead to the uncanny valley effect. Furthermore, the free movement of digital humans in the metaverse, including running, walking, and various other movements, makes them more human-like and truly moving.

[0005] However, digital humans come in a wide variety of types, with a wide range of skeletal topologies. Since motion generation models can only learn a limited number of these skeletal topologies, they face the challenge of motion transfer. In existing technologies, this process introduces numerous errors, leading to various motion issues, including inaccurate and distorted movements, unsmooth and unnatural movements, clipping, and sliding. This significantly impacts the visual quality of the digital human and results in poor quality of generated motion.

[0006] Summary of the Invention

[0007] This application provides a data processing method that can improve the quality of action generation.

[0008] In a first aspect, the present application provides a data processing method, the method comprising:

[0009] A feature representation of the skeleton topology of the first object (i.e., the first feature representation) and a feature representation of the skeleton topology of the second object (i.e., the second feature representation) are obtained. In order to be able to migrate the action data of the first object to the second object, the similarity information between the second skeleton node and each of the first skeleton nodes can be obtained based on the relationship between the first feature representation and the second feature representation, and the similarity information can be used as a weight (or the similarity information can be mapped to a weight through certain data processing) to fuse the feature representations of the action data corresponding to the multiple first skeleton nodes to obtain the feature representation of the action data of the second skeleton node; based on the feature representation of the second skeleton node, the action data of the second skeleton node is obtained, thereby migrating the action data of the first object to the second object.

[0010] In an embodiment of the present application, the skeleton topology of the first object includes a plurality of first skeleton nodes (which may be part or all of the skeleton nodes of the skeleton topology of the first object), and the skeleton topology of the second object includes a second skeleton node;

[0011] For example, the similarity information can be represented by a score, such as 0.1, 0.2.

[0012] In an embodiment of the present application, the similarity relationship between skeleton nodes of different skeleton topologies is determined based on the relationship between feature representations of different skeleton topologies, and action migration is performed based on the similarity relationship, thereby realizing action migration of different skeleton topologies.

[0013] In one possible implementation, there are multiple second skeleton nodes (which may be part or all of the skeleton nodes of the skeleton topology of the second object); based on the first feature representation and the second feature representation, the similarity information between the second skeleton node and each of the first skeleton nodes is obtained, and based on the weight of the first skeleton node, the feature representations of the action data corresponding to the multiple first skeleton nodes are fused to obtain the feature representation of the second skeleton node; based on the feature representation of the second skeleton node, the action data of the second skeleton node is obtained, including: based on the first feature representation and the second feature representation, obtaining the similarity information between each second skeleton node and each of the first skeleton nodes; based on the weight of the first skeleton node, the feature representations of the action data corresponding to the multiple first skeleton nodes are fused to obtain the feature representation of the action data of each skeleton node; based on the feature representation of the action data of each skeleton node, the action data of each skeleton node is obtained.

[0014] In one possible implementation, the feature representations of the action data corresponding to the multiple first skeleton nodes are fused according to the weights of the first skeleton nodes, including: using the similarity information as the weight and performing weighted summation of the feature representations of the action data corresponding to the multiple first skeleton nodes.

[0015] In one possible implementation, the similarity information between the second skeleton node and each of the first skeleton nodes is obtained based on the first feature representation and the second feature representation, including: obtaining the similarity information between the second skeleton node and each of the first skeleton nodes by performing dot multiplication and activation operations on the first feature representation and the second feature representation.

[0016] In a possible implementation, the skeleton topology includes positions of skeleton nodes and connection relationships between skeleton nodes.

[0017] In a possible implementation, the motion data includes a rotation angle of a skeleton node.

[0018] In order to determine the relationship between the skeleton topologies more accurately, the interference of other information except the structural information of the skeleton topology can be excluded (such as the difference in node positions due to different body postures, for example, one object is in a lying posture and the other is in a standing posture). Therefore, in one possible implementation, the skeleton topology of the first object and the skeleton topology of the second object correspond to the same body posture.

[0019] In one possible implementation, the first feature representation is obtained by encoding the skeleton topology of the first object through a first encoder, and the second feature representation is obtained by encoding the skeleton topology of the second object through a second encoder; the skeleton topology of the first object and the skeleton topology of the second object are the same; the first encoder and the second encoder can be updated based on the similarity information and the corresponding true value, wherein the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is greater than the similarity between the second skeleton node and the first skeleton node at the non-corresponding position.

[0020] The fact that the skeleton topology of the first object is the same as that of the second object can be understood as having the same topology type, for example, the number of joint points and the connection relationship are the same, and the lengths between the joints vary proportionally.

[0021] For example, the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is 1, and the similarity between the second skeleton node and the first skeleton node at a non-corresponding position is 0.

[0022] In one possible implementation, the first encoder and the second encoder are graph neural networks.

[0023] In a possible implementation, the first object and the second object are characters; and the method further includes:

[0024] The digital human corresponding to the second object is reconstructed according to the motion data of the second skeleton node.

[0025] In a second aspect, the present application provides a data processing device, the device comprising:

[0026] an acquisition module, configured to acquire a first feature representation and a second feature representation, wherein the first feature representation is a feature representation of a skeleton topology of a first object, and the second feature representation is a feature representation of a skeleton topology of a second object, wherein the skeleton topology of the first object includes a plurality of first skeleton nodes, and the skeleton topology of the second object includes a second skeleton node;

[0027] a processing module, configured to obtain similarity information between the second skeleton node and each of the first skeleton nodes based on the first feature representation and the second feature representation;

[0028] fusing feature representations of motion data corresponding to a plurality of first skeleton nodes according to the weights of the first skeleton nodes to obtain a feature representation of the motion data of the second skeleton node;

[0029] According to the feature representation of the second skeleton node, the action data of the second skeleton node is obtained.

[0030] In a possible implementation, there are multiple second skeleton nodes;

[0031] The processing module is specifically used to:

[0032] Based on the first feature representation and the second feature representation, the similarity information between each skeleton node included in the skeleton topology of the second object and each of the first skeleton nodes is obtained. According to the weight of the first skeleton node, the feature representations of the action data corresponding to multiple first skeleton nodes are fused to obtain the feature representation of the action data of each skeleton node. Based on the feature representation of the action data of each skeleton node, the action data of each skeleton node is obtained.

[0033] In a possible implementation, the processing module is specifically configured to:

[0034] The similarity information is used as a weight to perform weighted summation on the feature representations of the action data corresponding to the multiple first skeleton nodes.

[0035] In a possible implementation, the processing module is specifically configured to:

[0036] Similarity information between the second skeleton node and each of the first skeleton nodes is obtained by performing a dot product and an activation operation on the first feature representation and the second feature representation.

[0037] In a possible implementation, the skeleton topology includes positions of skeleton nodes and connection relationships between skeleton nodes.

[0038] In a possible implementation, the motion data includes a rotation angle of a skeleton node.

[0039] In one possible implementation, the skeleton topology of the first object and the skeleton topology of the second object correspond to the same body pose.

[0040] In one possible implementation, the first feature representation is obtained by encoding a skeleton topology of a first object using a first encoder, and the second feature representation is obtained by encoding a skeleton topology of a second object using a second encoder; the skeleton topology of the first object and the skeleton topology of the second object are the same;

[0041] The processing module is also used to:

[0042] The first encoder and the second encoder are updated according to the similarity information and the corresponding true value, wherein the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is greater than the similarity between the second skeleton node and the first skeleton node at the non-corresponding position.

[0043] In one possible implementation, the first encoder and the second encoder are graph neural networks.

[0044] In a possible implementation, the first object and the second object are people; and the processing module is further configured to:

[0045] The digital human corresponding to the second object is reconstructed according to the motion data of the second skeleton node.

[0046] In a third aspect, an embodiment of the present application provides a data processing device, which may include a memory, a processor, and a bus system, wherein the memory is used to store programs, and the processor is used to execute the programs in the memory to perform the first aspect and any optional method thereof.

[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned first aspect and any optional method thereof.

[0048] In a fifth aspect, an embodiment of the present application provides a computer program, which, when executed on a computer, enables the computer to execute the above-mentioned first aspect and any optional method thereof.

[0049] In a sixth aspect, the present application provides a chip system comprising a processor configured to support a data processing device in implementing the functions described in the aforementioned aspects, such as transmitting or processing the data or information described in the aforementioned methods. In one possible design, the chip system further comprises a memory configured to store program instructions and data necessary for executing or training the device. The chip system may consist solely of a chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] FIG1A is a schematic diagram of a structure of an artificial intelligence main framework;

[0051] FIG1B is a schematic diagram of the structure of the skeleton topology;

[0052] FIG1C, FIG1D and FIG2 are schematic diagrams of the application system framework of the present invention;

[0053] FIG3 is a schematic diagram of an optional hardware structure of a terminal;

[0054] FIG4 is a schematic diagram of the structure of a server;

[0055] FIG5 is a schematic diagram of a system architecture of the present application;

[0056] Figure 6 shows a cloud service process;

[0057] Figure 7 is a process for generating an action;

[0058] FIG8 is a flowchart of a data processing method provided in an embodiment of the present application;

[0059] FIG9 is a schematic diagram of skeleton topology information;

[0060] FIG10 is a schematic diagram of action data;

[0061] FIG11 is a schematic diagram of similarity information;

[0062] FIG12A is a flowchart of a data processing method provided in an embodiment of the present application;

[0063] FIG12B is a flowchart of a data processing method provided in an embodiment of the present application;

[0064] FIG12C is a flowchart of a data processing method provided in an embodiment of the present application;

[0065] FIG13 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0066] FIG14 is a schematic diagram of a structure of an execution device provided in an embodiment of the present application;

[0067] FIG15 is a schematic diagram of a structure of a training device provided in an embodiment of the present application;

[0068] FIG16 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0069] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.

[0070] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0071] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0072] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent variations in measurements or calculations that one of ordinary skill in the art would recognize. Furthermore, the use of "may" when describing embodiments of the present invention refers to "one or more possible embodiments." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.

[0073] First, let's describe the overall workflow of an AI system. See Figure 1A, which shows a schematic diagram of the main AI framework. This AI framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0074] (1) Infrastructure

[0075] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.

[0076] (2) Data

[0077] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0078] (3) Data processing

[0079] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0080] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0081] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0082] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0083] (4) General ability

[0084] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0085] (5) Smart products and industry applications

[0086] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.

[0087] First, we will introduce the application scenarios of this application. This application can be applied to, but is not limited to, applications that include human motion generation functions based on audio or text (hereinafter referred to as human motion generation applications) or cloud services provided by cloud-side servers. The following are introduced respectively:

[0088] 1. Human motion generation applications

[0089] The product form of the embodiment of the present application can be a human motion generation application. The human motion generation application can be run on a terminal device or a cloud-side server.

[0090] 1. Action generation based on text or audio

[0091] In one possible implementation, a human motion generation application can implement the task of generating human motion based on audio or text, wherein the human motion generation application can perform the task of generating human motion in response to the input audio or text, and obtain motion data corresponding to the human motion (such as 3D posture information). Based on the motion data, a virtual character can be restored, and the character's posture and audio (if it is a human motion generation task based on text, the audio can be generated based on text) can be aligned in rhythm and rhythm.

[0092] Here is an example application scenario:

[0093] The terminal receives user input, for example, a user voice-comments the intelligent assistant, "Are there any good movies available lately?" Upon receiving this input, the terminal performs recognition, perception, and understanding, converting the speech into text to determine the user's true intent. After understanding the user's question, the intelligent analysis and decision-making module generates an answer, such as "Recently released movies include Movie A, Movie B, and Movie C." This textual response is then fed into the character speech generation module and the character motion generation module to generate the corresponding speech, facial, and body motion parameters. Based on the speech and parameters generated in the previous stage, the audio and video synthesis and display module synthesizes the final presentation and displays it to the user through the terminal.

[0094] In one possible implementation, the user can open a human motion generation application installed on a terminal device and input (which can be active input or passive collection, for example, collected through sensors on the terminal device). The human motion generation application can generate human motions for audio or text using the method provided in the embodiment of the present application, and present the 3D posture information or a virtual character restored based on the 3D posture information to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).

[0095] In one possible implementation, a user can open a human motion generation application installed on a terminal device and input audio or text (which can be active input or passive collection, for example, collected through a camera on the terminal device). The human motion generation application can send the audio or text to a cloud-side server. The cloud-side server generates human motions for the audio or text using the method provided in an embodiment of the present application, and transmits the 3D posture information or the information of a virtual character restored based on the 3D posture information back to the terminal device. The terminal device can present the 3D posture information or the virtual character restored based on the 3D posture information to the user (the presentation method can be, but is not limited to, display, saving, uploading to the cloud side, etc.).

[0096] 2. Migration of motion data

[0097] In one possible implementation, a human motion generation application can migrate motion data between different skeletal topologies. This means assigning the motion of an object in skeletal topology A to an object in skeletal topology B (which is different from skeletal topology A), allowing the object in skeletal topology B to perform the same (or nearly the same) motion as the object in skeletal topology A. The term "different skeletal topologies" can be understood as, but is not limited to, different numbers of skeletal nodes, different connections between skeletal nodes, or disproportionate distances between skeletal nodes.

[0098] For example, when the number of nodes of skeleton topology A and skeleton topology B is inconsistent, it can be considered that skeleton topology A and skeleton topology B are different skeleton topologies.

[0099] For example, when the node connection relationships of skeleton topology A and skeleton topology B are inconsistent, it can be considered that skeleton topology A and skeleton topology B are different skeleton topologies.

[0100] For example, when skeleton topology A and skeleton topology B have the same number of nodes and connection relationships, but skeleton topology A is not the result of proportional enlargement or proportional reduction of skeleton topology B, it can be considered that skeleton topology A and skeleton topology B are different skeleton topologies.

[0101] For example, referring to FIG. 1B , FIG. 1B is a diagram illustrating two different skeletal topologies of a hand.

[0102] In some scenarios, due to limitations in training data and cost, the action generation network is often targeted at a specific skeleton topology. When the action of the skeleton topology to be generated is an action that the action generation network cannot generate, the action generation network can first generate an action data, and then migrate the action data to the object of the skeleton topology to be generated. Referring to Figure 1C, which is a schematic diagram of an architecture corresponding to this scenario, an exemplary application scenario is given:

[0103] The terminal receives input from the user. After receiving the input, the terminal recognizes, perceives and understands the virtual character's brain, converts the voice into text, and then finds out the user's true intention. After knowing the user's question, the intelligent analysis and decision-making module obtains the answer. After obtaining the answer text, the text is input into the character voice generation module and the character animation (or action) generation module to generate the corresponding voice and facial and body movement parameters. Among them, the character animation generation module may include a limb movement migration module for migrating the generated action to the object of the skeleton topology to be generated. Based on the voice and parameters obtained in the previous stage, the audio and video synthesis display module synthesizes the final display effect and displays it to the user through the terminal.

[0104] In some scenarios, it is necessary to generate a corresponding digital human based on the actions collected from real people. The digital human can make the same actions as the real person, and the skeletal topology of the digital human is different from that of the real person.

[0105] In one possible implementation, a user can open a human motion generation application installed on a terminal device and input data of a certain character (the data can be collected by a sensor, such as an image, and the data can be actively input or passively collected, such as collected by a sensor on the terminal device). The human motion generation application can migrate the motion data of the character to another skeletal topology character through the method provided in the embodiment of the present application, generate human motion, and present the 3D posture information or the virtual character restored based on the 3D posture information to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).

[0106] In one possible implementation, a user can open a human motion generation application installed on a terminal device and input data of a certain character (which can be active input or passive collection, such as collection through a camera on the terminal device). The human motion generation application can send the data to a server on the cloud side. The server on the cloud side migrates the motion data of the character to another skeletal topology character through the method provided in an embodiment of the present application, generates human motions, and transmits the 3D posture information or the information of the virtual character restored based on the 3D posture information back to the terminal device. The terminal device can present the 3D posture information or the virtual character restored based on the 3D posture information to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).

[0107] In one possible implementation, the human motion generation function implemented by the human motion generation application can be specifically used to enable the driving of virtual characters in application scenarios such as augmented reality (AR), virtual reality (VR), and mixed reality (MR) remote conferencing, sports health, and the metaverse.

[0108] It should be understood that the generation of the above-mentioned action objects is not limited to human beings, but can also be non-human beings (such as animals, robots), etc. In particular, in the action transfer task, the action of a human being can be transferred to a non-human being.

[0109] Next, the human motion generation application in the embodiment of this application is introduced from the functional architecture and the product architecture that realizes the function.

[0110] Referring to FIG. 1D , FIG. 1D is a schematic diagram of the functional architecture of a human motion generation application in an embodiment of the present application:

[0111] In one possible implementation, as shown in FIG1D , a human motion generation application 102 may receive input parameters 101 (e.g., audio or text) and generate 3D pose information 103 (or information of a virtual character restored based on the 3D pose information). The human motion generation application 102 may be executed on, for example, at least one computer system and may include computer code that, when executed by one or more computers, causes the computers to execute the data processing method described herein.

[0112] Referring to FIG. 2 , FIG. 2 is a schematic diagram of the physical architecture of a human motion generation application program according to an embodiment of the present application:

[0113] Referring to Figure 2, a schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (one server is used as an example in Figure 2), and the server 200 may provide an action generation service for one or more terminals.

[0114] Among them, the terminal 100 can be installed with a human motion generation application, or a web page related to the motion generation function can be opened. The above application and web page can provide an interface. The terminal 100 can receive the relevant parameters entered by the user on the motion generation function interface and send the above parameters to the server 200. The server 200 can obtain the processing results based on the received parameters and return the processing results to the terminal 100.

[0115] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the data processing result based on the received parameters by itself without the cooperation of the server, and the embodiments of the present application are not limited to this.

[0116] Next, the product form of the terminal 100 in FIG2 is described;

[0117] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.

[0118] FIG3 shows a schematic diagram of an optional hardware structure of the terminal 100 .

[0119] 3 , the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, and a power supply 190. Those skilled in the art will appreciate that FIG3 is merely an example of a terminal or multi-function device and does not limit the terminal or multi-function device. The terminal or multi-function device may include more or fewer components than shown, or may combine certain components or have different components.

[0120] The input unit 130 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the portable multifunction device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can detect user touch operations on or near it (for example, operations performed on or near the touch screen using a finger, joint, stylus, or any other suitable object) and drive corresponding connected devices according to pre-set programs. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and transmit them to the processor 170. It can also receive and execute commands sent by the processor 170; the touch signals include at least touch point coordinate information. The touch screen 131 provides an input interface and an output interface between the terminal 100 and the user. Touch screens can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control button 132 , a switch button 133 , etc.), a trackball, a mouse, a joystick, and the like.

[0121] The input device 132 may receive input audio or text, etc.

[0122] The display unit 140 may be used to display information input by or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In an embodiment of the present application, the display unit 140 may be used to display the interface of a human motion generation application, a virtual human generated based on audio or text, and the like.

[0123] Memory 120 can be used to store instructions and data. It primarily includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files and text. The instruction storage area can store software units such as the operating system, applications, and instructions required for at least one function, or subsets or extensions thereof. It may also include non-volatile random access memory (RAM). It provides processor 170 with management functions for the hardware, software, and data resources within the computing and processing device, supporting control software and applications. It is also used to store multimedia files and running programs and applications.

[0124] The processor 170 is the control center of the terminal 100. It connects all components of the terminal 100 using various interfaces and circuits. By executing instructions stored in the memory 120 and accessing data stored therein, it executes various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units. Preferably, the processor 170 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory may be implemented on a single chip; in other embodiments, they may be implemented on separate chips. The processor 170 may also generate corresponding operational control signals and send them to the corresponding components of the computing and processing device. It may also read and process data in the software, particularly the data and programs in the memory 120, to enable the various functional modules therein to perform their corresponding functions, thereby controlling the corresponding components to operate as instructed.

[0125] Among them, the memory 120 can be used to store software codes related to the data processing method, the processor 170 can execute the steps of the chip's data processing method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve corresponding functions.

[0126] The RF unit 110 (optional) can be used to send and receive information or receive and send signals during a call. For example, after receiving downlink information from the base station, it is passed to the processor 170 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices via wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0127] In this embodiment of the present application, the RF unit 110 can send audio or text to the server 200, and receive 3D posture information sent by the server 200 or information of a virtual character restored based on the 3D posture information.

[0128] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.

[0129] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0130] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100 .

[0131] Although not shown, the terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied to the terminal 100 shown in FIG3 .

[0132] Next, the product form of the server 200 in FIG2 is described;

[0133] FIG4 is a schematic diagram of the structure of a server 200. As shown in FIG4, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.

[0134] Bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified as address buses, data buses, control buses, etc. For ease of illustration, FIG4 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0135] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0136] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard drive (HDD), or solid state drive (SSD).

[0137] The memory 204 may be used to store software codes related to the data processing method, and the processor 202 may execute the steps of the data processing method of the chip, and may also schedule other units to implement corresponding functions.

[0138] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0139] It should be understood that the steps related to the model reasoning process in the embodiments of this application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and server is not limited to the processor-memory architecture described above. The system architecture provided in the embodiments of this application is described in detail below with reference to Figure 5.

[0140] FIG5 is a schematic diagram of the system architecture provided by an embodiment of the present application. As shown in FIG5 , the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data acquisition system 560 .

[0141] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.

[0142] The execution device 510 may be a terminal device or a server that runs the above-mentioned human body motion generation application.

[0143] Data acquisition device 560 is used to collect training samples. Training samples can be audio or text, as well as annotations of character motion data within the audio or text. Furthermore, when training a motion transfer network, training samples can include information such as skeletal topology. After collecting the training samples, data acquisition device 560 stores them in database 530.

[0144] The training device 520 can train the neural network to be trained (such as the encoder, decoder, neural network, etc. in the embodiment of the present application) based on the training samples maintained in the database 530 to obtain the target model / rule 501.

[0145] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.

[0146] The target model / rule 501 obtained through training with the training device 520 can be applied to different systems or devices, such as the execution device 510 shown in FIG5 . The execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, an in-vehicle terminal, etc., or a server, etc.

[0147] Specifically, the training device 520 may transfer the trained model to the execution device 510 .

[0148] In Figure 5, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices. The user can input data (such as audio or text in the embodiment of the present application) to the I / O interface 512 through the client device 540.

[0149] Preprocessing module 513 and preprocessing module 514 are used to preprocess the input data received by I / O interface 512. It should be understood that preprocessing module 513 and preprocessing module 514 may be absent or only one preprocessing module may be present. If preprocessing module 513 and preprocessing module 514 are absent, computing module 511 may be used directly to process the input data.

[0150] When the execution device 510 preprocesses the input data, or when the computing module 511 of the execution device 510 performs calculations and other related processing, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.

[0151] Finally, the I / O interface 512 provides the processing results (such as the skeleton posture of the character or the information of the virtual character restored based on the skeleton posture of the character, etc.) to the client device 540, and then provides it to the user.

[0152] In the scenario shown in FIG5 , the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another scenario, client device 540 can automatically send input data to I / O interface 512. If user authorization is required for client device 540 to automatically send input data, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, which can be presented in a specific form such as display, sound, or action. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results from I / O interface 512 as new sample data and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 can directly store the input data input into I / O interface 512 and output results from I / O interface 512 as new sample data in database 530.

[0153] It is worth noting that FIG5 is merely a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in FIG5 , the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.

[0154] From the inference side of the model:

[0155] In the embodiment of the present application, the computing module 511 of the above-mentioned execution device 510 can obtain the code stored in the data storage system 550 to implement the steps related to the model reasoning process in the embodiment of the present application.

[0156] In an embodiment of the present application, the computing module 511 of the execution device 510 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0157] Specifically, the computing module 511 of the execution device 510 can be a hardware system with an execution instruction function, and the steps related to the model reasoning process provided in the embodiment of the present application can be software codes stored in the memory. The computing module 511 of the execution device 510 can obtain the software code from the memory and execute the obtained software code to implement the steps related to the model reasoning process provided in the embodiment of the present application.

[0158] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to the model reasoning process provided in the embodiment of the present application can also be implemented by the hardware system that does not have the function of executing instructions in the computing module 511 of the execution device 510, which is not limited here.

[0159] From the training side of the model:

[0160] In an embodiment of the present application, the above-mentioned training device 520 can obtain the code stored in the memory (not shown in Figure 5, which can be integrated into the training device 520 or deployed separately from the training device 520) to implement the steps related to model training in the embodiment of the present application.

[0161] In an embodiment of the present application, the training device 520 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0162] It should be understood that the training device 520 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to model training provided in the embodiments of the present application can also be implemented by the hardware system in the training device 520 that does not have the function of executing instructions, which is not limited here.

[0163] 2. Action generation cloud services provided by the server:

[0164] In a possible implementation, the server may provide the terminal with an action generation function service through an application programming interface (API).

[0165] Among them, the terminal device can send relevant parameters (such as audio or text containing the human body) to the server through the API provided by the cloud. The server can obtain processing results (such as the skeleton posture of the character or the information of the virtual character restored based on the skeleton posture of the character, etc.) based on the received parameters, and return the processing results to the terminal.

[0166] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.

[0167] FIG6 shows a process of using an action generation function cloud service provided by a cloud platform.

[0168] 1. Activate and purchase content review services.

[0169] 2. Users can download the software development kit (SDK) corresponding to the content review service. Usually, the cloud platform provides multiple development versions of the SDK for users to choose according to the requirements of the development environment, such as JAVA version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.

[0170] 3. After the user downloads the corresponding version of the SDK to the local computer as needed, import the SDK project into the local development environment, configure and debug it in the local development environment. The local development environment can also be used to develop other functions, forming an application that integrates action generation functional capabilities.

[0171] 4. When an action generation application is in use and needs to perform an action generation function, it can trigger an API call for the action generation function. When the application triggers the action generation function, it initiates an API request to the running instance of the action generation service in the cloud environment. The API request carries audio or text, and the running instance in the cloud environment processes the audio or text to obtain the processing results (such as the skeletal posture of a character or information about a virtual character restored based on the skeletal posture of the character).

[0172] 5. The cloud environment returns the processing results to the application, thereby completing an action generation function service call.

[0173] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.

[0174] (1) Neural Network

[0175] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be:

[0176] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0177] (2) Deep Neural Networks

[0178] A deep neural network (DNN) can be understood as a neural network with many hidden layers. There is no special metric for "many" here. The multi-layer neural networks and deep neural networks we often talk about are essentially the same thing. According to the position of different layers of DNN, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector So how are the specific parameters defined in DNN? First, let's look at the definition of coefficient W. Take a three-layer DNN as an example, for example: the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as Note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters is more complex and has greater "capacity," meaning it can handle more complex learning tasks.

[0179] (3) Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor consisting of a convolution layer and a subsampling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve an input image or convolution feature plane (feature map). The convolution layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolution layer of the convolutional neural network, a neuron can only be connected to some neurons in the adjacent layer. A convolution layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are the convolution kernels. Shared weights can be understood as the way to extract image information is independent of position. The implicit principle is that the statistical information of a part of the image is the same as that of other parts. This means that the image information learned in a part can also be used in another part. Therefore, for all positions on the image, we can use the same learned image information. In the same convolutional layer, multiple convolution kernels can be used to extract different image information. Generally speaking, the more convolution kernels there are, the richer the image information reflected by the convolution operation.

[0180] Convolution kernels can be initialized as matrices of random size, and during the training process of the convolutional neural network, the convolution kernels can be learned to obtain reasonable weights. In addition, the direct benefit of shared weights is that they reduce the number of connections between the layers of the convolutional neural network, while also reducing the risk of overfitting.

[0181] (4) Backpropagation algorithm

[0182] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.

[0183] (5) Loss function

[0184] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.

[0185] (6) Backpropagation algorithm

[0186] Neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial neural network model during training, reducing the reconstruction error loss of the neural network model. Specifically, the forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial neural network model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0187] With technological advancements, digital humans are increasingly being used in a variety of scenarios, including digital idols, virtual social interactions, intelligent diagnosis and treatment, hosts, and even digital astronauts. The interaction and cognition of digital humans in these scenarios require the creation of virtual humans with natural movements, rich expressions, and intelligent brains.

[0188] For example, in conversational digital humans, body language can enhance the rhythm of a speech, making it more vivid and persuasive. Research has shown that body language plays a crucial role in communication. First, body language more accurately conveys intent and emotion, complementing the message conveyed by speech. Second, body language helps users focus on the content of the digital human communication. Third, it enhances the digital human's persuasiveness, credibility, and authenticity. Finally, it reflects the speaker's intentions and personality. A lack of body language or rigid body language in communication can lead to the uncanny valley effect. Furthermore, the free movement of digital humans in the metaverse, including running, walking, and various other movements, makes them more human-like and truly moving.

[0189] Digital humans come in a wide variety of types, with a vast array of skeletal topologies. However, the motion generation model can only learn a limited number of these skeletal topologies. Consequently, the process of motion transfer becomes challenging. In existing technologies, this process introduces numerous motion errors, leading to various motion issues, including inaccurate and distorted movements, unsmooth and unnatural movements, model clipping, and slippage. This significantly impacts the visual quality of the digital human and significantly hinders the development and implementation of digital human products. For example, Figure 7 illustrates a framework for digital human generation.

[0190] Based on the above problems, an embodiment of the present application proposes a method for action migration across skeleton topologies, which can realize action migration between different skeleton topologies.

[0191] The present invention provides a data processing method. The data processing method of the present invention is described in detail below with reference to the accompanying drawings.

[0192] Referring to Figure 8, Figure 8 is a flow chart of a data processing method provided in an embodiment of the present application. As shown in Figure 8, a data processing method provided in an embodiment of the present application may include steps 801 to 804, and these steps are described in detail below.

[0193] 801. Obtain a first feature representation and a second feature representation, wherein the first feature representation is a feature representation of the skeleton topology of a first object, and the second feature representation is a feature representation of the skeleton topology of a second object, the skeleton topology of the first object includes multiple first skeleton nodes, and the skeleton topology of the second object includes second skeleton nodes.

[0194] The first object and the second object may be a person, an animal, or a robot, etc.

[0195] In an embodiment of the present application, in order to migrate the action of the skeletal topology of a first object to the skeletal topology of a second object, a feature representation of the skeletal topology of the first object and a feature representation of the skeletal topology of the second object can be obtained, and the structural relationship between the skeletal topology of the first object and the skeletal topology of the second object can be determined based on the relationship between the feature representation of the skeletal topology of the first object and the feature representation of the skeletal topology of the second object, and the action data of the first object can be migrated to the second object based on the structural relationship.

[0196] In a possible implementation, skeleton topology information of the first object and skeleton topology information of the second object may be obtained.

[0197] The skeleton topology may include multiple key nodes (which may be referred to as skeleton nodes in the embodiment of the present application). For example, the skeleton nodes may include a head, a left hand, a right hand, a left foot, and a right foot.

[0198] In one possible implementation, the skeleton topology information may include the positions of the skeleton nodes and the connection relationships between the skeleton nodes. For example, the positions of the skeleton nodes may represent the relative positions between the nodes, such as by (x, y, z) coordinates, and the connection relationships between the skeleton nodes may include whether the nodes are connected and the parent-child relationship between the nodes with a connection relationship, that is, the connection relationship may be directional, such as the node closer to the torso facing the node relatively far from the torso, that is, the node closer to the torso is the parent node of the node relatively far from the torso.

[0199] In one possible implementation, skeleton topology can be represented as graph information, specifically the connection relationships between nodes. Each node corresponds to information about a skeleton node, and each edge represents the connection relationship between skeleton nodes. Node features can be used to represent joint definition locations, while edge features are used to represent node connection relationships and child-parent relationships. For example, see Figure 9, which illustrates a schematic representation of skeleton topology information.

[0200] It should be understood that in order to determine the relationship between the skeleton topologies more accurately, interference from other information besides the structural information of the skeleton topologies can be excluded (for example, differences in node positions due to different body postures, for example, one object is in a lying posture and the other is in a standing posture). Therefore, in one possible implementation, the skeleton topology of the first object and the skeleton topology of the second object correspond to the same body posture. For example, referring to Figure 11, Figure 11 shows a schematic diagram of two different skeleton topologies in the same body posture.

[0201] It should be understood that in order to determine the relationship between skeleton topologies more accurately, interference from other information other than the structural information of the skeleton topology (such as differences in node positions due to body shape, height, etc.) can be excluded. Therefore, in one possible implementation, the skeleton topology of the first object and the skeleton topology of the second object are topologies normalized by height or body shape.

[0202] In one possible implementation, the skeleton topology (information) of the first object and the skeleton topology (information) of the second object can be encoded by an encoder respectively to obtain a first feature representation and a second feature representation. Specifically, the skeleton topology of the first object can be processed by a first encoder to obtain a first feature representation, and the skeleton topology of the second object can be processed by a second encoder to obtain a second feature representation, wherein the first encoder and the second encoder can be the same or different encoders, and the embodiments of the present application are not limited thereto.

[0203] Among them, the multiple first skeleton nodes can be part or all of the nodes in the skeleton topology of the first object, which is not limited in this application.

[0204] In one possible implementation, the skeleton topology (information) of the first object and the skeleton topology (information) of the second object are represented as graph information, and accordingly, the first encoder and the second encoder can be graph neural networks.

[0205] 802. Obtain similarity information between the second skeleton node and each of the first skeleton nodes based on the first feature representation and the second feature representation;

[0206] In one possible implementation, the first feature representation and the second feature representation may respectively carry structural information of the skeleton topology of the first object and the skeleton topology of the second object. Based on an analysis of the relationship between the first feature representation and the second feature representation, a structural relationship between the skeleton topology and the skeleton topology of the second object may be obtained. For example, the structural relationship may include the similarity between each node in the skeleton topology of the second object and each node in the skeleton topology of the first object.

[0207] In a possible implementation, similarity information between each skeleton node included in the skeleton topology of the second object and each first skeleton node can be obtained based on the first feature representation and the second feature representation, and based on the similarity information.

[0208] Taking the example that the skeleton topology of the first object includes multiple first skeleton nodes and the skeleton topology of the second object includes second skeleton nodes, the similarity information between the second skeleton nodes and each of the first skeleton nodes can be obtained based on the first feature representation and the second feature representation.

[0209] When calculating the similarity information, the similarity information between the second skeleton node and each of the first skeleton nodes can be obtained by performing a dot product (ie, calculating the feature vector angle) and an activation operation on the first feature representation and the second feature representation.

[0210] For example, similarity information can be calculated as follows: s =GNN(x s ,e s )∈[N,C] z t =GNN(x t ,e t )∈[N,C]

[0211] Among them, the first feature is represented as z_s, the second feature is represented as z_t, and the attention weight matrix (that is, similarity information) is W.

[0212] When two skeleton topologies (for example, a first skeleton topology and a second skeleton topology) are consistent, it can be considered that the skeleton nodes included in the two skeleton topologies correspond one to one, and the similarity between the skeleton nodes in corresponding positions is higher than the similarity between the skeleton nodes in non-corresponding positions. For example, the first skeleton topology includes skeleton node A, and the second skeleton topology includes skeleton node B and skeleton node C. Skeleton node A and skeleton node B correspond to each other. Therefore, the similarity between skeleton node A and skeleton node B is greater than the similarity between skeleton node A and skeleton node C. When training the encoder, based on this idea, the true value corresponding to the similarity information can be constructed.

[0213] In one possible implementation, the first feature representation is obtained by encoding the skeleton topology of the first object through a first encoder, and the second feature representation is obtained by encoding the skeleton topology of the second object through a second encoder; the skeleton topology of the first object and the skeleton topology of the second object are the same; the first encoder and the second encoder can be updated according to the similarity information and the corresponding true value, wherein the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is greater than the similarity between the second skeleton node and the first skeleton node at the non-corresponding position.

[0214] The fact that the skeleton topology of the first object and the skeleton topology of the second object are the same can be understood as having the same topology type, for example, the number of joint points and the connection relationship are the same, and the lengths between the joints vary proportionally.

[0215] For example, the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is 1, and the similarity between the second skeleton node and the first skeleton node at a non-corresponding position is 0.

[0216] For example, when the topology types of the input skeleton (the skeleton topology of the first object) and the target skeleton (the skeleton topology of the second object) are the same, since the attention weight matrix representing the similarity information is the identity matrix, the topological consistency loss can be measured by the mean square error. Let the first feature be represented as z_s, the second feature be represented as z_t, and the attention weight matrix (that is, the similarity information) be W, then L_att is defined as:

[0217] 803. Based on the weights of the first skeleton nodes, the feature representations of the motion data corresponding to the plurality of first skeleton nodes are fused to obtain the feature representation of the motion data of the second skeleton node, wherein the weights are related to the similarity information.

[0218] In one possible implementation, motion data of the first object may be obtained, wherein the motion data may describe the dynamic motion of the first object at multiple moments, for example, the motion data may include the rotation angle of a skeleton node. For example, reference may be made to FIG10 , which is a schematic diagram of the motion data.

[0219] When migrating the motion data of the first object to the second object, the motion data of the first object can be acquired and migrated to the second object based on the structural relationship between the skeleton topology of the first object and the skeleton topology of the second object.

[0220] Specifically, based on the similarity between the skeleton nodes of the skeleton topology of the second object and each skeleton node of the skeleton topology of the first object, the action features of the skeleton nodes of the skeleton topology of the first object can be fused to obtain the action features of the skeleton nodes of the skeleton topology of the second object (the action features can be reconstructed into action data through a decoder).

[0221] Taking the calculation of the action data of the second skeleton node in the skeleton topology of the second object as an example, the feature representations of the action data corresponding to multiple first skeleton nodes can be fused according to the weights of the first skeleton nodes to obtain the feature representation of the action data of the second skeleton node, and the weight is related to the similarity information.

[0222] In a possible implementation, the similarity information can be used as a weight to perform weighted summation on the feature representations of the action data corresponding to the plurality of first skeleton nodes. The action features of the input structure are aggregated according to the similarity, that is, the weighted summation can refer to the following formula: s ∈[N s ,C m ] F t =w·F s ∈[N t ,C m];

[0223] 804. Obtain action data of the second skeleton node according to the feature representation of the second skeleton node.

[0224] In one possible implementation, the motion data of each skeleton node (including the second skeleton node) can be obtained based on the feature representation of the motion data of each skeleton node. For example, reference can be made to FIG12A, which shows that similarity is obtained based on the skeleton topology features of the source object and the skeleton topology features of the target object, and the motion features of the source object are fused based on the similarity to obtain the motion features of the target object.

[0225] In a possible implementation, the feature representation of the second skeleton node may be processed by a decoder to obtain the action data of the second skeleton node.

[0226] After obtaining the motion data, post-processing can be performed to obtain the final reconstructed animation. For example, dynamic constraints can be applied through forward kinematics (FK) or linear blending skinning (LBS) modules, combined with skinning for linear skinning, and optimization methods for threading and sliding can be performed through the post-optimization module.

[0227] In an embodiment of the present application, the similarity relationship between skeleton nodes of different skeleton topologies is determined based on the relationship between feature representations of different skeleton topologies, and action migration is performed based on the similarity relationship, thereby realizing action migration of different skeleton topologies.

[0228] For example, referring to Figures 12B and 12C, Figure 12B is a flow chart of the data processing method, and Figure 12C is a flow chart of the data processing method using graph information. Compared with Figure 12B, Figure 12C adds a motion whole-body integrated graph representation module and a topology whole-body integrated graph representation module.

[0229] 13 , which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in FIG13 , a data processing device 1300 provided in an embodiment of the present application includes:

[0230] An acquisition module 1301 is configured to acquire a first feature representation and a second feature representation, wherein the first feature representation is a feature representation of a skeleton topology of a first object, and the second feature representation is a feature representation of a skeleton topology of a second object, wherein the skeleton topology of the first object includes a plurality of first skeleton nodes, and the skeleton topology of the second object includes a second skeleton node;

[0231] The specific description of the acquisition module 1301 can refer to the introduction of step 801 in the above embodiment, which will not be repeated here.

[0232] A processing module 1302 is configured to obtain similarity information between the second skeleton node and each of the first skeleton nodes based on the first feature representation and the second feature representation;

[0233] fusing feature representations of motion data corresponding to a plurality of first skeleton nodes according to weights of the first skeleton nodes to obtain feature representations of motion data of a second skeleton node, wherein the weights are related to the similarity information;

[0234] The action data of the second skeleton node is obtained according to the feature representation of the second skeleton node.

[0235] The detailed description of the processing module 1302 may refer to the introduction of steps 802 to 804 in the above embodiment, which will not be repeated here.

[0236] In a possible implementation, there are multiple second skeleton nodes;

[0237] The processing module 1302 is specifically configured to:

[0238] Based on the first feature representation and the second feature representation, similarity information between each skeleton node included in the skeleton topology of the second object and each of the first skeleton nodes is obtained. Based on the weight of the first skeleton node, the feature representations of the action data corresponding to multiple first skeleton nodes are fused to obtain the feature representation of the action data of each skeleton node. Based on the feature representation of the action data of each skeleton node, the action data of each skeleton node is obtained.

[0239] In a possible implementation, the processing module 1302 is specifically configured to:

[0240] The similarity information is used as a weight to perform weighted summation on feature representations of the action data corresponding to the plurality of first skeleton nodes.

[0241] In a possible implementation, the processing module 1302 is specifically configured to:

[0242] Similarity information between the second skeleton node and each of the first skeleton nodes is obtained by performing a dot product and an activation operation on the first feature representation and the second feature representation.

[0243] In a possible implementation, the skeleton topology includes positions of skeleton nodes and connection relationships between skeleton nodes.

[0244] In a possible implementation, the motion data includes a rotation angle of a skeleton node.

[0245] In one possible implementation, the skeleton topology of the first object and the skeleton topology of the second object correspond to the same body pose.

[0246] In a possible implementation, the first feature representation is obtained by encoding the skeleton topology of the first object through a first encoder, and the second feature representation is obtained by encoding the skeleton topology of the second object through a second encoder; the skeleton topology of the first object and the skeleton topology of the second object are the same;

[0247] The processing module 1302 is further configured to:

[0248] The first encoder and the second encoder are updated according to the similarity information and the corresponding true value, wherein the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is greater than the similarity between the second skeleton node and the first skeleton node at a non-corresponding position.

[0249] In one possible implementation, the first encoder and the second encoder are graph neural networks.

[0250] In a possible implementation, the first object and the second object are people; the processing module 1302 is further configured to reconstruct a digital human corresponding to the second object based on the motion data of the second skeleton node.

[0251] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 14. Figure 14 is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1400 can be specifically manifested as a virtual reality VR device, a mobile phone, a tablet, a laptop computer, a smart wearable device, a monitoring data processing device or a server, etc., which is not limited here. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403 and a memory 1404 (wherein the number of processors 1403 in the execution device 1400 can be one or more, and Figure 14 takes one processor as an example), wherein the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403 and the memory 1404 may be connected via a bus or other means.

[0252] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0253] Processor 1403 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0254] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1403. Processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1403. The above processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1403 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in memory 1404. Processor 1403 reads information from memory 1404 and, in conjunction with its hardware, completes the steps involved in the model inference process in the above method.

[0255] Receiver 1401 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1402 can be used to output digital or character information through the first interface. Transmitter 1402 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1402 can also include a display device such as a display screen.

[0256] The present application also provides a training device. Please refer to FIG. 15 , which is a schematic diagram of the structure of a training device provided by an embodiment of the present application. Specifically, the training device 1500 is implemented by one or more servers. The training device 1500 may vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1515 (e.g., one or more processors), a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) storing application programs 1542 or data 1544. The memory 1532 and storage medium 1530 may be either ephemeral or persistent storage. The program stored in the storage medium 1530 may include one or more modules (not shown), each module including a series of instruction operations on the training device. Furthermore, the CPU 1515 may be configured to communicate with the storage medium 1530 to execute the series of instruction operations in the storage medium 1530 on the training device 1500.

[0257] The training device 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input and output interfaces 1558; or, one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0258] In the embodiment of the present application, the central processing unit 1515 is used to execute actions related to model training in the above embodiment.

[0259] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0260] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0261] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0262] Specifically, see Figure 16 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor (NPU) 1600. NPU 1600 is mounted on a host CPU (host CPU) as a coprocessor, with tasks assigned by the host CPU. The core of the NPU is arithmetic circuit 1603, which is controlled by controller 1604 to extract matrix data from memory and perform multiplication operations.

[0263] In some implementations, arithmetic circuit 1603 includes multiple processing units (PEs). In some implementations, arithmetic circuit 1603 is a two-dimensional systolic array. Arithmetic circuit 1603 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1603 is a general-purpose matrix processor.

[0264] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1602 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1601 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1608.

[0265] Unified memory 1606 is used to store input and output data. Weight data is directly transferred to weight memory 1602 through the Direct Memory Access Controller (DMAC) 1605. Input data is also transferred to unified memory 1606 through the DMAC.

[0266] BIU stands for Bus Interface Unit 1610 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1609 .

[0267] The bus interface unit 1610 (BIU) is used for the instruction fetch memory 1609 to obtain instructions from the external memory, and is also used for the storage unit access controller 1605 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0268] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1606 or move weight data to the weight memory 1602 or move input data to the input memory 1601.

[0269] The vector calculation unit 1607 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1603, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0270] In some implementations, vector calculation unit 1607 can store the processed output vector to unified memory 1606. For example, vector calculation unit 1607 can apply a linear function or a nonlinear function to the output of operation circuit 1603, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, vector calculation unit 1607 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to operation circuit 1603, for example, for use in subsequent layers in a neural network.

[0271] An instruction fetch buffer 1609 connected to the controller 1604 is used to store instructions used by the controller 1604;

[0272] Unified memory 1606, input memory 1601, weight memory 1602, and instruction fetch memory 1609 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0273] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0274] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0275] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0276] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0277] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. A data processing method, characterized in that: The method comprises: Acquire a first feature representation and a second feature representation, wherein the first feature representation is a feature representation of a skeleton topology of a first object, and the second feature representation is a feature representation of a skeleton topology of a second object, wherein the skeleton topology of the first object includes a plurality of first skeleton nodes, and the skeleton topology of the second object includes a second skeleton node; Obtaining similarity information between the second skeleton node and each of the first skeleton nodes according to the first feature representation and the second feature representation; According to the weights of the plurality of first skeleton nodes, the feature representations of the action data corresponding to the plurality of first skeleton nodes are fused to obtain the feature representation of the action data of the second skeleton node, wherein the weights are related to the similarity information; The action data of the second skeleton node is obtained according to the feature representation of the second skeleton node.

2. The method according to claim 1, characterized in that There are multiple second skeleton nodes; The method of obtaining similarity information between the second skeleton node and each of the first skeleton nodes according to the first feature representation and the second feature representation, fusing feature representations of action data corresponding to a plurality of first skeleton nodes according to weights of the first skeleton nodes to obtain a feature representation of the second skeleton node, and obtaining action data of the second skeleton node according to the feature representation of the second skeleton node includes: Based on the first feature representation and the second feature representation, the similarity information between each second skeleton node and each first skeleton node is obtained; based on the weight of the first skeleton node, the feature representations of the action data corresponding to multiple first skeleton nodes are fused to obtain the feature representation of the action data of each skeleton node; based on the feature representation of the action data of each skeleton node, the action data of each skeleton node is obtained.

3. The method according to claim 1 or 2, characterized in that: The step of fusing feature representations of action data corresponding to a plurality of first skeleton nodes according to the weights of the first skeleton nodes comprises: According to the weights of the plurality of first skeleton nodes, feature representations of the action data corresponding to the plurality of first skeleton nodes are weighted summed.

4. The method according to any one of claims 1 to 3, characterized in that: The obtaining, according to the first feature representation and the second feature representation, similarity information between the second skeleton node and each of the first skeleton nodes includes: By performing dot multiplication and activation operations on the first feature representation and the second feature representation, similarity information between the second skeleton node and each of the first skeleton nodes is obtained.

5. The method according to claim 4, characterized in that The skeleton topology includes the positions of skeleton nodes and the connection relationships between skeleton nodes.

6. The method according to any one of claims 1 to 5, characterized in that: The motion data includes the rotation angle of the skeleton node.

7. The method according to any one of claims 1 to 6, characterized in that: The skeletal topology of the first object and the skeletal topology of the second object correspond to a same body pose.

8. The method according to any one of claims 1 to 7, characterized in that: The first feature representation is obtained by encoding the skeleton topology of the first object through a first encoder, and the second feature representation is obtained by encoding the skeleton topology of the second object through a second encoder; the skeleton topology of the first object and the skeleton topology of the second object are the same; The method further comprises: The first encoder and the second encoder are updated according to the similarity information and the corresponding true value, wherein the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is greater than the similarity between the second skeleton node and the first skeleton node at the non-corresponding position.

9. The method according to claim 8, characterized in that The first encoder and the second encoder are graph neural networks.

10. The method according to any one of claims 1 to 9, characterized in that: The first object and the second object are characters; the method further includes: Reconstruct a digital human corresponding to the second object according to the action data of the second skeleton node.

11. A data processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a first feature representation and a second feature representation, wherein the first feature representation is a feature representation of a skeleton topology of a first object, and the second feature representation is a feature representation of a skeleton topology of a second object, wherein the skeleton topology of the first object includes a plurality of first skeleton nodes, and the skeleton topology of the second object includes a second skeleton node; A processing module, configured to obtain similarity information between the second skeleton node and each of the first skeleton nodes according to the first feature representation and the second feature representation; According to the weight of the first skeleton node, the feature representations of the action data corresponding to the plurality of first skeleton nodes are fused to obtain the feature representation of the action data of the second skeleton node, wherein the weight is related to the similarity information; The action data of the second skeleton node is obtained according to the feature representation of the second skeleton node.

12. The device according to claim 11, characterized in that There are multiple second skeleton nodes; The processing module is specifically used for: According to the first feature representation and the second feature representation, obtaining similarity information between each second skeleton node and each first skeleton node; according to the weight of the first skeleton node, fusing the feature representations of the action data corresponding to the plurality of first skeleton nodes to obtain the feature representation of the action data of each skeleton node; The action data of each skeleton node is obtained according to the feature representation of the action data of each skeleton node.

13. The device according to claim 11 or 12, characterized in that The processing module is specifically used for: According to the weights of the plurality of first skeleton nodes, feature representations of the action data corresponding to the plurality of first skeleton nodes are weighted summed.

14. The device according to any one of claims 11 to 13, characterized in that: The processing module is specifically used for: By performing dot multiplication and activation operations on the first feature representation and the second feature representation, similarity information between the second skeleton node and each of the first skeleton nodes is obtained.

15. The device according to claim 14, characterized in that The skeleton topology includes the positions of skeleton nodes and the connection relationships between skeleton nodes.

16. The device according to any one of claims 11 to 15, characterized in that The motion data includes the rotation angle of the skeleton node.

17. The device according to any one of claims 11 to 16, characterized in that The skeletal topology of the first object and the skeletal topology of the second object correspond to a same body pose.

18. The device according to any one of claims 11 to 17, characterized in that The first feature representation is obtained by encoding the skeleton topology of the first object through a first encoder, and the second feature representation is obtained by encoding the skeleton topology of the second object through a second encoder; the skeleton topology of the first object and the skeleton topology of the second object are the same; The processing module is further used for: The first encoder and the second encoder are updated according to the similarity information and the corresponding true value, wherein the true value indicates that the similarity between the second skeleton node and the first skeleton node at the corresponding position is greater than the similarity between the second skeleton node and the first skeleton node at the non-corresponding position.

19. The device according to claim 18, characterized in that The first encoder and the second encoder are graph neural networks.

20. The device according to any one of claims 11 to 19, characterized in that The first object and the second object are people; the processing module is further used for: Reconstruct a digital human corresponding to the second object according to the action data of the second skeleton node.

21. A computer storage medium, characterized in that The computer storage medium stores one or more instructions which, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 10.

22. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 10.

23. A system comprising at least one processor and at least one memory; the processor and the memory are connected via a communication bus and communicate with each other; The at least one memory is used to store code; The at least one processor is configured to execute the code to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Pairing-free action migration method based on deep convolutional network

    CN112164129A

  • Animation migration method and device, equipment and storage medium

    CN113313794A

  • Skeleton mapping method and device, equipment and storage medium

    CN113592987A

  • Motion redirection method and device, electronic equipment and computer readable storage medium

    CN116977502A

  • Data processing method and device

    CN117689778A