Characteristic extraction network training method and device, characteristic extraction network retrieval method and device, equipment, medium and program

By performing multimodal cross-training in action animation retrieval, the obtained target feature extraction network can realize accurate retrieval of text and actions without rendering the action file as video animation, solving the problem of large resource overhead and the need to render video in the prior art, and improving the efficiency and accuracy of the retrieval.

CN119941936APending Publication Date: 2025-05-06NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997304.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has a high resource overhead in action animation retrieval, and it is necessary to render the action file into a video animation for retrieval, resulting in high storage and calculation consumption.

Method used

By obtaining multiple sample action animation data and text annotation information, a preset action feature extraction module and text feature extraction module are used for feature extraction, and multimodal cross-training is performed by combining sample text features and action features to obtain a target feature extraction network. This network can realize both text retrieval and action retrieval without rendering action files as video animations.

Benefits of technology

It reduces the storage and calculation consumption of action animation retrieval, improves the accuracy and efficiency of retrieval, and can realize accurate text and actions without rendering the action file as video animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941936A_ABST
    Figure CN119941936A_ABST
Patent Text Reader

Abstract

The invention provides a feature extraction network training method and device, a feature extraction network retrieval method and device, equipment, a medium and a program, and relates to the technical field of games. According to the method, multiple pieces of sample action animation data are obtained, each piece of sample action animation data has corresponding text labeling information, and the text labeling information is used for describing action attributes of the sample action animation data; respectively adopting a preset action feature extraction module and a preset text feature extraction module to perform feature extraction on each piece of sample action animation data and the corresponding text labeling information to obtain a sample action feature and a sample text feature corresponding to each piece of sample action animation data; and performing parameter adjustment on a preset text feature extraction module and a preset action feature extraction module according to the sample text features and the sample action features corresponding to the multiple pieces of sample action animation data to obtain a target feature extraction network comprising a target text feature extraction module and a target action feature extraction module. Therefore, storage and calculation consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of game technology, and in particular to a feature extraction network training, retrieval method, device, equipment, medium and program. Background Art

[0002] In the fields of games, 3D movies, animation, etc., the production of character body motion animation (hereinafter referred to as "animation") is a very important link. Character animation is usually composed of three-dimensional rotation and displacement of each skeletal joint of a 3D character. When making character animation, creators can choose to make it frame by frame from the initial posture of the character, but this usually takes a lot of time. You can also choose to process or use some existing motion files as the basis. The latter can greatly shorten the production time and is also the most commonly used solution. Existing motion files can come from motion capture, or be screened from an existing motion library.

[0003] Among them, screening from the action library is also called "action retrieval". This is a technology that retrieves similar action animations from the database based on a given action animation (hereinafter referred to as "animation"). It can help users obtain a large amount of action animation data of the same type in a short time, so that users can quickly filter out the animation that is most similar to the specified animation; previous retrieval was mostly through single-modal retrieval from action to action, requiring users to first have a target action to be retrieved. However, in many cases, users cannot provide a ready-made target action animation in advance, which makes it difficult to meet user needs and greatly limits the application scenarios. In the past, action retrieval usually required rendering the action file into an animation video first, and then searching based on the video, which was resource-intensive and time-consuming. Summary of the invention

[0004] In view of this, the embodiments of the present application provide a feature extraction network training, retrieval method, device, equipment, medium and program to solve problems such as large resource overhead.

[0005] In a first aspect, an embodiment of the present application provides a method for training a feature extraction network, the method comprising:

[0006] Acquire multiple pieces of sample action animation data, each piece of sample action animation data having corresponding text annotation information, wherein the text annotation information is used to describe action attributes of the sample action animation data;

[0007] Using a preset action feature extraction module and a preset text feature extraction module respectively, extracting features from each piece of sample action animation data and the corresponding text annotation information, and obtaining a sample action feature and a sample text feature corresponding to each piece of sample action animation data;

[0008] According to the sample text features and sample action features corresponding to the multiple sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

[0009] In a second aspect, an embodiment of the present application provides an action animation retrieval method, the method comprising:

[0010] Get the input text action description information;

[0011] Using a target text feature extraction module in a target feature extraction network to extract features from the text action description information, and obtain target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the method described in any one of the first aspects above;

[0012] According to the target text feature, first target action animation data matching the target text feature is determined from a preset action animation database.

[0013] In a third aspect, an embodiment of the present application provides a feature extraction network training device, the device comprising:

[0014] A first acquisition module is used to acquire a plurality of sample action animation data, each of which has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data;

[0015] A first extraction module is used to respectively use a preset action feature extraction module and a preset text feature extraction module to perform feature extraction on each piece of sample action animation data and the corresponding text annotation information to obtain a sample action feature and a sample text feature corresponding to each piece of sample action animation data;

[0016] The training module is used to adjust the parameters of the preset text feature extraction module and the preset action feature extraction module according to the sample text features and sample action features corresponding to the multiple sample action animation data, so as to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

[0017] In a fourth aspect, an embodiment of the present application provides an action animation retrieval device, the device comprising:

[0018] The second acquisition module is used to acquire input text action description information;

[0019] A second extraction module is used to use a target text feature extraction module in a target feature extraction network to perform feature extraction on the text action description information to obtain target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the method described in any one of the first aspects above;

[0020] The determination module is used to determine, according to the target text feature, first target action animation data matching the target text feature from a preset action animation database.

[0021] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the feature extraction network training method as described in any one of the first aspects, or the steps of the action animation retrieval method as described in any one of the second aspects.

[0022] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the feature extraction network training method as described in any one of the first aspect, or the steps of the action animation retrieval method as described in any one of the second aspect are executed.

[0023] In a seventh aspect, an embodiment of the present application provides a computer program instruction, wherein the program instruction is executed by a processor to perform the steps of the feature extraction network training method as described in any one of the first aspect, or the steps of the action animation retrieval method as described in any one of the second aspect.

[0024] Compared with the prior art, this application has the following beneficial effects:

[0025] The present application provides a method, device, equipment, medium and program for training and retrieving a feature extraction network. The method obtains multiple sample action animation data, each of which has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data; a preset action feature extraction module and a preset text feature extraction module are respectively used to extract features from each sample action animation data and the corresponding text annotation information, so as to obtain sample action features and sample text features corresponding to each sample action animation data; according to the sample text features and sample action features corresponding to the multiple sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module. Thus, by combining the sample text features and the sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, so as to obtain a more accurate target feature extraction network, which can realize both text retrieval and action retrieval, without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0027] Figure 1 A flow chart of a feature extraction network training method provided in an embodiment of the present application;

[0028] Figure 2 A flowchart of a method for obtaining a target feature extraction network provided in an embodiment of the present application;

[0029] Figure 3 A schematic diagram of a flow chart of a method for calculating a first loss function value provided in an embodiment of the present application;

[0030] Figure 4 A flowchart of a method for calculating a target loss function value provided in an embodiment of the present application;

[0031] Figure 5 A flowchart of another method for calculating the target loss function value provided in an embodiment of the present application;

[0032] Figure 6 A flowchart of a method for calculating a second loss function value provided in an embodiment of the present application;

[0033] Figure 7 A flowchart of a method for calculating a third loss function value provided in an embodiment of the present application;

[0034] Figure 8 A flowchart of a method for calculating a loss function value for each action style provided in an embodiment of the present application;

[0035] Fig. 9 A model training architecture diagram provided for an embodiment of the present application;

[0036] Fig.10 A flowchart of an action animation retrieval method provided in an embodiment of the present application;

[0037] Fig.11 A schematic diagram of a flow chart of another action animation retrieval method provided in an embodiment of the present application;

[0038] Fig.12 A flowchart of another action animation retrieval method provided in an embodiment of the present application;

[0039] Fig.13 A schematic diagram of a feature extraction network training device provided in an embodiment of the present application;

[0040] Fig.14 A schematic diagram of an action animation retrieval device provided in an embodiment of the present application;

[0041] Fig.15 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0042] Icons: 1301 - first acquisition module, 1302 - first extraction module, 1303 - training module, 1401 - second acquisition module, 1402 - second extraction module, 1403 - determination module, 1501 - processor, 1502 - storage medium, 1503 - bus. DETAILED DESCRIPTION

[0043] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0044] In one embodiment of the present application, a feature extraction network training method can be run on a local terminal device or a server. When a feature extraction network training method is run on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.

[0045] Figure 1 A flowchart of a feature extraction network training method provided in an embodiment of the present application.

[0046] like Figure 1 As shown, the method includes:

[0047] S101. Acquire multiple pieces of sample action animation data, each piece of sample action animation data having corresponding text annotation information.

[0048] The text annotation information is used to describe the action attributes of the sample action animation data.

[0049] By annotating the action animation data with text annotation information, the action attributes of the action animation data can be known through the text annotation information. The action attribute includes an attribute name and an attribute value. The attribute name and the attribute value together describe the characteristics of the action, and the attribute value determines the specific value or specific state of the action attribute. The action animation data is composed of at least one action attribute. That is, the text annotation information can be used to know what kind of action the character in the action animation is doing in what state, so that the action description is more accurate.

[0050] For example, in order to facilitate network training, it is necessary to extract data feature formats suitable for neural network training from action animation files. This application adopts action features in the commonly used invariant features format, which has a total of 626 dimensions; assuming that the human skeleton in the action file mainly includes the following components: root node features, body joint features, and feet touching the ground labels. Root node features: vertical height of the root node, linear velocity in the xz plane, angular velocity of rotation around the y axis, and 6D rotation features. Body joint features: 6D rotation, local position coordinates, and local linear velocity of body joints. Feet touching the ground label: a label for whether the feet are touching the ground or not (1 if touching the ground, otherwise 0). The above features can be used to accurately obtain the action style of the action animation.

[0051] First, the action animation files are identified to obtain the action animation data. In order to increase the amount of training data, the action animation data is enhanced by action mirroring, which doubles the amount of data.

[0052] For example, the text annotation information may be: "A person is running forward; the gender is male; the age is adult; the emotion is happy", "A person raises his hands and shakes them from side to side; the gender is female; the age is young; the emotion is happy".

[0053] S102 , respectively using a preset action feature extraction module and a preset text feature extraction module to extract features from each piece of sample action animation data and the corresponding text annotation information, to obtain a sample action feature and a sample text feature corresponding to each piece of sample action animation data.

[0054] The feature extraction network provided in this application includes: an action feature extraction module and a text feature extraction module.

[0055] For example, the action feature extraction module and the text feature extraction module are both multi-head action feature extraction modules. The action feature extraction module and the text feature extraction module can be the same network model or different network models, as long as the feature extraction function can be achieved.

[0056] S103. According to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

[0057] Therefore, by combining sample text features and sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, and obtain a more accurate target feature extraction network. The target feature extraction network can realize both text retrieval and action retrieval, without rendering action files as video animations, thereby greatly reducing storage and computing consumption.

[0058] In summary, in this embodiment, a plurality of sample action animation data are obtained, each of which has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data; a preset action feature extraction module and a preset text feature extraction module are respectively used to extract features from each sample action animation data and the corresponding text annotation information, and obtain sample action features and sample text features corresponding to each sample action animation data; according to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module. Thus, by combining the sample text features and the sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, and obtain a more accurate target feature extraction network, which can realize both text retrieval and action retrieval, without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0059] Based on the above embodiments, the embodiments of the present application also provide a method for obtaining a target feature extraction network. Figure 2 A flow chart of a method for obtaining a target feature extraction network provided in an embodiment of the present application. Figure 2 As shown, in S103, according to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module, including:

[0060] S201. Calculate a first loss function value according to sample text features and sample action features corresponding to a plurality of sample action animation data.

[0061] For model training, the purpose of calculating the first loss function value is to make the sample text features and sample action features corresponding to the same action animation data as similar as possible or have a high degree of similarity, and to make the sample text features of the action animation data and the sample action features of other action animation data have a low degree of similarity. If this goal can be achieved, it means that the features extracted by the preset text feature extraction module and the preset action feature extraction module are the same, which meets the model training requirements.

[0062] For example, a similarity matrix can be formed in sequence by the sample text features and sample action features corresponding to multiple sample action animation data. Each row of the similarity matrix corresponds to a similar combination of a sample action feature and all sample text features, and each column corresponds to a similar combination of a sample text feature and all sample action features. The diagonal position of the similarity matrix is ​​the similar combination of sample text features and sample action features corresponding to the same action animation data, and the other positions are the similar combinations of sample text features and sample action features corresponding to different action animation data.

[0063] S202. Calculate a target loss function value according to the first loss function value.

[0064] For example, the first loss function value may be directly determined as the target loss function value.

[0065] S203. According to the target loss function value, adjust the parameters of the preset text feature extraction module and the preset action feature extraction module to obtain a target feature extraction network.

[0066] A preset training cutoff condition is set in advance, for example, the preset training cutoff condition is: the target loss function value is less than or equal to the preset threshold. If the calculated target loss function value is less than or equal to the preset threshold, the training is stopped to obtain the target feature extraction network. If the calculated target loss function value is greater than the preset threshold, the preset text feature extraction module and the preset action feature extraction module are continuously adjusted until the target loss function value is less than or equal to the preset threshold, and the target feature extraction network is obtained.

[0067] For example, the Adam optimizer can be used to set the data batch size to 32 and the initial learning rate to 0.0001 for training.

[0068] Therefore, by combining sample text features and sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, and obtain a more accurate target feature extraction network.

[0069] In summary, in this embodiment, the first loss function value is calculated according to the sample text features and sample action features corresponding to the plurality of sample action animation data; the target loss function value is calculated according to the first loss function value; and the preset text feature extraction module and the preset action feature extraction module are adjusted according to the target loss function value to obtain the target feature extraction network. Thus, by combining the sample text features and the sample action features for multimodal cross-training, a more accurate target feature extraction network is obtained.

[0070] Based on the above embodiments, the embodiments of the present application also provide a method for calculating the value of a first loss function. Figure 3 A flow chart of a method for calculating a first loss function value provided in an embodiment of the present application. Figure 3 As shown, in S201, the first loss function value is calculated according to the sample text features and sample action features corresponding to the plurality of sample action animation data, including:

[0071] S301. Calculate first action loss function values ​​of multiple pieces of sample action animation data according to sample text features and sample action features corresponding to multiple pieces of sample action animation data.

[0072] The first action loss function value is the loss function value of the sample action feature relative to the sample text feature.

[0073] For example, the first loss function value can be calculated using InfoNCE loss, and the loss function values ​​of the action and content are calculated, so that the first loss function value is more accurate.

[0074] Specifically, the calculation method of the first action loss function value is shown in the following formula (1):

[0075]

[0076] in, is the first action loss function value, N is the number of samples, sim(.) is the cosine similarity calculation (i.e., the dot product of two feature vectors), f i cM is the sample action feature of the i-th sample, f j cT is the sample text feature of the jth sample.

[0077] S302: Calculate first text loss function values ​​of the plurality of sample action animation data according to sample text features and sample action features corresponding to the plurality of sample action animation data.

[0078] The first text loss function value is the loss function value of the sample text feature relative to the sample action feature.

[0079] Specifically, the calculation method of the first text loss function value is shown in the following formula (2):

[0080]

[0081] Among them, f i cT is the sample text feature of the i-th sample, f j cM is the sample action feature of the jth sample.

[0082] S303. Calculate a first loss function value according to the first action loss function value and the first text loss function value.

[0083] Specifically, the calculation method of the first loss function value is shown in the following formula (3):

[0084]

[0085] Among them, L_concate is the first loss function value.

[0086] In summary, in this embodiment, the first action loss function values ​​of the plurality of sample action animation data are calculated according to the sample text features and sample action features corresponding to the plurality of sample action animation data; the first text loss function values ​​of the plurality of sample action animation data are calculated according to the sample text features and sample action features corresponding to the plurality of sample action animation data; and the first loss function value is calculated according to the first action loss function value and the first text loss function value. Thus, the first loss function value is accurately calculated.

[0087] Based on the above embodiments, the embodiments of the present application also provide a method for calculating the target loss function value. Figure 4 A flow chart of a method for calculating a target loss function value provided in an embodiment of the present application. Figure 4 As shown, the text annotation information includes: action content annotation information, the sample action features include: first action content information, and the sample text features include: second action content information.

[0088] Before calculating the target loss function value according to the first loss function value in S202, the method further includes:

[0089] S401. Calculate a second loss function value according to first action content information and second action content information corresponding to a plurality of sample action animation data.

[0090] For example, the action content annotation information can be: "A game character is running forward", "A game character is swinging left and right with his hands raised high". By introducing action content, the characteristics of the action are enriched, the model training accuracy is improved, and the model has the ability to retrieve action content.

[0091] The first action style information represents action content information extracted from action animation data; the second action style information represents action content information extracted from text annotation information.

[0092] For model training, the purpose of calculating the second loss function value is to make the first action content information and the second action content information corresponding to the same action animation data as identical as possible or have a high degree of similarity, and to make the first action content information and the second action content information corresponding to different action animation data have a low degree of similarity. If this purpose can be achieved, it means that the features extracted by the preset text feature extraction module and the preset action feature extraction module are the same, which meets the model training requirements.

[0093] For example, a content similarity matrix may be formed in sequence by first action content information and second action content information corresponding to a plurality of sample action animation data. Each row of the content similarity matrix corresponds to a similar combination of a first action content information and all second action content information, and each column corresponds to a similar combination of a second action content information and all first action content information. The diagonal positions of the similarity matrix are similar combinations of first action content information and second action content information corresponding to the same action animation data, and other positions are similar combinations of first action content information and second action content information corresponding to different action animation data.

[0094] Calculating the target loss function value according to the first loss function value in S202 includes:

[0095] S402. Calculate a target loss function value according to the first loss function value and the second loss function value.

[0096] For example, the first loss function value and the second loss function value may be added to obtain a target loss function value.

[0097] In summary, in this embodiment, the text annotation information includes: action content annotation information, the sample action features include: first action content information, and the sample text features include: second action content information; the second loss function value is calculated based on the first action content information and the second action content information corresponding to the plurality of sample action animation data; the target loss function value is calculated based on the first loss function value and the second loss function value. Thus, the target loss function value is accurately calculated.

[0098] Based on the above embodiments, the embodiments of the present application also provide another method for calculating the target loss function value. Figure 5 A flow chart of another method for calculating the target loss function value provided in an embodiment of the present application. Figure 5 As shown, the text annotation information also includes: action style annotation information; the sample action feature also includes: first action style information, and the sample text feature also includes: second action style information.

[0099] Before calculating the target loss function value according to the first loss function value in S202, the method further includes:

[0100] S501. Calculate a third loss function value according to first action style information and second action style information corresponding to a plurality of sample action animation data.

[0101] For example, action style annotation information can be: gender, age, clarity and other information. By introducing action style, the characteristics of the action are enriched, the model training accuracy is improved, and the model has the ability to retrieve action styles. For example, gender includes: male, female; age includes: childhood, youth, adult, old age; emotions include: happy, angry, sad, afraid, confused, shy, not obvious.

[0102] The first action style information represents the action style information extracted from the action animation data; the second action style information represents the action style information extracted from the text annotation information.

[0103] For model training, the purpose of calculating the third loss function value is to make the first action style information and the second action style information corresponding to the same action animation data as identical as possible or have a high degree of similarity. If this goal can be achieved, it means that the features extracted by the preset text feature extraction module and the preset action feature extraction module are identical, which meets the model training requirements.

[0104] For example, a style similarity matrix may be formed in sequence by first action style information and second action style information corresponding to a plurality of sample action animation data. Each row of the style similarity matrix corresponds to a similar combination of a first action style information and all second action style information, and each column corresponds to a similar combination of a second action style information and all first action style information. The diagonal positions of the style similarity matrix are similar combinations of the first action style information and the second action style information corresponding to the same action animation data, and the other positions are similar combinations of the first action style information and the second action style information corresponding to different action animation data.

[0105] It should be noted that if in the same row or column of a similar combination of the first action style information and the second action style information corresponding to the same action animation data, there exists a similar combination with the same style as the first action style information and the second action style information corresponding to the action animation data, then the same similar combination should also be made as identical as possible or as highly similar as possible.

[0106] Calculating the target loss function value according to the first loss function value in S202 includes:

[0107] S502. Calculate a target loss function value according to the first loss function value and the third loss function value.

[0108] Specifically, the target loss function value is calculated as shown in the following formula (4):

[0109] Loss = α·L style +β·L concate (4)

[0110] Among them, L concate is the first loss function value, L style is the third loss function value, α and β are the weights of the third loss function value and the first loss function value respectively.

[0111] In summary, in this embodiment, the text annotation information also includes: action style annotation information; the sample action feature also includes: first action style information, and the sample text feature also includes: second action style information; the third loss function value is calculated according to the first action style information and the second action style information corresponding to the plurality of sample action animation data; the target loss function value is calculated according to the first loss function value and the third loss function value. Thus, the target loss function value is accurately calculated.

[0112] Furthermore, based on the above embodiments, the embodiments of the present application also provide a method for calculating a target loss function value based on a first loss function value, a second loss value and a third loss function value.

[0113] The specific calculation method is shown in the following formula (5):

[0114] Loss = L content +α·L style +β·L conccate (5)

[0115] Among them, L content is the second loss value.

[0116] Based on the above embodiments, the embodiments of the present application also provide a method for calculating the value of a second loss function. Figure 6 A flow chart of a method for calculating a second loss function value provided in an embodiment of the present application. Figure 6 As shown, in S401, according to the first action content information and the second action content information corresponding to the plurality of sample action animation data, calculating the second loss function value includes:

[0117] S601. Calculate second action loss function values ​​of multiple pieces of sample action animation data according to first action content information and second action content information corresponding to multiple pieces of sample action animation data.

[0118] The second action loss function value is a loss function value of the first action content information relative to the second action content information.

[0119] For example, the first action content information and the second action content information are respectively a first action content feature and a second action content feature.

[0120] Specifically, the calculation method of the second action loss function value is shown in the following formula (6):

[0121]

[0122] Among them, L m is the second action loss function value, f i M is the first action content feature of the i-th sample, is the second action content feature of the jth sample.

[0123] S602: Calculate second text loss function values ​​of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data.

[0124] The second text loss function value is a loss function value of the second action content information relative to the first action content information.

[0125] Specifically, the second text loss function value is calculated as shown in the following formula (7):

[0126]

[0127] Among them, L t is the second text loss function value, f i T is the second action content feature of the i-th sample, is the first action content feature of the jth sample.

[0128] S603. Calculate a second loss function value according to the second action loss function value and the second text loss function value.

[0129] Specifically, the calculation method of the second loss function value is shown in the following formula:

[0130]

[0131] In summary, in this embodiment, the second action loss function values ​​of the plurality of sample action animation data are calculated according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second text loss function values ​​of the plurality of sample action animation data are calculated according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; and the second loss function value is calculated according to the second action loss function value and the second text loss function value. Thus, the second loss function value is accurately calculated.

[0132] Based on the above embodiments, the embodiments of the present application also provide a method for calculating the value of a third loss function. Figure 7 A flow chart of a method for calculating the third loss function value provided in an embodiment of the present application. Figure 7 As shown, the action style annotation information includes: annotation information of at least one action style, and correspondingly, the first action style information includes: a first action style feature of at least one action style, and the second action style information includes: a second action style feature of at least one action style;

[0133] Calculating a third loss function value according to the first action style information and the second action style information corresponding to the plurality of sample action animation data in S501 includes:

[0134] S701 : Calculate a loss function value of each action style according to a first action style feature and a second action style feature of each action style corresponding to a plurality of sample action animation data.

[0135] In order to accurately calculate the third loss function value, the loss function value of each action style is calculated separately first.

[0136] S702: Calculate a third loss function value according to the loss function value of at least one action style.

[0137] For example, the loss function value of at least one action style is added to obtain a third loss function value.

[0138] In summary, in this embodiment, the action style annotation information includes: at least one action style annotation information, and correspondingly, the first action style information includes: at least one first action style feature of the action style, and the second action style information includes: at least one second action style feature of the action style; according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data, the loss function value of each action style is calculated; according to the loss function value of at least one action style, the third loss function value is calculated. Thus, the target loss function value is accurately calculated.

[0139] Based on the above embodiment, the embodiment of the present application also provides a method for calculating the loss function value of each action style. Figure 8 A flowchart of a method for calculating the loss function value of each action style provided in an embodiment of the present application. Figure 8 As shown, in S701, according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data, the loss function value of each action style is calculated, including:

[0140] S801. Calculate a third action loss function value for each action style according to first action style features and second action style features of each action style corresponding to a plurality of sample action animation data.

[0141] The third action loss function value is the loss function value of the first action style feature relative to the second action style feature.

[0142] For example, the third action loss function may adopt binary cross entropy loss (BCE Loss).

[0143] Specifically, the calculation method of the third action loss function value is shown in the following formula (8):

[0144]

[0145] in, is the loss function value of the third action, y ij is the true value of the i-th row and j-th column of the style similarity matrix (0 or 1, if the first action style feature of the i-th row and the second action style feature of the j-th column are the same, then it is 1, if they are different, then it is 0). ij is the predicted value of the i-th row and j-th column of the style similarity matrix, that is, p ij is the similarity between the first action style feature in the i-th row and the second action style feature in the j-th column.

[0146] S802: Calculate a third text loss function value for each action style according to a first action style feature and a second action style feature of each action style corresponding to a plurality of sample action animation data.

[0147] The third text loss function value is the loss function value of the second action style feature relative to the first action style feature.

[0148] Specifically, the calculation method of the third text loss function value is shown in the following formula (9):

[0149]

[0150] in, is the third text loss function value, y ji is the true value of the jth row and ith column of the style similarity matrix. ji is the predicted value of the jth row and ith column of the style similarity matrix, that is, p ji is the similarity between the first action style feature in the jth row and the second action style feature in the i-th column.

[0151] S803. Calculate the loss function value of each action style according to the third action loss function value and the third text loss function value of each action style.

[0152] Specifically, the loss function value of each action style is calculated as shown in the following formula (10):

[0153]

[0154] In summary, in this embodiment, the third action loss function value of each action style is calculated according to the first action style feature and the second action style feature of each action style corresponding to multiple sample action animation data; the third action loss function value is the loss function value of the first action style feature relative to the second action style feature; the third text loss function value of each action style is calculated according to the first action style feature and the second action style feature of each action style corresponding to multiple sample action animation data; the third text loss function value is the loss function value of the second action style feature relative to the first action style feature; the loss function value of each action style is calculated according to the third action loss function value and the third text loss function value of each action style. Thus, the loss function value of each action style is accurately calculated.

[0155] On the basis of the above-mentioned embodiment, in another embodiment of the present application, in addition to directly extracting sample text features and sample action features. The first action content information, the second action content information, the first action style information and the second style content information corresponding to the sample action animation data may also be extracted first. The second loss function value is calculated according to the first action content information and the second action content information, and the third loss function value is calculated according to the first action style information and the second style content information. The first action content information and the first action style information are then spliced ​​to obtain the sample action features, and the second action content information and the second style content information are spliced ​​to obtain the sample text features. The first loss function value is calculated according to the sample text features and the sample action features. The specific splicing method may be feature vector addition, which will not be described in detail here.

[0156] Based on the above embodiments, Fig. 9 A model training architecture diagram provided for an embodiment of the present application.

[0157] like Fig. 9 As shown, under this architecture, the feature extraction network model includes: an action feature extraction module and a text feature extraction module.

[0158] The action feature extraction module consists of a multi-layer Transformer encoding layer, whose input is the original action data (626-dimensional invariant features feature vector), and the output contains four heads, namely action content, action gender style, action age style, and action emotion style. Similarly, the text feature extraction module is a text encoder of a pre-trained CLIP model, whose input is four text data, namely action content description text, gender description text, age description text, and emotion description text; the output also contains four heads, corresponding to the input text content and three style features.

[0159] The multi-head action feature extractor consists of a general feature extractor and four multi-head feature extractors; each part is a Transformer layer network. Given an action data x m , and its corresponding text label is Among them, x m It is the key information of animation skeleton extracted from the animation file, including the rotation of each joint bone of the character and the displacement information of the root node; the content and style information of the action are coupled in the rotation and displacement information of these joints, and the multi-head action feature extractor is responsible for decoupling the action content and style information respectively.

[0160] The motion feature extraction module first uses a common encoder M(x m ) Extract general action features f g :

[0161]

[0162] Then will The four input heads extract 512-dimensional action content and three style features of gender, age, and emotion.

[0163]

[0164] Multi-head text feature extraction module. In order to obtain various features under the text modality, a multi-head text feature extraction module is also designed, which includes a general text feature extractor T and four multi-head feature extractors. The general text feature extraction part is composed of a text encoder of a pre-trained CLIP model, denoted as T(.). The encoder loads the CLIP pre-training parameters for initialization and updates the network parameters during the training process. The input of the feature extraction module is four types of text annotation information, namely action content text and three styles of text. The final content and gender, age, and emotion text features are extracted by T(.) and the multi-head text feature extractor respectively:

[0165]

[0166] In addition to the feature extraction network model, this architecture also includes a contrastive learning module, the underlying layer of which is implemented through mathematical calculations (calculating the loss function value).

[0167] After the above feature extraction module, the content and three styles (gender, age, emotion) of the two modes of action and text are obtained respectively; in order to support the retrieval of action content and action style respectively, the contrastive learning module is divided into two parts, namely, action content contrastive learning and action style contrastive learning. On this basis, the content in the action mode and the features of the three styles are combined to obtain action features, the content in the text mode and the features of the three styles are combined to obtain text features, and the action features and text features are combined for contrastive learning.

[0168] Based on the above embodiments, the embodiments of the present application also provide an action animation retrieval method. Fig.10 This is a flow chart of an action animation retrieval method provided in an embodiment of the present application. Fig.10 As shown, the method includes:

[0169] S901. Obtain input text action description information.

[0170] Users enter text action description information based on their own search needs.

[0171] S902: Using the target text feature extraction module in the target feature extraction network, extract features from the text action description information to obtain target text features corresponding to the text action description information.

[0172] Among them, the target feature extraction network is a network model trained by using any method of the above embodiments.

[0173] S903: Determine, according to the target text feature, first target action animation data matching the target text feature from a preset action animation database.

[0174] The network model pre-trained by using any of the methods in the above-mentioned embodiments in the preset action animation database stores a plurality of action animations and action features of the action animations.

[0175] The action animation with the highest similarity between the action feature and the target text feature is used to determine the first target action animation data matching the target text feature.

[0176] In summary, in this embodiment, the input text action description information is obtained; the target text feature extraction module in the target feature extraction network is used to extract features from the text action description information to obtain the target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained using the above method; according to the target text features, the first target action animation data matching the target text features is determined from a preset action animation database. Thus, by combining sample text features and sample action features for multimodal cross-training, a more accurate target feature extraction network is obtained, which can realize both text retrieval and action retrieval without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0177] Based on the above embodiment, the embodiment of the present application also provides another action animation retrieval method. Fig.11 A flowchart of another method for retrieving action animations provided in an embodiment of the present application is shown below. Fig.11 As shown, the text action description information includes: action content information and / or action style information.

[0178] In S902, the target text feature extraction module in the target feature extraction network is used to extract features from the text action description information to obtain target text features corresponding to the text action description information, including:

[0179] S1001. Using a target text feature extraction module, extract features from action content information and / or action style information to obtain target action content features and / or target action style features.

[0180] The target text feature extraction module can extract action content information and / or action style information from the text action description information.

[0181] In S903, according to the target text feature, determining first target action animation data matching the target text feature from a preset action animation database includes:

[0182] S1002: Determine, from a preset action animation database, first target action animation data matching the target action content feature and / or the target action style feature according to the target action content feature and / or the target action style feature.

[0183] The action features of the action animation include: action content features and / or action style features. The action content features and / or action style features of the action animation are compared with the action animation with the highest similarity to the target action content features and / or target action style features to determine the first target action animation data matching the target action content features and / or target action style features.

[0184] Therefore, the use of action content information and / or action style information for animation retrieval improves retrieval accuracy and facilitates user use.

[0185] In summary, in this embodiment, a target text feature extraction module is used to extract features of action content information and / or action style information to obtain target action content features and / or target action style features; based on the target action content features and / or target action style features, first target action animation data matching the target action content features and / or target action style features is determined from a preset action animation database. Thus, the use of action content information and / or action style information for animation retrieval improves retrieval accuracy and is convenient for users.

[0186] Based on the above embodiments, the embodiments of the present application also provide another action animation retrieval method. Fig.12 A flowchart of another method for retrieving action animations provided in an embodiment of the present application is shown below. Fig.12 As shown, the method also includes:

[0187] S1101. Obtain input given action animation data.

[0188] In the action animation retrieval method provided by the present application, in addition to text retrieval, action animation retrieval can also be performed, wherein the given action animation data is an image or video after rendering the action animation, so that animation retrieval can be performed based on the rendered image or video.

[0189] S1102, using the target action feature extraction module in the target feature extraction network to perform feature extraction on given action animation data to obtain target action features corresponding to the given action animation data.

[0190] Among them, the target action feature includes the content information and style information of the action.

[0191] S1103: Determine, according to the target action feature, second target action animation data matching the target action feature from a preset action animation database.

[0192] The action animation whose action feature has the highest similarity with the target action feature is used to determine second target action animation data that matches the target action feature.

[0193] In summary, in this embodiment, given action animation data is input; the target action feature extraction module in the target feature extraction network is used to extract features from the given action animation data to obtain target action features corresponding to the given action animation data; and according to the target action features, second target action animation data matching the target action features is determined from a preset action animation database. Thus, the action animation data is used for animation retrieval, which improves the retrieval accuracy and is convenient for users.

[0194] The following is a description of a feature extraction network training device, equipment, storage medium and program provided by the present application for execution. The specific implementation process and technical effects are mentioned above and will not be repeated below.

[0195] Fig.13 A schematic diagram of a feature extraction network training device provided in an embodiment of the present application is shown in FIG. Fig.13 As shown, the device comprises:

[0196] The first acquisition module 1301 is used to acquire multiple pieces of sample action animation data, each piece of sample action animation data has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data.

[0197] The first extraction module 1302 is used to respectively use a preset action feature extraction module and a preset text feature extraction module to extract features from each sample action animation data and the corresponding text annotation information to obtain a sample action feature and a sample text feature corresponding to each sample action animation data.

[0198] The training module 1303 is used to adjust the parameters of the preset text feature extraction module and the preset action feature extraction module according to the sample text features and sample action features corresponding to multiple sample action animation data, so as to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

[0199] Furthermore, the first extraction module 1302 is specifically used to calculate a first loss function value based on sample text features and sample action features corresponding to multiple sample action animation data; calculate a target loss function value based on the first loss function value; and adjust parameters of a preset text feature extraction module and a preset action feature extraction module based on the target loss function value to obtain a target feature extraction network.

[0200] Furthermore, the first extraction module 1302 is specifically used to calculate the first action loss function values ​​of multiple sample action animation data based on the sample text features and sample action features corresponding to the multiple sample action animation data; the first action loss function value is the loss function value of the sample action feature relative to the sample text feature; calculate the first text loss function values ​​of the multiple sample action animation data based on the sample text features and sample action features corresponding to the multiple sample action animation data; the first text loss function value is the loss function value of the sample text feature relative to the sample action feature; calculate the first loss function value based on the first action loss function value and the first text loss function value.

[0201] Furthermore, the first extraction module 1302 is specifically used for text annotation information including: action content annotation information, sample action features including: first action content information, and sample text features including: second action content information; calculating the second loss function value based on the first action content information and the second action content information corresponding to multiple sample action animation data; calculating the target loss function value based on the first loss function value and the second loss function value.

[0202] Furthermore, the first extraction module 1302 is specifically used for text annotation information also including: action style annotation information; the sample action features also include: first action style information, and the sample text features also include: second action style information; according to the first action style information and the second action style information corresponding to multiple sample action animation data, the third loss function value is calculated; according to the first loss function value and the third loss function value, the target loss function value is calculated.

[0203] Furthermore, the first extraction module 1302 is specifically used to calculate the second action loss function value of multiple sample action animation data based on the first action content information and the second action content information corresponding to the multiple sample action animation data; the second action loss function value is the loss function value of the first action content information relative to the second action content information; calculate the second text loss function value of the multiple sample action animation data based on the first action content information and the second action content information corresponding to the multiple sample action animation data; the second text loss function value is the loss function value of the second action content information relative to the first action content information; calculate the second loss function value based on the second action loss function value and the second text loss function value.

[0204] Further, the first extraction module 1302 is specifically used for action style annotation information including: annotation information of at least one action style, and correspondingly, the first action style information includes: a first action style feature of at least one action style, and the second action style information includes: a second action style feature of at least one action style; according to the first action style feature and the second action style feature of each action style corresponding to multiple sample action animation data, the loss function value of each action style is calculated; according to the loss function value of at least one action style, a third loss function value is calculated.

[0205] Furthermore, the first extraction module 1302 is specifically used to calculate the third action loss function value of each action style according to the first action style features and the second action style features of each action style corresponding to multiple sample action animation data; the third action loss function value is the loss function value of the first action style feature relative to the second action style feature; according to the first action style features and the second action style features of each action style corresponding to multiple sample action animation data, calculate the third text loss function value of each action style; the third text loss function value is the loss function value of the second action style feature relative to the first action style feature; according to the third action loss function value and the third text loss function value of each action style, calculate the loss function value of each action style.

[0206] Through the above method, by obtaining multiple sample action animation data, each sample action animation data has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data; using a preset action feature extraction module and a preset text feature extraction module respectively, feature extraction is performed on each sample action animation data and the corresponding text annotation information, and the sample action feature and sample text feature corresponding to each sample action animation data are obtained; according to the sample text features and sample action features corresponding to the multiple sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module. Thus, by combining the sample text features and the sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, and obtain a more accurate target feature extraction network, which can realize both text retrieval and action retrieval, without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0207] The following is a description of an action animation retrieval device, equipment, storage medium and program provided by the present application for execution. The specific implementation process and technical effects are described above and will not be repeated below.

[0208] Fig.14A schematic diagram of an action animation retrieval device provided in an embodiment of the present application is shown in FIG. Fig.14 As shown, the device comprises:

[0209] The second acquisition module 1401 is used to acquire input text action description information.

[0210] The second extraction module 1402 is used to use the target text feature extraction module in the target feature extraction network to perform feature extraction on the text action description information to obtain the target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by any of the methods in the above embodiments.

[0211] The determination module 1403 is used to determine, according to the target text feature, first target action animation data matching the target text feature from a preset action animation database.

[0212] Furthermore, the second extraction module 1402 is specifically used for text action description information including: action content information and / or action style information; using a target text feature extraction module, extracting features from the action content information and / or action style information to obtain target action content features and / or target action style features.

[0213] Furthermore, the determination module 1403 is specifically configured to determine, from a preset action animation database, first target action animation data matching the target action content feature and / or the target action style feature according to the target action content feature and / or the target action style feature.

[0214] Furthermore, the second acquisition module 1401 is also used to acquire input given action animation data;

[0215] Furthermore, the second extraction module 1402 is further used to use the target action feature extraction module in the target feature extraction network to perform feature extraction on the given action animation data to obtain the target action feature corresponding to the given action animation data;

[0216] Furthermore, the determination module 1403 is further configured to determine, according to the target action feature, second target action animation data matching the target action feature from a preset action animation database.

[0217] Through the above method, the input text action description information is obtained; the target text feature extraction module in the target feature extraction network is used to extract features of the text action description information to obtain the target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the above method; according to the target text features, the first target action animation data matching the target text features is determined from the preset action animation database. Thus, by combining the sample text features and the sample action features for multimodal cross-training, a more accurate target feature extraction network is obtained, which can realize both text retrieval and action retrieval without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0218] Fig.15 A structural diagram of an electronic device provided in an embodiment of the present application includes: a processor 1501, a storage medium 1502 and a bus 1503, the storage medium 1502 stores machine-readable instructions executable by the processor 1501, the processor 1501 communicates with the storage medium 1502 via the bus 1503, and the processor 1501 executes the machine-readable instructions to perform the above method.

[0219] Acquire multiple pieces of sample action animation data, each piece of sample action animation data having corresponding text annotation information, and the text annotation information is used to describe action attributes of the sample action animation data;

[0220] Using a preset action feature extraction module and a preset text feature extraction module respectively, feature extraction is performed on each sample action animation data and the corresponding text annotation information to obtain a sample action feature and a sample text feature corresponding to each sample action animation data;

[0221] According to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

[0222] Optionally, according to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module, including:

[0223] Calculating a first loss function value according to sample text features and sample action features corresponding to the plurality of sample action animation data;

[0224] Calculate the target loss function value according to the first loss function value;

[0225] According to the target loss function value, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain the target feature extraction network.

[0226] Optionally, calculating the first loss function value according to the sample text features and the sample action features corresponding to the plurality of sample action animation data includes:

[0227] Calculate the first action loss function value of the plurality of sample action animation data according to the sample text features and the sample action features corresponding to the plurality of sample action animation data; the first action loss function value is the loss function value of the sample action feature relative to the sample text feature;

[0228] Calculate the first text loss function value of the plurality of sample action animation data according to the sample text features and the sample action features corresponding to the plurality of sample action animation data; the first text loss function value is the loss function value of the sample text feature relative to the sample action feature;

[0229] A first loss function value is calculated according to the first action loss function value and the first text loss function value.

[0230] Optionally, the text annotation information includes: action content annotation information, the sample action feature includes: first action content information, and the sample text feature includes: second action content information;

[0231] Before calculating the target loss function value according to the first loss function value, the method further includes:

[0232] Calculate a second loss function value according to first action content information and second action content information corresponding to the plurality of sample action animation data;

[0233] According to the first loss function value, the target loss function value is calculated, including:

[0234] Calculate the target loss function value based on the first loss function value and the second loss function value.

[0235] Optionally, the text annotation information further includes: action style annotation information; the sample action feature further includes: first action style information, and the sample text feature further includes: second action style information;

[0236] Before calculating the target loss function value according to the first loss function value, the method further includes:

[0237] Calculating a third loss function value according to first action style information and second action style information corresponding to the plurality of sample action animation data;

[0238] According to the first loss function value, the target loss function value is calculated, including:

[0239] Calculate the target loss function value based on the first loss function value and the third loss function value.

[0240] Optionally, calculating the second loss function value according to the first action content information and the second action content information corresponding to the plurality of sample action animation data includes:

[0241] Calculate the second action loss function value of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second action loss function value is the loss function value of the first action content information relative to the second action content information;

[0242] Calculate the second text loss function value of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second text loss function value is the loss function value of the second action content information relative to the first action content information;

[0243] A second loss function value is calculated based on the second action loss function value and the second text loss function value.

[0244] Optionally, the action style annotation information includes: annotation information of at least one action style, and correspondingly, the first action style information includes: a first action style feature of at least one action style, and the second action style information includes: a second action style feature of at least one action style;

[0245] Calculating a third loss function value according to the first action style information and the second action style information corresponding to the plurality of sample action animation data includes:

[0246] Calculating a loss function value of each action style according to a first action style feature and a second action style feature of each action style corresponding to the plurality of sample action animation data;

[0247] A third loss function value is calculated based on the loss function value of at least one action style.

[0248] Optionally, calculating the loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data includes:

[0249] Calculate a third action loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data; the third action loss function value is the loss function value of the first action style feature relative to the second action style feature;

[0250] Calculate a third text loss function value for each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data; the third text loss function value is a loss function value of the second action style feature relative to the first action style feature;

[0251] The loss function value of each action style is calculated according to the third action loss function value and the third text loss function value of each action style.

[0252] Through the above method, by obtaining multiple sample action animation data, each sample action animation data has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data; using a preset action feature extraction module and a preset text feature extraction module respectively, feature extraction is performed on each sample action animation data and the corresponding text annotation information, and the sample action feature and sample text feature corresponding to each sample action animation data are obtained; according to the sample text features and sample action features corresponding to the multiple sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module. Thus, by combining the sample text features and the sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, and obtain a more accurate target feature extraction network, which can realize both text retrieval and action retrieval, without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0253] Get the input text action description information;

[0254] Using the target text feature extraction module in the target feature extraction network, feature extraction is performed on the text action description information to obtain target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by any of the methods in the above embodiments;

[0255] According to the target text feature, first target action animation data matching the target text feature is determined from a preset action animation database.

[0256] Optionally, the text action description information includes: action content information and / or action style information;

[0257] The target text feature extraction module in the target feature extraction network is used to extract features from the text action description information to obtain target text features corresponding to the text action description information, including:

[0258] Using a target text feature extraction module to extract features of the action content information and / or the action style information to obtain target action content features and / or target action style features;

[0259] According to the target text feature, determining first target action animation data matching the target text feature from a preset action animation database includes:

[0260] According to the target action content feature and / or the target action style feature, first target action animation data matching the target action content feature and / or the target action style feature is determined from a preset action animation database.

[0261] Optionally, the method further comprises:

[0262] Get the input given action animation data;

[0263] Using the target action feature extraction module in the target feature extraction network, feature extraction is performed on the given action animation data to obtain the target action feature corresponding to the given action animation data;

[0264] According to the target action feature, second target action animation data matching the target action feature is determined from a preset action animation database.

[0265] Through the above method, the input text action description information is obtained; the target text feature extraction module in the target feature extraction network is used to extract features of the text action description information to obtain the target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the above method; according to the target text features, the first target action animation data matching the target text features is determined from the preset action animation database. Thus, by combining the sample text features and the sample action features for multimodal cross-training, a more accurate target feature extraction network is obtained, which can realize both text retrieval and action retrieval without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0266] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method embodiment is executed.

[0267] Acquire multiple pieces of sample action animation data, each piece of sample action animation data having corresponding text annotation information, and the text annotation information is used to describe action attributes of the sample action animation data;

[0268] Using a preset action feature extraction module and a preset text feature extraction module respectively, feature extraction is performed on each sample action animation data and the corresponding text annotation information to obtain a sample action feature and a sample text feature corresponding to each sample action animation data;

[0269] According to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

[0270] Optionally, according to the sample text features and sample action features corresponding to the plurality of sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module, including:

[0271] Calculating a first loss function value according to sample text features and sample action features corresponding to the plurality of sample action animation data;

[0272] Calculate the target loss function value according to the first loss function value;

[0273] According to the target loss function value, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain the target feature extraction network.

[0274] Optionally, calculating the first loss function value according to the sample text features and the sample action features corresponding to the plurality of sample action animation data includes:

[0275] Calculate the first action loss function value of the plurality of sample action animation data according to the sample text features and the sample action features corresponding to the plurality of sample action animation data; the first action loss function value is the loss function value of the sample action feature relative to the sample text feature;

[0276] Calculate the first text loss function value of the plurality of sample action animation data according to the sample text features and the sample action features corresponding to the plurality of sample action animation data; the first text loss function value is the loss function value of the sample text feature relative to the sample action feature;

[0277] A first loss function value is calculated according to the first action loss function value and the first text loss function value.

[0278] Optionally, the text annotation information includes: action content annotation information, the sample action feature includes: first action content information, and the sample text feature includes: second action content information;

[0279] Before calculating the target loss function value according to the first loss function value, the method further includes:

[0280] Calculate a second loss function value according to first action content information and second action content information corresponding to the plurality of sample action animation data;

[0281] According to the first loss function value, the target loss function value is calculated, including:

[0282] Calculate the target loss function value based on the first loss function value and the second loss function value.

[0283] Optionally, the text annotation information further includes: action style annotation information; the sample action feature further includes: first action style information, and the sample text feature further includes: second action style information;

[0284] Before calculating the target loss function value according to the first loss function value, the method further includes:

[0285] Calculating a third loss function value according to first action style information and second action style information corresponding to the plurality of sample action animation data;

[0286] According to the first loss function value, the target loss function value is calculated, including:

[0287] Calculate the target loss function value based on the first loss function value and the third loss function value.

[0288] Optionally, calculating the second loss function value according to the first action content information and the second action content information corresponding to the plurality of sample action animation data includes:

[0289] Calculate the second action loss function value of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second action loss function value is the loss function value of the first action content information relative to the second action content information;

[0290] Calculate the second text loss function value of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second text loss function value is the loss function value of the second action content information relative to the first action content information;

[0291] A second loss function value is calculated based on the second action loss function value and the second text loss function value.

[0292] Optionally, the action style annotation information includes: annotation information of at least one action style, and correspondingly, the first action style information includes: a first action style feature of at least one action style, and the second action style information includes: a second action style feature of at least one action style;

[0293] Calculating a third loss function value according to the first action style information and the second action style information corresponding to the plurality of sample action animation data includes:

[0294] Calculating a loss function value of each action style according to a first action style feature and a second action style feature of each action style corresponding to the plurality of sample action animation data;

[0295] A third loss function value is calculated based on the loss function value of at least one action style.

[0296] Optionally, calculating the loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data includes:

[0297] Calculate a third action loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data; the third action loss function value is the loss function value of the first action style feature relative to the second action style feature;

[0298] Calculate a third text loss function value for each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data; the third text loss function value is a loss function value of the second action style feature relative to the first action style feature;

[0299] The loss function value of each action style is calculated according to the third action loss function value and the third text loss function value of each action style.

[0300] Through the above method, by obtaining multiple sample action animation data, each sample action animation data has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data; using a preset action feature extraction module and a preset text feature extraction module respectively, feature extraction is performed on each sample action animation data and the corresponding text annotation information, and the sample action feature and sample text feature corresponding to each sample action animation data are obtained; according to the sample text features and sample action features corresponding to the multiple sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module. Thus, by combining the sample text features and the sample action features for multimodal cross-training, the target feature extraction network can accurately extract text features and action features, and obtain a more accurate target feature extraction network, which can realize both text retrieval and action retrieval, without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0301] Get the input text action description information;

[0302] Using the target text feature extraction module in the target feature extraction network, feature extraction is performed on the text action description information to obtain target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by any of the methods in the above embodiments;

[0303] According to the target text feature, first target action animation data matching the target text feature is determined from a preset action animation database.

[0304] Optionally, the text action description information includes: action content information and / or action style information;

[0305] The target text feature extraction module in the target feature extraction network is used to extract features from the text action description information to obtain target text features corresponding to the text action description information, including:

[0306] Using a target text feature extraction module to extract features of the action content information and / or the action style information to obtain target action content features and / or target action style features;

[0307] According to the target text feature, determining first target action animation data matching the target text feature from a preset action animation database includes:

[0308] According to the target action content feature and / or the target action style feature, first target action animation data matching the target action content feature and / or the target action style feature is determined from a preset action animation database.

[0309] Optionally, the method further comprises:

[0310] Get the input given action animation data;

[0311] Using the target action feature extraction module in the target feature extraction network, feature extraction is performed on the given action animation data to obtain the target action feature corresponding to the given action animation data;

[0312] According to the target action feature, second target action animation data matching the target action feature is determined from a preset action animation database.

[0313] Through the above method, the input text action description information is obtained; the target text feature extraction module in the target feature extraction network is used to extract features of the text action description information to obtain the target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the above method; according to the target text features, the first target action animation data matching the target text features is determined from the preset action animation database. Thus, by combining the sample text features and the sample action features for multimodal cross-training, a more accurate target feature extraction network is obtained, which can realize both text retrieval and action retrieval without rendering the action file as a video animation, thereby greatly reducing storage and computing consumption.

[0314] In the embodiment of the present application, the computer program can also execute other machine-readable instructions when run by the processor to execute other methods described in the embodiment. For the specific execution method steps and principles, please refer to the description of the embodiment, which will not be repeated here.

[0315] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0316] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0317] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0318] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0319] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.

[0320] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A feature extraction network training method, characterized in that: The method comprises: Acquire multiple pieces of sample action animation data, each piece of sample action animation data having corresponding text annotation information, wherein the text annotation information is used to describe action attributes of the sample action animation data; Using a preset action feature extraction module and a preset text feature extraction module respectively, extracting features from each piece of sample action animation data and the corresponding text annotation information, and obtaining a sample action feature and a sample text feature corresponding to each piece of sample action animation data; According to the sample text features and sample action features corresponding to the multiple sample action animation data, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

2. The method according to claim 1, characterized in that The preset text feature extraction module and the preset action feature extraction module are adjusted according to the sample text features and the sample action features corresponding to the plurality of sample action animation data to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module, including: Calculating a first loss function value according to sample text features and sample action features corresponding to the plurality of sample action animation data; Calculate a target loss function value according to the first loss function value; According to the target loss function value, the preset text feature extraction module and the preset action feature extraction module are adjusted to obtain the target feature extraction network.

3. The method according to claim 2, characterized in that The calculating the first loss function value according to the sample text features and the sample action features corresponding to the plurality of sample action animation data includes: Calculate the first action loss function value of the plurality of sample action animation data according to the sample text features and the sample action features corresponding to the plurality of sample action animation data; the first action loss function value is the loss function value of the sample action feature relative to the sample text feature; Calculating first text loss function values ​​of the plurality of sample action animation data according to sample text features and sample action features corresponding to the plurality of sample action animation data; the first text loss function value is a loss function value of the sample text feature relative to the sample action feature; The first loss function value is calculated according to the first action loss function value and the first text loss function value.

4. The method according to claim 2, characterized in that: The text annotation information includes: action content annotation information, the sample action features include: first action content information, and the sample text features include: second action content information; Before calculating the target loss function value according to the first loss function value, the method further includes: Calculate a second loss function value according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; The calculating a target loss function value according to the first loss function value includes: The target loss function value is calculated based on the first loss function value and the second loss function value.

5. The method according to claim 2 or 4, characterized in that: The text annotation information also includes: action style annotation information; the sample action feature also includes: first action style information, and the sample text feature also includes: second action style information; Before calculating the target loss function value according to the first loss function value, the method further includes: Calculating a third loss function value according to the first action style information and the second action style information corresponding to the plurality of sample action animation data; The calculating a target loss function value according to the first loss function value includes: The target loss function value is calculated based on the first loss function value and the third loss function value.

6. The method according to claim 4, characterized in that The calculating the second loss function value according to the first action content information and the second action content information corresponding to the plurality of sample action animation data includes: Calculate the second action loss function value of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second action loss function value is the loss function value of the first action content information relative to the second action content information; Calculate the second text loss function value of the plurality of sample action animation data according to the first action content information and the second action content information corresponding to the plurality of sample action animation data; the second text loss function value is the loss function value of the second action content information relative to the first action content information; The second loss function value is calculated according to the second action loss function value and the second text loss function value.

7. The method according to claim 5, characterized in that The action style annotation information includes: annotation information of at least one action style, and correspondingly, the first action style information includes: a first action style feature of the at least one action style, and the second action style information includes: a second action style feature of the at least one action style; The calculating the third loss function value according to the first action style information and the second action style information corresponding to the plurality of sample action animation data comprises: Calculating a loss function value of each action style according to a first action style feature and a second action style feature of each action style corresponding to the plurality of sample action animation data; The third loss function value is calculated according to the loss function value of the at least one action style.

8. The method according to claim 7, characterized in that The calculating the loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data comprises: Calculate a third action loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data; the third action loss function value is a loss function value of the first action style feature relative to the second action style feature; Calculate a third text loss function value of each action style according to the first action style feature and the second action style feature of each action style corresponding to the plurality of sample action animation data; the third text loss function value is a loss function value of the second action style feature relative to the first action style feature; The loss function value of each action style is calculated according to the third action loss function value and the third text loss function value of each action style.

9. A method for retrieving action animations, characterized in that: The method comprises: Get the input text action description information; Using a target text feature extraction module in a target feature extraction network, feature extraction is performed on the text action description information to obtain target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the method described in any one of claims 1 to 8 above; According to the target text feature, first target action animation data matching the target text feature is determined from a preset action animation database.

10. The method according to claim 9, characterized in that The text action description information includes: action content information and / or action style information; The target text feature extraction module in the target feature extraction network is used to extract features from the text action description information to obtain target text features corresponding to the text action description information, including: Using the target text feature extraction module, extracting features from the action content information and / or the action style information to obtain target action content features and / or target action style features; The step of determining, according to the target text feature, first target action animation data matching the target text feature from a preset action animation database comprises: According to the target action content feature and / or the target action style feature, the first target action animation data matching the target action content feature and / or the target action style feature is determined from the preset action animation database.

11. The method according to claim 9, characterized in that The method further comprises: Get the input given action animation data; Using the target action feature extraction module in the target feature extraction network to extract features from the given action animation data, and obtaining target action features corresponding to the given action animation data; According to the target action feature, second target action animation data matching the target action feature is determined from the preset action animation database.

12. A feature extraction network training device, characterized in that: The device comprises: A first acquisition module is used to acquire a plurality of sample action animation data, each of which has corresponding text annotation information, and the text annotation information is used to describe the action attributes of the sample action animation data; A first extraction module is used to respectively use a preset action feature extraction module and a preset text feature extraction module to perform feature extraction on each piece of sample action animation data and the corresponding text annotation information to obtain a sample action feature and a sample text feature corresponding to each piece of sample action animation data; The training module is used to adjust the parameters of the preset text feature extraction module and the preset action feature extraction module according to the sample text features and sample action features corresponding to the multiple sample action animation data, so as to obtain a target feature extraction network including a target text feature extraction module and a target action feature extraction module.

13. An action animation retrieval device, characterized in that: The device comprises: The second acquisition module is used to acquire input text action description information; A second extraction module is used to use a target text feature extraction module in a target feature extraction network to perform feature extraction on the text action description information to obtain target text features corresponding to the text action description information; wherein the target feature extraction network is a network model trained by the method described in any one of claims 1 to 8 above; The determination module is used to determine, according to the target text feature, first target action animation data matching the target text feature from a preset action animation database.

14. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the feature extraction network training method as described in any one of claims 1 to 8, or the steps of the action animation retrieval method as described in any one of claims 9 to 11.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the feature extraction network training method according to any one of claims 1 to 8, or the steps of the action animation retrieval method according to any one of claims 9 to 11.

16. A computer program instruction, characterized in that The program instructions are executed by the processor to perform the steps of the feature extraction network training method as described in any one of claims 1 to 8, or the steps of the action animation retrieval method as described in any one of claims 9 to 11.