Prompt voice posture generation method, related device and medium

Generate pose descriptions through large language models and use diffusion model guidance, the problem of rigidity in the prior art prompted speech pose generation is solved, more accurate and detailed pose generation is achieved, and the ability of hearing-impaired people to understand speech or text is improved.

CN120409484APending Publication Date: 2025-08-01SHENZHEN TENCENT COMP SYST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410129489.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing cue speech pose generation techniques lack semantic connections, resulting in the generated pose rigidity, insufficient accuracy and precision, and inaccuracy to accurately help hearing-impaired people understand speech or text.

Method used

By obtaining the target input, using the large language model to generate pose descriptions, and encoding them into pose description guidance vectors, guiding them in combination with the diffusion model, gradually adding semantic related details to generate target prompt speech poses.

Benefits of technology

Improve the accuracy and precision of the prompted speech posture, so that the generated posture has a clear semantic connection with the input speech or text, and enhance the ability of hearing-impaired people to understand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409484A_ABST
    Figure CN120409484A_ABST
Patent Text Reader

Abstract

The invention provides a prompt voice posture generation method, a related device and a medium. The method comprises the steps of obtaining target input, wherein the target input comprises at least one of a target text and target voice; adding the target input into the guide language input large language model to obtain a posture description for describing a target prompt voice posture; generating an attitude description guide vector corresponding to the attitude description; guiding a diffusion model by using the attitude description guide vector, so that the diffusion model generates a target prompt voice attitude vector; and generating a target prompt voice attitude based on the target prompt voice attitude vector. According to the embodiment of the invention, the accuracy and fineness of generating the prompt voice posture can be improved. The embodiment of the invention can be applied to various scenes such as online education, online communication, video processing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a method for generating prompting speech gestures, related devices, and media. Background Art

[0002] Prompting speech (CS) is a technology that uses hand and lip gestures to assist deaf people in communication. Prompting speech gestures are the gestures made by hands and lips to help deaf people understand text. Deaf people cannot hear speech or understand text normally. Through prompting speech gestures, they can accurately understand speech or text. In the existing technologies for helping deaf people understand speech or text, generally, the rules for mapping text to gestures are solidified in the model, and then the text or speech is input into the model to obtain prompting speech gestures, etc., or the prompting speech gestures are generated through a machine learning model. However, no matter which method is used, there is no semantic connection between the text or speech and the generated gestures. For example, when representing "tree", a gesture of spreading two fingers and placing them at the neck is generated, and the semantics of this gesture has no connection with the semantics of "tree", and the mapping is completely realized by mechanical rules. Therefore, the generated gestures are rigid, with poor accuracy and insufficient fineness.

[0003] Therefore, a technology for generating prompting speech gestures more accurately and finely is needed. Summary of the Invention

[0004] The present disclosure provides a method for generating prompting speech gestures, related devices, and media, which can improve the accuracy and fineness of generating prompting speech gestures.

[0005] According to one aspect of the present disclosure, a method for generating prompting speech gestures is provided, including:

[0006] Obtaining a target input, where the target input includes at least one of a target text and a target speech;

[0007] Adding the target input to a guiding language input large language model to obtain a gesture description describing the target prompting speech gesture;

[0008] Generating a gesture description guiding vector corresponding to the gesture description;

[0009] Guiding a diffusion model by using the gesture description guiding vector to enable the diffusion model to generate a target prompting speech gesture vector;

[0010] Generating a target prompting speech gesture based on the target prompting speech gesture vector.

[0011] According to one aspect of the present disclosure, a prompting speech gesture generating device is provided, including:

[0012] A first acquisition unit, configured to acquire a target input, where the target input includes at least one of a target text and a target voice;

[0013] A first generation unit, configured to add the target input to a guiding language input into a large language model to obtain a gesture description for describing a target prompt voice gesture;

[0014] A second generation unit, configured to generate a gesture description guiding vector corresponding to the gesture description;

[0015] A diffusion unit, configured to guide a diffusion model by using the gesture description guiding vector, so that the diffusion model generates a target prompt voice gesture vector;

[0016] A third generation unit, configured to generate a target prompt voice gesture based on the target prompt voice gesture vector.

[0017] Optionally, the first generation unit is specifically configured to:

[0018] Acquire an overall rule for converting the target input into the gesture description;

[0019] Add the target input to the guiding language and the overall rule, and input the result into the large language model to obtain the gesture description.

[0020] Optionally, the prompt voice gesture generation device further includes a first training unit, configured to jointly train the guiding language and the overall rule in the following manner:

[0021] Acquire a first sample set, where the first sample set has a first sample text and a gesture label corresponding to the first sample text;

[0022] Add each first sample text in the first sample set to the guiding language and the overall rule, and input the result into the large language model to obtain a prediction result corresponding to the first sample text;

[0023] Calculate a first loss function based on the prediction result corresponding to the first sample text and the gesture label;

[0024] If the first loss function is less than a first threshold, stop training the guiding language and the overall rule; otherwise, adjust the guiding language and the overall rule, and return to the step of adding each first sample text in the first sample set to the guiding language and the overall rule, and inputting the result into the large language model to obtain a prediction result corresponding to the sample text.

[0025] Optionally, in an embodiment, the target input includes a target text and a target voice, and the first generation unit is specifically configured to: add the target text to the guiding language and input the result into the large language model;

[0026] The prompt voice gesture generation device further includes a fourth generation unit for generating a voice guidance vector corresponding to the target voice;

[0027] Specifically, the diffusion unit is configured to: guide the diffusion model by using the pose description guidance vector and the voice guidance vector, so that the diffusion model generates a target prompt voice pose vector;

[0028] In one embodiment, the second generation unit is specifically configured to: input the pose description into a first vector generation model to obtain the pose description guidance vector;

[0029] The fourth generation unit is specifically configured to: input the target voice into a second vector generation model to obtain the voice guidance vector;

[0030] Optionally, the prompt voice gesture generation device further includes a second training unit for training the first vector generation model in the following manner:

[0031] Obtain a sample pair set, where the sample pairs in the sample pair set include a plurality of first sample pairs and a plurality of second sample pairs. The first sample pairs include matching sample poses and sample pose descriptions, and the second sample pairs include the unmatched sample poses and the sample pose descriptions;

[0032] Convert the sample pose description into a sample pose description vector through the first vector generation model, and convert the sample pose into a sample pose vector through a third vector generation model;

[0033] Jointly train the first vector generation model and the third vector generation model, so that the distance between the sample pose description vector and the sample pose vector of the first sample pair becomes smaller, and the distance between the sample pose description vector and the sample pose vector of the second sample pair becomes larger.

[0034] In one embodiment, the second training unit is specifically configured to:

[0035] Set the matching label of the first sample pair to 1 and the matching label of the second sample pair to 0;

[0036] Through the first vector generation model, based on the sample pose description vector and the sample pose vector of the sample pair, predict a first probability that the sample pose description and the sample pose in the sample pair match;

[0037] Through the third vector generation model, based on the sample pose description vector and the sample pose vector of the sample pair, predict a second probability that the sample pose and the sample pose description in the sample pair match;

[0038] Based on the matching tags, the first probability, and the second probability of each of the sample pairs, calculate a second loss function, and jointly train the first vector generation model and the third vector generation model based on the second loss function.

[0039] In one embodiment, the diffusion unit is specifically configured to:

[0040] In some embodiments, the guiding of the diffusion model using the pose description guiding vector and the speech guiding vector to enable the diffusion model to generate a target prompt speech pose vector includes:

[0041] Initialize the prompt speech pose vector to be diffused as an initial prompt speech pose vector, and initialize the step number to 1;

[0042] Input the prompt speech pose vector to be diffused, the step number, the pose description guiding vector, and the speech guiding vector into the diffusion model to obtain the diffusion noise corresponding to the step number;

[0043] Subtract the diffusion noise corresponding to the step number from the prompt speech pose vector to be diffused, increment the step number by 1, and return to the step of inputting the prompt speech pose vector to be diffused, the step number, the pose description guiding vector, and the speech guiding vector into the diffusion model until the step number is incremented to a preset maximum number of steps.

[0044] Optionally, the prompt speech pose generation device further includes a third training unit for training the diffusion model in the following manner:

[0045] Obtain a second sample set, where each second sample in the second sample set includes a sample reference prompt speech pose vector, a sample pose description vector, and a sample speech guiding vector;

[0046] Perform multi-step noise addition processing on the sample reference prompt speech pose vector, and record the noise added at each step number as the noise label corresponding to the step number;

[0047] Generate a prompt speech pose label corresponding to the step number based on the noise-added prompt speech pose vector obtained after each step of the noise addition processing;

[0048] Obtain a rhythm adjustment pose difference label corresponding to the step number;

[0049] Input the sample speech pose vector obtained after the multi-step noise addition processing into the diffusion model, and perform multi-step denoising processing under the guidance of the sample pose description vector and the sample speech guidance vector. Calculate the third loss function based on the denoising result of each step, the noise label corresponding to the step number, the prompt speech pose label, and the rhythm adjustment pose difference label, and train the diffusion model based on the third loss function.

[0050] In one embodiment, the third training unit is specifically configured to:

[0051] Calculating the third loss function based on the denoising result of each step, the noise label corresponding to the step number, the prompt speech pose label, and the rhythm adjustment pose difference label includes:

[0052] Calculate a first loss sub-function based on the predicted diffusion noise corresponding to the step number and the noise label;

[0053] Calculate a second loss sub-function based on the predicted prompt speech pose corresponding to the step number and the prompt speech pose label;

[0054] Calculate a third loss sub-function based on the predicted rhythm adjustment pose difference corresponding to the step number and the rhythm adjustment pose difference label;

[0055] Calculate the third loss function based on the first loss sub-function, the second loss sub-function, and the third loss sub-function.

[0056] In one embodiment, the third training unit is specifically configured to:

[0057] Calculate the first difference between the predicted diffusion noise corresponding to the step number and the noise label;

[0058] Calculate the square of the norm of the first difference as the first loss sub-function.

[0059] In one embodiment, the third training unit is specifically configured to:

[0060] Calculate the cosine distance between the predicted prompt speech pose corresponding to the step number and the prompt speech pose label;

[0061] Take the difference between 1 and the cosine distance as the second loss sub-function.

[0062] In one embodiment, the third training unit is specifically configured to:

[0063] Obtain the first weight of the first loss sub-function, the second weight of the second loss sub-function, and the third weight of the third loss sub-function;

[0064] Using the first weight, the second weight, and the third weight, perform a weighted sum on the first loss sub-function, the second loss sub-function, and the third loss sub-function to obtain the third loss function.

[0065] In one embodiment, the diffusion model is used to predict diffusion noise in the following manner:

[0066] Input a first proportion of the sample prompt speech pose vectors, the step number, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain a first sub-predicted diffusion noise corresponding to the step number;

[0067] Input a second proportion of the sample prompt speech pose vectors, the step number, and the speech guidance vector into the diffusion model to obtain a second sub-predicted diffusion noise corresponding to the step number, where the sum of the first proportion and the second proportion is 1;

[0068] Perform a weighted sum of the first sub-predicted diffusion noise and the second sub-predicted diffusion noise according to the first proportion and the second proportion to obtain the predicted diffusion noise.

[0069] In one embodiment, the diffusion unit is specifically configured to:

[0070] Input a concatenated vector of the to-be-diffused prompt speech pose vector and the step number into a multi-head attention model to obtain a multi-head attention vector;

[0071] Superimpose and normalize the concatenated vector and the multi-head attention vector to obtain a first normalized vector;

[0072] Input the first normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain a diffusion noise corresponding to the step number.

[0073] In one embodiment, the diffusion unit is specifically configured to:

[0074] Input the first normalized vector into a first feed-forward neural network to obtain a first feed-forward vector;

[0075] Superimpose and normalize the first feed-forward vector and the first normalized vector to obtain a second normalized vector;

[0076] Input the second normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain a diffused vector;

[0077] Input the diffused vector into a second feed-forward neural network to obtain a second feed-forward vector;

[0078] Superimpose and normalize the second feedforward vector and the diffused vector to obtain the diffused noise.

[0079] Optionally, the prompt voice pose generation device further includes a rhythm adjustment unit, and the rhythm adjustment unit is specifically configured to:

[0080] Input the target voice into the rhythm adjustment pose difference generation model to obtain a rhythm adjustment pose difference;

[0081] Adjust the target prompt voice pose with the rhythm adjustment pose difference to obtain an adjusted prompt voice pose.

[0082] Optionally, the prompt voice pose generation device further includes a fourth training unit, and the fourth training unit is used to train the rhythm adjustment pose difference generation model in the following manner:

[0083] Obtain a third sample set, including multiple sample video segments obtained by splitting a target sample video and sample voice segments corresponding to the sample video segments;

[0084] Extract sample object motion representations from the sample video segments;

[0085] Determine the average motion representation of the sample object motion representations;

[0086] Determine the difference between the sample object motion representation and the average motion representation as the rhythm adjustment pose difference label;

[0087] Input the sample voice segments corresponding to the sample video segments into the rhythm adjustment pose difference generation model to obtain the predicted rhythm adjustment pose difference;

[0088] Calculate a fourth loss function based on the predicted rhythm adjustment pose difference and the rhythm adjustment pose difference label, and train the rhythm adjustment pose difference generation model based on the fourth loss function.

[0089] In one embodiment, the third training unit and the fourth training unit are specifically configured to jointly train the diffusion model and the rhythm adjustment pose difference generation model in the following manner:

[0090] Obtain a fourth sample set, and the fourth sample set includes multiple multimodal samples, and each multimodal sample includes a second sample text, a sample voice, and a sample video;

[0091] Input the second sample text into the large language model with the guiding language to obtain a sample pose description;

[0092] Based on the first guiding vector corresponding to the sample pose description and the second guiding vector corresponding to the sample speech, guide the diffusion model to generate a sample prompt speech pose vector, and generate a sample prompt speech pose based on the sample prompt speech pose vector;

[0093] Input the sample speech into the rhythm adjustment pose difference generation model to obtain a sample rhythm adjustment pose difference, and use the sample rhythm adjustment pose difference to adjust the sample prompt speech pose to obtain an adjusted sample prompt speech pose;

[0094] Based on the comparison between the adjusted sample prompt speech pose and the prompt speech pose label extracted from the sample video, calculate a fifth loss function, and train the rhythm adjustment pose difference generation model and the diffusion model based on the fifth loss function.

[0095] According to an aspect of the present disclosure, there is provided an electronic device including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned speech pose generation method is implemented.

[0096] According to an aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned speech pose generation method is implemented.

[0097] According to an aspect of the present disclosure, there is provided a computer program product including a computer program, which is read and executed by a processor of a computer device, so that the computer device executes the above-mentioned speech pose generation method.

[0098] In the embodiments of the present disclosure, considering that the semantics of the generated prompt speech pose has no connection with the semantics of the initial text or initial speech, an intermediate link of pose description for describing the prompt speech pose to be generated is introduced. This pose description is a description of the prompt speech pose, and there is a one-to-one semantic correspondence relationship with the prompt speech pose. The embodiments of the present disclosure utilize a mature large language model plus prompt engineering to generate a pose description for a target input. Then, according to the semantic correspondence between the pose description and the prompt speech pose, a pose description guiding vector corresponding to the pose description and a speech guiding vector corresponding to the target speech are generated, and the diffusion model is guided by these two vectors to finally generate a target prompt speech pose. Since the semantics between the pose description and the prompt speech pose are one-to-one corresponding, the generated target prompt speech pose is not rigid, can achieve fine-grainedness, and improves the accuracy and fineness of generating the prompt speech pose.

[0099] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present disclosure. The objectives and other advantages of the present disclosure may be realized and attained by the structure particularly pointed out in the specification, claims as well as the drawings. Description of the Drawings

[0100] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0101] Figure 1 is a framework diagram of the system to which the prompting voice gesture generation method according to the embodiment of the present disclosure is applied;

[0102] Figure 2 is a schematic diagram of the application of the embodiment of the present disclosure in a specific scenario of communicating with hearing-impaired persons;

[0103] Figure 3 is a flowchart of the prompting voice gesture generation method according to an embodiment of the present disclosure;

[0104] Figure 4 is a schematic diagram of the hardware topology for executing the prompting voice gesture generation method according to an embodiment of the present disclosure;

[0105] Figure 5 is a specific schematic diagram of the overall rules for Chinese prompting voice conversion provided by an embodiment of the present disclosure and the generation of prompting voice gestures using the prompting voice gesture generation method of the embodiment of the present disclosure;

[0106] Figure 6 is a specific implementation schematic diagram of steps 320 and 340 in the case where the target input includes target text and target voice;

[0107] Figure 7 is Figure 3 a specific implementation schematic diagram of generating gesture descriptions in step 320;

[0108] Figure 8 is a specific implementation schematic diagram of the training guiding language and overall rules according to an embodiment of the present disclosure;

[0109] Figure 9 is a flowchart of the prompting voice gesture generation method according to an embodiment of the present disclosure;

[0110] Figure 10 is a flowchart of training a first vector generation model according to an embodiment of the present disclosure;

[0111] Figure 11Schematic diagram of a specific implementation for training a first vector generation model according to an embodiment of the present disclosure;

[0112] Figure 12 is Figure 10 A refined flowchart of step 1030 in

[0113] Figure 13 is Figure 9 A refined flowchart of step 950 in

[0114] Figure 14 is associated with Figure 13 Corresponding process schematic diagram;

[0115] Figure 15 Schematic diagram of a specific implementation for training a diffusion model according to an embodiment of the present disclosure;

[0116] Figure 16 Schematic diagram of the process for training a diffusion model by forward noise addition and backward denoising according to an embodiment of the present disclosure;

[0117] Figure 17 is Figure 15 Sub - flowchart of step 1550 in

[0118] Figure 18 Schematic diagram of a specific implementation for generating predicted diffusion noise according to an embodiment of the present disclosure;

[0119] Figure 19 is Figure 13 Flowchart of step 1320 in

[0120] Figure 20 is Figure 19 Flowchart of step 1930 in

[0121] Figure 21 Schematic diagram of a specific implementation for adjusting the target prompt voice pose by adjusting the pose difference with rhythm;

[0122] Figure 22 Schematic diagram of a specific implementation for training a model for generating pose difference with rhythm adjustment according to an embodiment of the present disclosure;

[0123] Figure 23 Topological schematic diagram of a model for generating pose difference with rhythm adjustment according to an embodiment of the present disclosure;

[0124] Figure 24 Schematic diagram of a specific implementation for jointly training a diffusion model and a model for generating pose difference with rhythm adjustment according to an embodiment of the present disclosure;

[0125] Figure 25Schematic topology diagram for jointly training a diffusion model and a rhythm-adjusted pose difference generation model according to an embodiment of the present disclosure;

[0126] Figure 26 Schematic parameter diagram of the dataset for training the diffusion model used in an embodiment of the present disclosure;

[0127] Figure 27 Parameter comparison diagram of the open-source dataset for training the diffusion model according to an embodiment of the present disclosure;

[0128] Figure 28 Shows the experimental comparison data of the prompt voice pose generation method using the present disclosure embodiment to generate prompt voice poses and the existing prompt voice pose generation method;

[0129] Figure 29 Shows the experimental comparison data of generating prompt voice poses using the prompt voice pose generation method of the present disclosure embodiment after training the diffusion model with different types of datasets;

[0130] Figure 30 Shows the experimental comparison data of using different methods to extract audio features from the target voice and then generating prompt voice poses using the prompt voice pose generation method of the present disclosure embodiment;

[0131] Figure 31 Is a specific real-time process diagram of the prompt voice pose generation method of the present disclosure embodiment.

[0132] Figure 32 Block diagram of the prompt voice pose generation device according to an embodiment of the present disclosure;

[0133] Figure 33 Is to execute according to an embodiment of the present disclosure Figure 3 Terminal structure diagram of the prompt voice pose generation method shown;

[0134] Figure 34 Is to execute according to an embodiment of the present disclosure Figure 3 Server structure diagram of the prompt voice pose generation method shown. Detailed implementation manners

[0135] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0136] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0137] Cued Speech (CS): Also translated as "Cued Language", it is a method that helps deaf people understand and communicate using spoken language by encoding vowels with a small number of hand positions and consonants with a small number of handshapes. Different from traditional Sign Language, when expressing content through Sign Language, there is a direct semantic connection between the handshape gestures made and the content to be expressed. However, in Cued Speech, the gestures to be made depend on the pronunciation of the content to be expressed, rather than the semantics of the content to be expressed. That is, in Cued Speech, there is no semantic connection between the gesture and the content to be expressed. For example, in Sign Language, when expressing the word "tree", the corresponding gesture is "putting the thumbs and index fingers of both hands together to form a circle and moving upwards", and this gesture describes the shape of the tree trunk, and there is a certain connection between this gesture and the semantics of the word "tree". Similarly, in Sign Language, when expressing the word "book", the corresponding gesture is "putting the palms of both hands together and then opening them", and this gesture itself simulates the action of turning a book, and there is also a certain connection between this gesture and the semantics of the word "book". But in Cued Speech, when expressing the word "tree", the pinyin of "tree" is shu, its vowel is u [u], and the corresponding hand position is to put the hand on the neck position, and the consonant is The corresponding handshape is "extending the index finger and middle finger, with the index finger and middle finger separated". Then, in Cued Speech, when expressing the word "tree", the corresponding gesture is "putting the hand on the neck position, extending the index finger and middle finger, with the index finger and middle finger separated". It can be seen that there is an obvious lack of semantic connection between this gesture and the word "tree". Similarly, when expressing the word "book" through Cued Speech, since the vowels and consonants of the words "tree" and "book" are the same, the Cued Speech gesture for expressing the word "book" is the same as the Cued Speech gesture for expressing the word "tree", which is also "putting the hand on the neck position, extending the index finger and middle finger, with the index finger and middle finger separated", and there is also an obvious lack of semantic connection between this gesture and the word "book". As can be seen from the above, compared with Sign Language, which makes corresponding gestures directly according to the semantics of the content to be expressed for communication, Cued Speech focuses more on making corresponding gestures according to the pronunciation of the content to be expressed, and it can better help deaf people understand and use spoken language. Since Cued Speech makes corresponding gestures based on pronunciation, this means that in Cued Speech, only the mapping between the basic phonetic symbols of the pronunciation of the text and the gestures needs to be established. Although this makes the semantic connection between the content to be expressed and the corresponding gesture weaker, the conversion rule between the content to be expressed and the gesture is relatively simpler, the learning cost of Cued Speech becomes lower, and the expression of spoken language is more accurate, which can better help deaf people learn and understand spoken language and communicate.

[0138] Large Language Models (LLMs): They are deep learning models trained with a large amount of text data and can be used to generate natural language text. Currently, they are mainly applied in fields such as text summarization, question answering, and translation.

[0139] Cued Speech (CS) is a technology that uses hand and lip gestures to assist deaf and hard-of-hearing people in communication. Cued Speech gestures are the gestures made with hands and lips to help deaf and hard-of-hearing people understand text or speech. Deaf and hard-of-hearing people cannot hear speech normally. At the same time, due to the lack of hearing, it is very difficult for them to learn and understand spoken language. Through Cued Speech gestures, it is possible to help deaf and hard-of-hearing people accurately understand speech and learn and understand spoken language. In the existing technologies for helping deaf and hard-of-hearing people understand spoken language, generally, the rules for mapping text or speech to gestures are fixed in the model, and then the text or speech is input into the model to obtain Cued Speech gestures, etc., or the way of generating Cued Speech gestures through a machine learning model. However, no matter which way, there is no semantic connection between the text or speech and the generated gestures. For example, when representing "tree", a gesture of spreading two fingers and placing them at the neck is generated, and the semantics of this gesture have no connection with the semantics of "tree", and the mapping is completely achieved by mechanical rules. Therefore, the generated gestures are rigid, with poor accuracy and insufficient fineness.

[0140] Therefore, a technology for generating Cued Speech gestures more accurately and finely is needed.

[0141] System architecture and scenario description applied in the embodiments of the present disclosure

[0142] Figure 1 It is the system architecture diagram applied to the Cued Speech gesture generation method according to the embodiments of the present disclosure. It includes: object terminal 110, Internet 120, gateway 130, and server 140.

[0143] Object terminal 110 is a device for an object to input a target input and view the target Cued Speech gesture generated according to the target input. It includes various forms such as desktop computers, laptops, PDAs (Personal Digital Assistants), mobile phones, in-vehicle terminals, home theater terminals, and dedicated terminals. In addition, it can be a single device or a collection composed of multiple devices. For example, multiple devices are connected through a local area network and share a display device for collaborative work, jointly constituting a terminal. Object terminal 110 can also communicate with Internet 120 in a wired or wireless manner to exchange data.

[0144] The gateway 130 is also known as an internetwork connector or protocol converter. The gateway 130 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, or even with completely different architectures, the gateway 130 is a translator. At the same time, the gateway 130 can also provide filtering and security functions. The messages sent by the object terminal 110 to the server 140 need to be sent to the corresponding server 140 through the gateway 130. The messages sent by the server 140 to the object terminal 110 also need to be sent to the corresponding object terminal 110 through the gateway 130.

[0145] The server 140 refers to a computer system that can provide a service for generating target prompt voice postures to the object terminal 110. Compared with the object terminal 110, the server 140 has higher requirements in terms of stability, security, performance, etc. The server 140 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. The server 140 can also communicate with the Internet 120 by wired or wireless means to exchange data.

[0146] The embodiments of the present disclosure can be applied in a variety of scenarios, such as Figure 2 shown, where it is necessary to convert the input target voice or text into a prompt voice posture. For example, scenarios for assisting hearing-impaired people to communicate with others, scenarios for non-hearing-impaired people to convey information to hearing-impaired people, scenarios for educating hearing-impaired people, etc.

[0147] (1) Scenarios for assisting hearing-impaired people to communicate with others

[0148] Hearing-impaired people cannot hear speech, so when communicating with hearing-impaired people, information is mainly transmitted through vision. In this scenario, if the hearing-impaired person communicates with others, if others communicate in the form of voice, the hearing-impaired person cannot obtain the information conveyed by others. At this time, the terminal accurately converts the voice messages sent by others or the recorded content of others' speech into corresponding prompt voice gestures, thereby visualizing the voice of others. In this way, the hearing-impaired person can accurately obtain the information contained in the voice of others. For example, in some online group chat scenarios, since there are a large number of people chatting in the group, everyone has their own chatting habits. Some people use voice messages to communicate for their own convenience or other reasons. At this time, it is difficult for the hearing-impaired to obtain the information that these people want to convey. This will cause the hearing-impaired to lose some important information during the communication and be unable to communicate with others smoothly. The embodiment of the present disclosure can solve this problem. It can be applied to the plug-in of the chat tool. When the hearing-impaired receives the voice message sent by others, the voice message sent by others is used as the target input, and the corresponding posture description is generated according to the pinyin of each word in the voice message. The corresponding prompt voice posture can be generated to visualize the voice message sent by others. At this time, the hearing-impaired can obtain the information conveyed by others. For example, when a hearing-impaired person receives a voice message saying "hello everyone," they use this voice message as the target input. After receiving the target input, the terminal analyzes it and finds that the voice message contains the three characters "大", "家", and "好". The pinyin for these three characters is "dà", "jiā", and "hǎo" respectively. The corresponding gestures for these three pinyins in the prompt voice message are "extend your index finger and place your hand next to your face", "extend your thumb, index finger, and middle finger, spread your thumb and index finger apart, put your index finger and middle finger together, place your hand next to your mouth, then extend your index finger and middle finger and spread them apart, place your hand next to your face", and "extend your middle finger, ring finger, and pinky finger together, place your hand on your neck". The terminal will then generate the above three prompt voice gestures in sequence. When the hearing-impaired person sees the prompt voice gestures, they will understand that the received voice message is "hello everyone". In addition, the hearing-impaired person can also convert the text messages received during the communication process into corresponding prompt voice gestures. In this way, compared to directly browsing text messages, the hearing-impaired person can more intuitively understand the messages sent by others.

[0149] (2) Scenarios where non-hearing-impaired people convey information to hearing-impaired people

[0150] Since non-hearing-impaired people do not have hearing impairments themselves, most of them do not specifically learn prompt voices or sign languages. When non-hearing-impaired people need to communicate with hearing-impaired people, it is very difficult to convey the content they want to express to hearing-impaired people. At this time, through the method proposed in the embodiments of the present disclosure, non-hearing-impaired people can input voices into the terminal and then generate corresponding prompt voice gestures by the terminal. At this time, non-hearing-impaired people can directly give the page of the terminal to hearing-impaired people to view, or non-hearing-impaired people can imitate and learn after viewing the prompt voice gestures generated by the terminal, and then gesture corresponding prompt voice gestures, so as to convey the information they want to express to hearing-impaired people. Specifically, such communication scenarios can include scenarios such as customer service responding to the inquiries of hearing-impaired people, online education, and news broadcasts. It can be understood that when customer service needs to respond to the inquiries of hearing-impaired people, customers can input the voices or texts of their answers to the inquiries of hearing-impaired people into the terminal, and then generate corresponding prompt voice gestures, and the customer service then gestures the corresponding prompt voice gestures to hearing-impaired people, so as to respond to the questions of hearing-impaired people. In the news broadcast scenario, since news broadcasts are for the public, it is necessary to consider hearing-impaired people among the public. At this time, the terminal can collect the voice of the news anchor as the target input, generate corresponding prompt voice gestures, and then insert the generated prompt voice gestures into the picture of the news broadcast. For example, a small window can be set in the lower right corner of the picture, and the generated prompt voice gestures are played in the small window. In this way, the content described by the news anchor can be made visual, and hearing-impaired people can receive and understand the content described by the news anchor.

[0151] (3) Scenarios of Educating Hearing-Impaired People

[0152] In an educational scenario, a course instructor generally uses a PPT or other documents as an aid and mainly relies on the instructor's own oral narration to explain knowledge. At this time, since hearing-impaired people can only see the content of the PPT and cannot hear the voice of the instructor teaching knowledge, this will cause hearing-impaired people to not be able to learn the knowledge taught by the instructor well. Especially when the instructor teaches professional knowledge, if the instructor does not record a complete knowledge system in auxiliary documents such as PPTs, the lack of a small amount of information may cause hearing-impaired people to not understand the knowledge taught by the instructor. For example, when the instructor is explaining the derivation of some mathematical formulas and only records the rough derivation results on the PPT without recording the derivation process, and even does not standardize the meanings of some symbols in the formula, then, in the case of not hearing the voice of the instructor teaching, hearing-impaired people may not be able to understand how this derivation result is obtained or the meaning of this derivation result itself. Moreover, since there are some professional terms in the formula, such as "differential", "integral", "Fourier transform", etc., the meanings of these professional terms cannot be well expressed through traditional sign language. At this time, how to accurately convey this professional knowledge to hearing-impaired people has become a problem to be solved. And through the method proposed in the embodiments of the present disclosure, taking the voice of the instructor's oral narration of knowledge as the target input and accurately generating the corresponding prompt voice gesture, the knowledge taught by the instructor can be conveyed to hearing-impaired people accurately down to each word.

[0153] General description of embodiments of the present disclosure

[0154] According to an embodiment of the present disclosure, a method for generating a prompt voice gesture is provided.

[0155] The method for generating a prompt voice gesture is a method for hearing-impaired people to convert voice messages and / or text messages input by non-hearing-impaired people and received by hearing-impaired people into corresponding prompt voice gestures and display them on the interface of the object terminal 110. The prompt voice gesture is a gesture corresponding to the input voice message and / or text message, and the gesture consists of two elements: the position of the hand and the shape of the hand.

[0156] In current methods for generating prompt voice gestures, the rule of mapping text to gestures is often solidified in the model, and then the text or voice is input into the model to obtain prompt voice gestures, etc., or the method of generating prompt voice gestures through a machine learning model. But no matter which method, there is no semantic connection between the text or voice and the generated gesture. For example, when representing "tree", a gesture of spreading two fingers and placing them at the neck is generated, and the semantics of this gesture has no connection with the semantics of "tree" at all, and the mapping is completely achieved by mechanical rules. Therefore, the generated gesture is rigid, with poor accuracy and insufficient fineness.

[0157] This method can be applied to such as Figure 2Scenarios that convert the input voice or text into a prompt voice gesture to assist hearing-impaired people in communication, scenarios that convey information to hearing-impaired people, and scenarios that educate hearing-impaired people can also be used in other scenarios in the daily life of hearing-impaired people.

[0158] As Figure 3 shown, the method for generating a prompt voice gesture according to an embodiment of the present disclosure may include:

[0159] Step 310: Obtain a target input, where the target input includes at least one of a target text and a target voice;

[0160] Step 320: Add the target input to the guiding language input large language model to obtain a gesture description describing the target prompt voice gesture;

[0161] Step 330: Generate a gesture description guiding vector corresponding to the gesture description;

[0162] Step 340: Guide the diffusion model by using the gesture description guiding vector to enable the diffusion model to generate a target prompt voice gesture vector;

[0163] Step 350: Generate a target prompt voice gesture based on the target prompt voice gesture vector.

[0164] The following will describe steps 310-350 in detail.

[0165] The target input in step 310 can be at least one of a target voice and a target text. For example, the target input includes only the target voice, or only the target text, or both the target voice and the target text. The target voice here can refer to voice information in various communication processes, such as the words spoken by one party in an offline face-to-face communication, a voice message sent by one party in an online communication, the words spoken by a host during a live broadcast, the dubbing of a video, etc. Similarly, the target text can also refer to text information in various communication processes, such as chat text in an online communication, the alphanumeric text of a video, etc. It should be noted that when the target input includes both the target voice and the target text, the target voice and the target text should correspond to the same expression content. For example, if the target voice input is the sentence "Everyone is good", then the target text is also composed of the three characters "big", "home", and "good". It is understandable that when the target input only includes one of the target text and the target voice, another form of target input can also be generated based on the obtained target input. For example, if the obtained target input only includes the target text, the target text can be read aloud by the text reading model to obtain the corresponding target voice. Similarly, when the obtained target input only includes the target voice, the target text corresponding to the target voice can be generated through speech recognition technology (Automatic Speech Recognition, ASR). In this way, the two forms of target input, voice and text, can be completed, so that the corresponding prompt voice gestures can be generated more finely and accurately in the future.

[0166] It can be understood that the target speech or target text may contain multiple characters, and in the prompt speech, the generated prompt speech posture is determined based on the pronunciation of the input characters. For the target speech or target text containing multiple characters, it is necessary to generate the corresponding prompt speech posture for each character in the target speech or target text in turn, that is, for the target speech or target text containing multiple characters, it is necessary to execute a round of steps 310-350 for each character in the target speech or target text in turn. For example, the sentence "Hello everyone" includes three characters, so it is necessary to execute three rounds of steps 310-350 to generate the prompt speech postures corresponding to the three characters "big", "home", and "good" in turn. It can be understood that in order to improve the coherence of the prompt speech posture, when the target input includes multiple characters, the first generated prompt speech posture vector can be used as the subsequent vector to be diffused. For example, when it is necessary to generate the prompt speech postures of the three characters "big", "home", and "good" in turn, the diffusion model will first generate the improved speech posture vector of the character "big", and then the target prompt speech posture vector corresponding to the character "big" can be used as the vector to be diffused for the character "home".

[0167] When a part of the above-mentioned prompting voice gesture generation method is implemented on the terminal 110 and another part is implemented in the server 140, the ways to obtain the target input may include obtaining the target input entered by the user through the terminal, receiving the target input sent by others through the terminal, and so on.

[0168] Obtaining the target input entered by the user through the terminal means recording the voice spoken by the user through the microphone of the terminal or obtaining the text typed by the user through the keyboard. When a non-hearing-impaired person needs to convey a message to a hearing-impaired person, since the hearing-impaired person cannot hear the voice, then the non-hearing-impaired person can record their own spoken voice or type text through the terminal, and then generate the corresponding prompting voice gesture based on the recorded voice and / or the typed text. In this way, the non-hearing-impaired person can convey the content they want to express to the hearing-impaired person by showing the prompting voice gesture to the hearing-impaired person or by learning and imitating the prompting voice gesture themselves and showing it to the hearing-impaired person. For example, the non-hearing-impaired person can be a news anchor. Since news broadcasts are for the public, when broadcasting, it is necessary to consider how to convey the content told by the news anchor to the hearing-impaired people in the public. At this time, the voice of the news anchor during the news broadcast can be recorded through the terminal as the target input, and the corresponding prompting voice gesture can be generated and inserted into the news broadcast screen, so as to convey the news content to the hearing-impaired people.

[0169] Receiving the target input sent by others through the terminal means receiving the target voice and / or target text sent by others through the network through the terminal. When a hearing-impaired person communicates with others, since the hearing-impaired person cannot hear the voice, if others send messages to the hearing-impaired person by sending voice messages and other means, the corresponding prompting voice gesture can be generated through the method of this embodiment of the present disclosure, so that the hearing-impaired person can obtain the content in the voice message sent by others. For example, in the scenario of an online group chat, a netizen sends a voice message. At this time, the hearing-impaired person can use the voice message sent by the netizen as the target input, and then generate the corresponding prompting voice gesture through the terminal, so that the hearing-impaired person can obtain the content in the voice message sent by others. Of course, since the hearing-impaired person is also more accustomed to using the method of prompting voice to communicate in daily life, for the text content sent by others, the corresponding prompting voice gesture can also be generated for reading, which is more in line with the daily life habits of the hearing-impaired person.

[0170] The large language model in step 320 can be obtained after fine-tuning the existing Generative Pre-Trained Transformer (GPT) model, such as after fine-tuning ChatGLM. Of course, it can also be obtained after fine-tuning other open-source large language model frameworks. In this regard, it is not limited in this embodiment. The guiding statement is a statement used to guide the large language model to generate an answer in a specific direction for the target input. It can be understood that as a text generation model, for a given input, the large language model will generate corresponding text. However, the direction of the text generated by the large language model is not fixed. For example, when the input is the sentence "Hello everyone", the answer of the large language model may be a greeting rather than the gesture description of the sentence "Hello everyone" in the prompt voice. Based on this, in this embodiment, by setting the guiding statement, the large language model is guided to generate the corresponding gesture description of the target input in the prompt voice. Specifically, a sample of the guiding statement can be "For <s>, what is the corresponding gesture description in the prompt voice?”, where <s>That is the target input. By adding the target input to the guiding text and then inputting it into the large language model, the large language model can perceive that the specific text generation task is to generate the pose description corresponding to it in the prompt speech. Further, since the prompt speech pose generally includes hand shape and hand position, the guiding text can be further refined to "For <s>What is the corresponding hand shape and the corresponding hand position in the prompt voice? Based on this guiding statement, the large language model can further perceive that the specific text generation task is to generate a description of the corresponding hand shape and hand position of the target input in the prompt voice; the guiding statement can also be a masked statement, such as "For <s>, the corresponding hand shape in the prompt voice is , the corresponding hand position is ”, where And is the masked part. Input this masked statement into the large language model to let the large language model predict the text at the masked position, so as to better control the direction of the text generated by the large language model and enable the large language model to output the corresponding hand shape and hand position according to the target input. The gesture description is used to describe the hand shape and hand position corresponding to the target input in the prompt speech, and this gesture description is determined based on the encoding rules of the prompt speech and the target input. Refer to Figure 5 For example, the target input is a "tree" character. According to the coding rules of the prompt voice, the pinyin of the character "tree" is shu, and its vowel is u[u]. The corresponding hand position is to put the hand on the neck position, and the consonant is The corresponding hand shape is "extend two fingers without closing them". Then, when the target input is the word "tree", the posture description output by the large language model is "place your hand on your neck, extend two fingers without closing them". It can be seen that the posture corresponding to this posture description and the word "tree" are obviously not related at the semantic level. If the word "tree" is directly input into the model for generating prompt voice posture, since the two have no semantic connection, the model cannot generate the corresponding prompt voice posture based on the word "tree" well. Through step 320 of this embodiment, the target input can be converted into a posture description that has a semantic connection with the corresponding posture. Therefore, the corresponding prompt voice posture can be accurately generated based on the posture description.

[0171] Next, in step 330, a posture description guidance vector corresponding to the posture description is generated. The posture description guidance vector is a feature vector obtained by encoding the posture description. Various encoding methods exist for encoding posture descriptions. As an example, an encoding method for generating a model using the first vector will be discussed below and described in detail.

[0172] The diffusion model in step 340 is a generative model that can gradually denoise the initial input by predicting noise in multiple steps, thereby generating a specific output. Specifically, in this embodiment, the diffusion model has been trained. It learns how to predict the corresponding noise according to the given pose description guidance vector, and then gradually denoise the speech pose vector to be diffused according to the predicted noise, so as to generate the target speech pose vector. The following describes the process of denoising the diffusion model. First, a speech pose vector to be diffused is given. This speech pose vector to be diffused is a random vector that does not carry any information at all. Specifically, the speech pose vector to be diffused is a random vector that follows a standard Gaussian distribution. It can be understood that for any vector, by adding random noise to the vector multiple times, as long as the number of times of adding random noise is sufficient, then this given vector can be converted into a vector that follows a standard Gaussian distribution. Therefore, it can be considered that the speech pose vector to be diffused is obtained after adding random noise to any pose vector multiple times, that is, it can be considered that the speech pose vector to be diffused can be a vector obtained after adding random noise to the target speech pose vector to be generated multiple times. Then, by gradually removing the added noise from the speech pose vector to be diffused, the speech pose vector to be diffused can be converted into the target speech pose vector to be generated. And this process of denoising the speech to be diffused is actually a process of gradually adding detailed information to the speech pose vector to be diffused, thereby gradually generating the corresponding vector. Based on this, in this embodiment, the diffusion model is guided by the pose description guidance vector to predict noise, and the speech pose vector to be diffused is denoised according to the predicted noise, so as to gradually add detailed information related to the pose description to the speech pose vector to be diffused, and finally generate the target speech pose vector corresponding to the pose description. The target speech pose vector here is a vector composed of the characteristic information of the coordinates of each human key skeletal point in the target pose.

[0173] The target speech pose in step 350 is the image data corresponding to the target speech pose vector. Specifically, the target speech pose vector generated by the diffusion model in step 340 carries the characteristic information of the coordinates of each human key skeletal point in the target speech pose to be finally generated. By decoding the target speech pose vector, the coordinates of each human key skeletal point are obtained, and the corresponding target speech pose is obtained.

[0174] In the embodiments of the present disclosure, by adding the target input to the guiding language and then inputting it into the large language model, the guiding language is used to guide the large language model to generate a posture description of the target input in the prompt voice. Thus, the target input that clearly has no semantic connection with the posture of the target prompt voice is converted into a posture description that has a strong semantic connection with the posture of the target prompt voice. Then, the encoder captures the semantic information in the posture description, encodes the posture description into a posture description guiding vector with a high similarity to the corresponding posture, and guides the diffusion process of the diffusion model through the posture description guiding vector. During the diffusion process, information related to the posture description is gradually added to the prompt voice posture vector to be diffused, and finally, the target prompt voice posture corresponding to the target input is accurately generated. At the same time, the characteristics of the diffusion model gradually adding relevant detailed information of the posture description during the diffusion process are utilized to improve the fineness of the generated target prompt voice posture.

[0175] Since the details of steps 310 and 350 have been fully described above, the details of steps 320 - 340 will be described in detail below.

[0176] Detailed description of step 320

[0177] In one embodiment, to use the large language model to generate a posture description describing the target prompt voice posture according to the target input, the large language model can first be made to perceive the conversion rule of the prompt voice, and this conversion rule indicates how to encode the vowels and consonants of the target input through hand shapes and hand positions.

[0178] Refer to Figure 7 , in one embodiment, step 320 may include:

[0179] Step 710, obtaining the overall rule for converting the target input into a posture description;

[0180] Step 720, adding the target input to the guiding language and the overall rule, inputting it into the large language model, and obtaining a posture description.

[0181] The overall rule in step 710 refers to the rule of how to encode consonants using hand shapes and vowels using hand positions in the prompt voice. Refer to Figure 5 In part (a), it gives the general rules for Chinese prompt voices. In Chinese prompt voices, eight hand shapes are used to encode consonants, namely "extend the index finger", "extend the index finger and middle finger, with the index finger and middle finger together", "extend the middle finger, ring finger and little finger, with the three fingers together", "extend the four fingers except the thumb, with the four fingers together", "extend the five fingers, with the thumb and index finger separated and the four fingers except the thumb together", "extend the index finger and middle finger, with the index finger and middle finger separated", "extend the thumb, index finger and middle finger, with the index finger and thumb separated and the index finger and middle finger together", "extend the index finger and middle finger, with the index finger and middle finger separated". Each hand shape corresponds to 2 - 3 consonants respectively. For example, the hand shape of "extend the index finger" corresponds to the three consonants d, p, and zh. Additionally, five hand positions, namely the eyes, beside the face, the mouth, the neck, and the throat, are used to encode vowels. Each hand position corresponds to 3 - 4 vowels. For example, placing the hand near the eyes corresponds to the three vowels an, e, and o.

[0182] In step 720, after obtaining the general rules, the target input is added to the general rules and the guiding language and then input into the large language model. In this way, with the prompts of the guiding language and the general rules, the large language model can perceive that the task to be processed is to convert the target input into the corresponding pose description according to the input general rules. Specifically, a prompt statement template can be constructed based on the guiding language and the general rules, such as "Please generate" <s>"The description of the corresponding gesture in the Chinese prompt voice, the general rule of the conversion between text and gesture in the Chinese prompt voice is: when the input pinyin vowel is d, p, zh, the corresponding hand shape is to extend the index finger; when the input pinyin vowel is k, z, q, the corresponding hand shape is to extend the index finger and middle finger, and put the index finger and middle finger together...", among which <s>This is the target input. For example, if the target input is "tree," then after adding the target input to the prompt template composed of the guide words and the overall rules, we get the corresponding gesture description of "Please generate "tree"" in the Chinese prompt voice. The overall rule for converting characters and gestures in the Chinese prompt voice is: when the input pinyin vowels are d, p, zh, the corresponding hand shape is to extend the index finger; when the input pinyin vowels are k, z, q, the corresponding hand shape is to extend the index finger and middle finger, and then put the index finger and middle finger together..." Such a task text is then input into the large language model. The large language model can perceive that the task to be processed is to generate the gesture description of the word "tree" according to the given overall rules. After that, the large language model performs the task and obtains the gesture description "Extend two fingers without closing them, and place your hand near your neck."

[0183] In this embodiment, by obtaining the overall rules for converting the target input into the posture description in the prompt voice, a prompt sentence template is constructed based on the overall rules and the guide words, the target input is inserted into the prompt sentence template to form a task text, and then the task text is input into the large language model. Through the prompt template composed of the guide words and the overall rules, the large language model can accurately perceive that the task to be processed is to generate a posture description of the target prompt voice posture corresponding to the target input according to the overall rules. As a result, the large language model can accurately generate the corresponding posture description based on the target input.

[0184] It is understandable that for a large language model, even if the input sentences are semantically consistent, but there are differences in the descriptions, the output of the large language model may also be different. Based on this, in order for the large language model to accurately generate a posture description of the target prompt voice posture based on the target input, a prompt template consisting of guide words and overall rules needs to be trained and optimized.

[0185] Reference Figure 8 , the guide words and overall rules are jointly trained in the following way:

[0186] Step 810: Acquire a first sample set, where the first sample set includes a first sample text and a gesture label corresponding to the first sample text;

[0187] Step 820 , adding a guide word and overall rules to each first sample text in the first sample set, inputting the result into the large language model, and obtaining a prediction result corresponding to the first sample text;

[0188] Step 830: Calculate a first loss function based on the prediction result and the posture label corresponding to the first sample text;

[0189] Step 840, if the first loss function is less than the first threshold, stop training the guiding language and the overall rules; otherwise, adjust the guiding language and the overall rules, and return to the step of adding each first sample text in the first sample set to the guiding language and the overall rules, inputting them into the large language model, and obtaining the prediction results corresponding to the sample texts.

[0190] The first sample set in step 810 is a pre-constructed data set for training the guiding language and the overall rules, which contains multiple first sample texts and the pose labels corresponding to each first sample text. The first sample text can be pre-collected text data containing a piece of text, and the pose label is the prompt voice pose corresponding to the text in the first sample text. For example, if the first sample text includes the word "tree", then the pose label corresponding to this first sample text is the pose of "extending the index finger and the middle finger, with the index finger and the middle finger separated". It can be understood that the pose label can be represented in the form of a vector.

[0191] The prediction result in step 820 refers to inputting the first sample text, the guiding language, and the overall rules into the large language model to generate a pose description, then encoding the pose description into a corresponding pose description vector, and based on the pose vector matched by the pose description vector. Specifically, in step 820, first add the first sample text to the prompt template constructed according to the guiding language and the overall rules to form a task text corresponding to the first sample text. This task text is used to represent the task of generating a pose description corresponding to the first sample text according to the overall rules. Input this task text into the large language model, and the large language model generates a pose description. Then, through the first vector generation model, encode the pose description generated by the large language model into a vector to obtain the corresponding pose description vector. Then, based on this pose description vector, match the corresponding pose vector as the corresponding prediction result. It should be noted that the trained first vector generation model can encode the input pose description text into a pose description vector with a relatively high similarity to the pose vector of the corresponding pose. Based on this, the pose description vector can be used to match the corresponding pose vector in the high-dimensional feature space as the prediction result. In some embodiments, because the similarity between the pose description vector generated by the first vector generation model and the pose vector of the corresponding pose is relatively high, the pose description vector can also be directly regarded as the pose vector of the corresponding pose, that is, directly use the pose description vector as the prediction result.

[0192] The first loss function in step 830 is used to measure the difference between the predicted result corresponding to the first sample text and the pose label. It can be understood that during the training process, since the prompt template constructed according to the guiding language and the overall rules is not perfect, at this time, the pose description output by the large language model based on the first sample text, the guiding language, and the overall rules may not describe the prompt voice pose corresponding to the first sample text. Therefore, it is necessary to calculate the first loss function to judge the difference between the pose matched according to the pose description output by the large language model and the pose label corresponding to the first sample text. It can be understood that it is expected that the predicted result obtained by the large language model based on the first sample text, the guiding language, and the overall rules is as similar as possible to the pose label corresponding to the first sample text. The higher the similarity between the two, the more perfect the guiding language and the overall rules are. Therefore, the first loss function can be determined by calculating the similarity between the predicted result and the pose label. Specifically, the cosine similarity between the predicted result and the pose label can be calculated, and then the corresponding first loss function can be determined based on the negative correlation function of the cosine similarity. For example, 1 minus the cosine similarity is used as the first loss function. It can be understood that the first sample set can include multiple first sample texts, and then the corresponding first loss function is the sum of the negative correlation functions of the cosine similarities between the predicted results and the pose labels corresponding to each first sample text. It should be noted that the trained first vector generation model can encode the input pose description into a pose description vector with a relatively high similarity to the pose vector of the corresponding pose. Therefore, the similarity between the pose description vector and the pose label can also be directly calculated to judge whether the predicted result obtained by the large language model based on the first sample text, the guiding language, and the overall rules matches the true pose label. The training process of the first vector generation model will be described in detail later and will not be elaborated here.

[0193] The first threshold in step 840 can be a preset threshold. It can be understood that in step 830, by using the negative correlation function of the cosine similarity between the predicted result corresponding to the first sample text and the pose description label as the first loss function, when the first loss function is greater than the set first threshold, it means that the cosine similarity between the predicted result corresponding to the first sample text and the pose description label is relatively low. At this time, it means that the guiding language and the overall rules cannot well enable the large language model to output the pose description corresponding to the first sample text, and then the guiding language and the overall rules need to be adjusted until the value of the first loss function is less than the first threshold. When the first loss function is less than the set first threshold, it means that the cosine similarity between the predicted result corresponding to the first sample text and the pose description label is relatively high, that is, the pose description generated by the large language model based on the first sample text, the guiding language, and the overall rules is already very close to the pose description label corresponding to the first sample text. At this time, the guiding language and the overall rules are relatively perfect, so the adjustment of the guiding language and the overall rules can be stopped.

[0194] In this embodiment, based on the prediction result generated by the large language model according to the first sample text, the guiding language, and the overall rules, and the pose description label corresponding to the first sample text, the first loss function is calculated. Based on the first loss function, it is determined whether the guiding language and the overall rules need to be adjusted. When the first loss function is less than the set first threshold, it is considered that the guiding language and the overall rules can guide the large language model to accurately generate the pose description. At this time, the adjustment of the guiding language and the overall rules is stopped. In this embodiment, by continuously adjusting the guiding language and the overall rules, the large language model can accurately generate the pose description corresponding to the input target text under the guidance of the guiding language and the overall rules, thereby improving the accuracy of generating the pose description.

[0195] Detailed description of step 330

[0196] As mentioned above, in the case where the target input includes both the target text and the target voice, referring to Figure 6 , the above step 320 may include: step 610, inputting the target text into the guiding language to the large language model to obtain a pose description describing the target prompt voice pose. This pose description is a text data. At this time, referring to Figure 4 and Figure 9 , step 330 may include: step 910, inputting the pose description into the first vector generation model to obtain a pose description guiding vector;

[0197] The first vector generation model in step 910 is trained. For the input pose description, the first vector generation model can generate a pose description guiding vector with a relatively high vector similarity to the pose corresponding to the pose description, so that the diffusion model can be guided by the pose description guiding vector later to accurately generate the pose corresponding to the pose description. Specifically, in this embodiment, the encoder of Transformer is used as the first vector generation model, which includes multiple identical encoding layers, and each encoding layer includes a self-attention layer and a feed-forward neural network. When the pose description is input into the first vector generation model, first, the pose description will be segmented into multiple words through the word segmentation technology. For example, if the pose description is "Put the hand on the neck position, extend the index finger and middle finger, and the index finger and middle finger are spread apart", then after input into the first vector generation model, this sentence will be first segmented into a series of words such as "Put", "hand", "on", "neck position", "extend", "index finger", "and", "middle finger", "index finger", "and", "middle finger", "spread apart", and this series of words will be represented as a sequence. Then, each word will be converted into a corresponding word vector. Specifically, here, the one-hot vector representation method or the word-embedding method can be used to convert the word into the corresponding word vector. At the same time, according to the position of each word in the sequence, the corresponding position encoding of each word is determined. The position encoding here is also a vector, and the position encoding is added to the word vector of the corresponding word, so as to add the position information to the word vector of each word; then, each word vector added with the position information is input into the encoder. In each encoding layer, first, according to the three weight matrices of the self-attention layer, a linear transformation is performed on the input word vector to obtain the corresponding query vector, key vector, and value vector of each word. Specifically, the linear transformation can be performed through the following formula:

[0198] Q = X * W Q ;

[0199] K = X * W K ;

[0200] V = X * W V ;

[0201] In the formula, X represents the word vector of the word, Q represents the query vector corresponding to the word, K represents the key vector corresponding to the word, V represents the value vector corresponding to the word, W Q 、W K 、W V They are three weight matrices used when calculating the query vector, key vector, and value vector respectively. These three weight matrices are parameters of the self-attention layer and are determined during the training of the encoder. They will not be elaborated here. The query vector represents the relevance of the corresponding word to other words in the pose description, that is, the similarity between the word and other words in the pose description. The key vector represents the key information of the corresponding word, and the value vector is used to represent the actual information of each word and the final output. Specifically, in the self-attention layer, by calculating the dot product of the query vector of each word and the key vector of each other word, the attention score of the word can be obtained. The higher this attention score, the greater the influence of the word on other words in the pose description, and the higher the importance of the word in the pose description. Then, the softmax function is used to normalize the attention score to ensure that the sum of the attention scores of all words in the pose description is 1, thereby converting the attention score into the corresponding attention weight. Then, based on this attention weight, the value vectors of each word are weighted, and thus the final vector representation of each word is obtained. By weighted summing the value vectors of all words according to the attention weight, the output of the pose description in this self-attention layer is obtained. After that, in each encoding layer, a feed-forward neural network is also used to perform non-linear processing on the output of the self-attention layer, and the feed-forward neural network is used to extract the feature information therein, and thus the output of each encoding layer is obtained. Through this design of the self-attention mechanism and the feed-forward neural network, each encoding layer can perform deeper abstraction and understanding of the input and capture the features in the input. Taking the output of this encoding layer as the input of the next encoding layer and repeating the above steps, the output of the last encoding layer is the pose description guidance vector generated according to the pose description. This multi-encoding layer structure can make the finally generated pose description guidance vector carry the semantic feature information of the pose description.

[0202] Since the encoder includes multiple encoding layers, and except for the first encoding layer, other encoding layers take the output of the previous encoding layer as the input. As the depth of the encoder increases, this easily leads to the loss of a lot of details when the input pose description is passed to the subsequent encoding layers. Therefore, for each encoding layer, a residual connection is used to add and normalize the input and output of each encoding layer and then input it into the next encoding layer, so that the input feature information can be better passed to the network with a higher depth, making the finally generated pose description guidance vector contain both the important feature information in the pose description and the details in the pose description.

[0203] In the above embodiments, the specific process of generating the pose description guidance vector using the first vector generation model is disclosed. However, since the pose description itself is a text data generated by a large language model, and the target prompt voice pose vector to be finally generated is a pose vector representing the coordinates of joint bone points such as the human hand, the pose description guidance vector generated based on the text data and the target prompt voice pose vector to be finally generated actually belong to different modal data. To use the pose description guidance vector to guide the diffusion process of the diffusion model, it is first necessary to ensure that there is a high similarity between the pose description guidance vector generated based on the pose description and the feature vector of the corresponding pose. Based on this, it is necessary to train the first vector generation model so that the pose description guidance vector generated according to the input pose description is as close as possible to the feature vector of the corresponding pose. Thus, the diffusion model can be better guided by the pose description guidance vector to generate the corresponding target prompt voice pose.

[0204] As Figure 10 shown, according to an embodiment of the present disclosure, the first vector generation model can be trained in the following manner:

[0205] Step 1010, obtain a sample pair set. The sample pairs in the sample pair set include a plurality of first sample pairs and a plurality of second sample pairs. The first sample pairs include matching sample poses and sample pose descriptions, and the second sample pairs include non-matching sample poses and sample pose descriptions.

[0206] Step 1020, convert the sample pose description into a sample pose description vector through the first vector generation model, and convert the sample pose into a sample pose vector through the third vector generation model.

[0207] Step 1030, jointly train the first vector generation model and the third vector generation model to make the distance between the sample pose description vector and the sample pose vector of the first sample pair smaller, and make the distance between the sample pose description vector and the sample pose vector of the second sample pair larger.

[0208] The sample pair set in step 1010 can be pre-constructed. When constructing the sample pair set, a video can be recorded when the prompt voice user communicates through the prompt voice, and the human body image of the user can be intercepted from the video. It should be noted that the human body image needs to contain at least the hand shape of the user and the position information of the hand relative to the user's head. Then, models such as HRNet, OpenPose, and PoseNet are used to extract the key skeletal points of the human body from the image, especially the key skeletal points of the five fingers of the hand. Then, these key skeletal points are connected to form a skeleton, and this skeleton is the required sample pose. At the same time, the sample pose corresponding to each human body image is labeled. For example, for a sample pose of stretching out the index finger and thumb, the index finger and thumb are spread apart, and the hand is placed near the neck, then the corresponding label is "stretching out the index finger and thumb, the index finger and thumb are spread apart, and the hand is placed near the neck", and this label is the sample pose description matching the sample pose, so multiple first sample pairs are obtained. At the same time, the sample pose in each first sample pair is paired with the sample pose description in other first sample pairs. For example, for a sample pose of stretching out the index finger and pointing the index finger at the eye, it is paired with the sample pose description of "stretching out the index finger and thumb, the index finger and thumb are spread apart, and the hand is placed at the throat position", so multiple pairs of unmatched sample poses and sample pose descriptions, that is, second sample pairs, are obtained. In this way, the construction of the sample pair set is completed. The sample pair set can be expressed as where E represents the sample pair set, B represents the number of the first sample pairs in the sample pair set, represents the sample pose description in the i-th first sample pair, represents the sample pose in the i-th first sample pair, and the first sample pair can be expressed as The second sample pair can be expressed as

[0209] It can be understood that the human body image can also be obtained by collecting existing images on the Internet. The process of constructing the corresponding first sample pairs and second sample pairs based on the collected human body images refers to the above steps and will not be elaborated here. It should be noted that in this embodiment, the first vector generation model is ultimately used to generate a pose description related to the target prompt voice pose, and the target prompt voice pose is mainly composed of two elements: hand shape and hand position. Then, the corresponding pose description is also used to describe the hand shape and hand position. Based on this, during the training of the first vector generation model, the labeled sample pose description is also used to describe the hand shape and hand position in the corresponding sample pose.

[0210] The sample pose description in step 1020 refers to all the sample pose descriptions in the sample pair set obtained in step 1010, and the sample pose is all the sample poses in the sample pair set obtained in step 1010. The first vector generation model is the first vector generation model mentioned in step 330 above. The encoder in Transformer can be used. The process of converting the sample pose description into a sample pose description vector through the first vector generation model refers to the process of generating a pose description guidance vector according to the pose description in step 330 of the above embodiment. Just replace the input from the pose description with the sample pose description, and details will not be repeated here. The third vector generation model is a pose vector encoder. Specifically, after extracting key skeleton points from a human body image through models such as OpenPose, HRNet, and PoseNet, the coordinates of these key skeleton points can be directly output. Each sample pose is composed of multiple key skeleton points. Based on the coordinates of all the key skeleton points in each sample pose, the vector corresponding to the sample pose can be determined. Specifically, the coordinates of each key skeleton point can be first mapped to the same coordinate system, and then the coordinates of these key skeleton points are converted into vector representations. Finally, the vector representations of all the key skeleton points in the same sample pose are combined to obtain the sample pose vector.

[0211] The distance between the sample pose description vector and the sample pose vector in step 1030 can be determined by calculating the dot product of these two vectors or calculating the cosine similarity of these two vectors. The calculation methods of the dot product and cosine similarity of two vectors are obvious to those skilled in the art and will not be elaborated here. It should be noted that when using the dot product method to calculate the similarity between the sample pose description vector and the sample pose vector, since the sample pose description and the sample pose belong to data of different modalities, there may be a difference in dimension between the two, which will lead to a large gap in the modulus lengths after they are converted into vectors. Therefore, before calculating the similarity between the sample pose vector and the sample pose description vector by the dot product method, the sample pose description vector and the sample pose vector can be normalized, that is, the modulus lengths of the sample pose vector and the sample pose description vector are unified to 1, so as to avoid the influence of the modulus length of the sample pose description vector and the modulus length of the sample pose vector on the calculated similarity, and thus ensure the performance of the first vector generation model obtained by training. It can be understood that the larger the dot product or cosine similarity between two vectors, the more similar these two vectors are, and then the smaller the distance between the corresponding two vectors. On the contrary, the smaller the dot product or cosine similarity between two vectors, the less similar these two vectors are, and then the larger the distance between the corresponding two vectors. Here, the result of subtracting the dot product or cosine similarity of two vectors from 1 can be used as the distance between the two vectors. In this embodiment, the first vector generation model is trained to make the pose description guidance vector generated by the first vector generation model according to the input pose description have a higher similarity with the pose vector of the corresponding pose, and at the same time make the pose description guidance vector and the pose vector of the mismatched pose have a lower similarity. Based on this, in the process of jointly training the first vector generation model and the third vector generation model, the distance between the sample pose description vector and the sample pose vector in the first sample pair should be made smaller, that is, their cosine similarity or dot product is larger, and at the same time make the distance between the sample pose description vector and the sample pose vector in the second sample pair larger, that is, their cosine similarity or dot product is smaller. Refer to Figure 11 For an example given, during the process of training the first vector generation model, the sample pose description is "extend two fingers, do not close them, and place the hand near the neck". The input sample pose is the pose that matches this sample pose description. The two are respectively converted into corresponding sample pose description vectors and sample pose vectors through the first vector generation model and the third vector generation model. During the joint training process, the first vector generation model and the third vector generation model are made to generate vectors with higher similarity for the mutually matching sample pose description and sample pose, and lower similarity for the non-matching sample pose description and sample pose. In this way, the first vector generation model can learn how to generate a pose description vector with a relatively high similarity to the pose vector of the corresponding pose according to the input pose description. After the first vector generation model is trained, the same pose description, that is, "extend two fingers, do not close them, and place the hand near the neck", is input into the first vector generation model, and a pose description guidance vector very close to the pose vector of the corresponding pose can be generated. Inputting this pose description guidance vector into the diffusion model to guide the diffusion process can accurately generate the corresponding target prompt voice pose.

[0212] In this embodiment, two different modalities of data, namely the sample pose description and the sample pose, are respectively input into the first vector generation model and the third vector generation model to generate corresponding vectors. Then, the first vector generation model and the third vector generation model are jointly trained so that the first vector generation model and the third vector generation model generate vectors with high similarity, that is, smaller distances, for the sample pose description and the sample pose in the first sample pair, and generate vectors with low similarity, that is, larger distances, for the sample pose description and the sample pose in the second sample pair. In this way, the unification of different modality inputs is completed. After the first vector surface model is trained, a pose description guidance vector with a relatively high similarity to the pose vector of the corresponding pose can be generated for the input pose description, thereby better guiding the diffusion process of the diffusion model to generate an accurate target prompt voice pose.

[0213] Refer to Figure 12 , in an embodiment, step 1030 may include:

[0214] Step 1210, set the matching label of the first sample pair to 1 and the matching label of the second sample pair to 0;

[0215] Step 1220, through the first vector generation model, based on the sample pose description vector and the sample pose vector of the sample pair, predict the first probability that the sample pose description and the sample pose in the sample pair match;

[0216] Step 1230, through the third vector generation model, based on the sample pose description vector and the sample pose vector of the sample pair, predict the second probability that the sample pose and the sample pose description in the sample pair match;

[0217] Step 1240: Calculate a second loss function based on the matching labels, the first probability, and the second probability of each sample pair, and jointly train a first vector generation model and a third vector generation model based on the second loss function.

[0218] The matching labels in Step 1210 are used to characterize whether the sample poses and the sample pose descriptions in the sample pair match. It can be understood that for the sample pose and the sample pose description in the first sample pair, the two match each other. Therefore, the matching labels of all first sample pairs can be set to 1; similarly, for the sample pose and the sample pose description in the second sample pair, they do not match. Therefore, the matching labels of all second sample pairs are set to 0. The matching labels set here characterize the true matching relationship between the sample pose description and the sample pose.

[0219] The first probability in Step 1220 refers to the probability distribution of the matching between the sample pose description vector predicted by the first vector generation model and each sample pose vector. Specifically, after encoding the sample pose description into a sample pose description vector by the first vector generation model and converting the sample pose into a sample pose vector by the third vector generation model, the first vector generation model will predict the probability of matching between each sample pose description vector and each sample pose vector. In one embodiment, the first probability can be obtained by calculating the similarity scores between the sample pose description and each sample pose and then normalizing these similarity scores. For example, for the i-th sample pose description vector, the first vector generation model can calculate the similarity scores between it and each sample pose vector by means of cosine similarity or dot product, and then normalize the similarity scores between the i-th sample pose description vector and each sample pose vector through the softmax function, which will map these similarity scores to the range of [0,1], and the sum of all the similarity scores after normalization is 1. Thus, the similarity scores after normalization can be regarded as the matching probabilities of the i-th sample pose description vector and each sample pose vector. Combining these matching probabilities into a vector or matrix form, this vector or matrix can be used to characterize the predicted probability distribution of the matching between the i-th sample pose description vector and all sample pose vectors. This predicted probability distribution is the first probability of the i-th sample pose description vector, and the first probability can be denoted as where represents the first vector generation model, represents the i-th sample pose description vector, is the matching probability between the i-th sample pose description vector predicted by the first vector generation model and the first sample pose vector, is the matching probability between the i-th sample pose description vector predicted by the first vector generation model and the second sample pose vector. is the probability that the i-th sample pose description vector predicted by the first vector generation model matches the i-th sample pose vector. It can be understood that the i-th sample pose description vector in the above steps refers to any one of all the sample pose description vectors. In step 1220, the above operations will be performed on each sample pose description vector, and thus the corresponding first probability of each sample pose description vector is obtained.

[0220] The second probability in step 1230 refers to the probability distribution of the matching between the sample pose vectors predicted by the third vector generation model and each sample pose description vector. Specifically, after encoding the sample pose description into a sample pose vector through the sample pose vector generation model and converting the sample pose into a sample pose description vector through the sample pose vector generation model, the sample pose vector generation model will predict the matching probability between each sample pose vector and each sample pose description vector. In one embodiment, the second probability can be obtained by calculating the similarity scores between the sample pose descriptions and each sample pose, and then normalizing these similarity scores. For example, for the j-th sample pose vector, the sample pose vector generation model can calculate the similarity scores between it and each sample pose description vector by means of cosine similarity or dot product, and then normalize the similarity scores between the j-th sample pose vector and each sample pose description vector through the softmax function. After normalization, the matching probabilities between the j-th sample pose vector and each sample pose description vector are obtained. Combining these matching probabilities into a vector or matrix, this vector or matrix can be used to represent the predicted probability distribution of the matching between the j-th sample pose vector and all sample pose description vectors, and this predicted probability distribution can be used as the second probability of the j-th sample pose vector. The second probability can be denoted as It can be understood that the j-th sample pose vector in the above steps refers to any one of all the sample pose vectors. In step 1230, the above operations will be performed on each sample pose vector, and thus the corresponding second probability of each sample pose vector is obtained.

[0221] The second loss function in step 1240 can be in the form of cross-entropy, which can be used to measure the difference between two probability distributions. It can be understood that in step 1210, the matching labels are set based on the true matching relationship between the sample poses and the sample pose descriptions. These matching labels can be used as the true matching probabilities of the sample pose vectors and the sample pose description vectors. Based on these true matching probabilities, the true matching probability distribution of each sample pose description vector can be constructed. It can be understood that since the i-th sample pose description matches the i-th sample pose and does not match other sample poses, in the true matching probability distribution representing the i-th sample pose description vector, only the value of the i-th element is 1, and the values of the other elements are 0. For example, the true matching probability distribution for the first sample pose description can be expressed as Similarly, the true matching probability distribution of each sample pose vector can also be obtained based on the above matching labels. For example, if the second sample pose and the second sample pose description match each other, then the true matching probability distribution of the second sample pose can be expressed as In this way, for each sample pose description vector and each sample pose vector, their true matching probability distributions and predicted probability distributions are obtained. At this time, the second loss function can be calculated through the cross-entropy formula. Specifically, the second loss function can be expressed as:

[0222]

[0223] where L2 is the second loss function, and H() represents the cross-entropy function. represents the sum of the cross-entropy function values calculated for all sample pose vectors and sample pose description vectors. is the true matching probability distribution of the i-th sample pose description vector. is the first probability of the i-th sample pose description vector obtained in step 1220. is the true matching probability distribution of the j-th sample pose vector. is the second probability of the j-th sample pose vector obtained in step 1230. Specifically, the cross-entropy function can be expressed as:

[0224]

[0225]

[0226] It can be understood that since the sample pose descriptions and sample poses in the sample pair set are in one-to-one correspondence, that is, the number of sample pose descriptions and sample poses in the sample pair set is equal. Taking the number of sample poses and sample pose descriptions both being B as an example, then for any It represents the probability of the i-th sample pose description vector matching each sample pose vector. The number of elements in it is the number of sample pose vectors, that is, the number B of sample poses in the sample pair set. Similarly, for any one The number of elements in it is also B, and the cross-entropy function can be expressed as follows:

[0227]

[0228]

[0229] where is the true matching probability distribution of the i-th sample pose description vector The x-th element in, that is, the matching label between the i-th sample pose description and the x-th sample pose; is the predicted probability distribution of the i-th sample pose description vector The x-th element in; is the true matching probability distribution of the j-th sample pose vector The x-th element in, that is, the matching label between the j-th sample pose and the x-th sample pose description; log represents the logarithmic function; B is the number of sample poses, that is, the number of sample pose descriptions. Since the true matching probability distributions are all probability distributions with only one element being 1 and the rest being 0, for example, in only the i-th element takes 1 and the rest of the elements are 0. Similarly, in only the j-th element takes 1 and the rest of the elements are 0. Then substituting into the above formula, we can get:

[0230]

[0231]

[0232] Since the elements in the predicted probability distributions and are all obtained after normalization processing, and their values are in the range of [0,1]. From the above formula, it is not difficult to know that when The larger it is, that is, the higher the similarity between the sample pose vectors and sample pose description vectors generated by the first vector generation model and the third vector generation model for the matching label of 1, that is, the sample pose and sample pose description in the first sample pair, the smaller the value of the corresponding cross-entropy function, and at this time the second loss function is also smaller. Based on this, after calculating the second loss function through cross-entropy, the first vector generation model and the third vector generation model can be jointly trained based on the second loss function, and the parameters of the first vector generation model and the third vector generation model can be continuously optimized to minimize the second loss function, so that the first vector generation model and the third vector generation model can generate pose description vectors and pose vectors with relatively high similarity for mutually matching pose descriptions and pose generations.

[0233] Detailed description of step 340

[0234] In one embodiment, when the target input includes both the target speech and the target text, referring to Figure 6 , step 340 may include:

[0235] Step 620, generating a speech guidance vector corresponding to the target speech;

[0236] Step 630, using the pose description guidance vector and the speech guidance vector to guide the diffusion model, so that the diffusion model generates a target prompt speech pose vector.

[0237] The speech guidance vector in step 620 refers to a vector generated based on the target speech and used to guide the output of the diffusion model. It can be obtained by encoding the target speech. Specifically, referring to Figure 9 , step 620 includes:

[0238] Step 920, inputting the target speech into the second vector generation model to obtain the speech guidance vector.

[0239] The second vector generation model in step 920 can be a pre-trained audio encoder. In one embodiment, the WavLM model can be used as the second vector generation model to extract audio features from the target speech and generate the speech guidance vector. Of course, a pre-trained model that can extract audio features through the method of mel-frequency cepstral coefficients can also be used as the second vector generation model.

[0240] In step 630, when the target input includes both the target voice and the target text, by inputting the target text into the large language model, a pose description of the target prompt voice pose can be obtained. Then, a pose description guidance vector is generated to guide the diffusion model. However, the target text is data in text format and cannot carry complex information such as speech rate. When the target input includes multiple characters, this will result in relatively poor coherence of the generated target prompt voice pose and lack of rhythm in the normal communication process. At this time, by generating a voice guidance vector corresponding to the target voice and jointly guiding the diffusion process of the diffusion model through the voice guidance vector and the pose description guidance vector, the rhythm of the target prompt voice pose can be further adjusted, and the coherence of the prompt voice pose can be improved.

[0241] The diffusion process refers to inputting the guidance vector into the diffusion model, enabling the diffusion model to predict the noise at multiple time steps based on the input guidance vector. Then, denoising is performed on the initial vector composed entirely of noise data given based on the noise at each time step, so that the detailed information in the guidance vector is gradually added to the initial vector through this denoising process, and finally the output corresponding to the guidance vector is generated.

[0242] Refer to Figure 13 , in one embodiment, step 630 includes:

[0243] Step 1310, initialize the pose vector of the prompt voice to be diffused as the initial pose vector of the prompt voice, and initialize the step number to 1;

[0244] Step 1320, input the pose vector of the prompt voice to be diffused, the step number, the pose description guidance vector, and the voice guidance vector into the diffusion model to obtain the diffusion noise corresponding to the step number;

[0245] Step 1330, subtract the diffusion noise corresponding to the step number from the pose vector of the prompt voice to be diffused, increment the step number by 1, and return to the step of inputting the pose vector of the prompt voice to be diffused, the step number, the pose description guidance vector, and the voice guidance vector into the diffusion model until the step number increases to the preset maximum number of steps.

[0246] The pose vector of the prompt voice to be diffused in step 1310 is the vector that will be gradually denoised when input into the diffusion model later. The initial pose vector of the prompt voice is a vector that follows the standard Gaussian distribution and is completely composed of random noise. Specifically, the initial pose vector of the prompt voice can be expressed as: This represents the initial pose vector of the prompt voice The elements in it follow the standard Gaussian distribution with a mathematical expectation of 0 and a variance of 1. The step number is used to record how many times the denoising process of the pose vector of the prompt voice to be diffused is currently being performed. Refer to Figure 14 , in this embodiment, first, the to-be-diffused prompt voice pose vector Z0 is initialized as a vector completely composed of noise data without carrying any information, and at the same time, the step number t is initialized to 1, indicating that the first denoising process of the to-be-diffused prompt voice pose vector will be performed next;

[0247] In step 1320, referring to Figure 14 , after completing the initialization of the to-be-diffused prompt voice pose vector and the step number, the to-be-diffused prompt voice pose vector Z0, the step number t, the pose description guidance vector g, and the voice guidance vector A are input into the diffusion model. Specifically, referring to Figure 4 , after concatenating the to-be-diffused prompt voice pose vector Z0 and the step number t, the process of concatenating the step number t and the to-be-diffused prompt voice pose vector Z0 is equivalent to adding the information of the step number to the to-be-diffused prompt voice pose vector. It can be understood that since the diffusion process of the diffusion model is multi-step, and each step of diffusion is to denoise the to-be-diffused prompt voice pose vector of the previous time step. Therefore, by concatenating the step number and the to-be-diffused prompt voice pose vector of the corresponding time step, the information of the step number is added to the to-be-diffused prompt voice pose vector, and at the same time, it is sensed how many times of denoising have been performed, which time step of denoising is being performed, and which to-be-diffused prompt voice pose vector corresponding to which step number should be denoised. It can be understood that after the to-be-diffused prompt voice pose vector is input into the diffusion model, the diffusion model itself samples from a randomly distributed Gaussian noise as the diffusion noise, but this sampling process is completely random, so it is impossible to generate a target prompt voice pose vector corresponding to the target input directionally. Therefore, the pose description guidance vector and the voice guidance vector need to be input into the diffusion model together. After the pose description guidance vector and the voice guidance vector are input into the diffusion model, the sampling center when the diffusion model samples Gaussian noise will be changed, so that the sampling center when the diffusion model samples noise is close to the corresponding guidance centers of the pose description guidance vector and the voice guidance vector, so that the target prompt voice pose vector can be generated directionally. It can be understood that the diffusion noise ∈ t predicted by the diffusion model is a vector with the same dimension as the to-be-diffused prompt voice pose vector Z0. Specifically, the expression for predicting the diffusion noise corresponding to each step number in the diffusion model is as follows:

[0248] ∈ t =G D (Z t ,t,g,A);

[0249] In the formula, ∈ t is the diffusion noise predicted by the diffusion model corresponding to the step number t, and Z t Denote the speech gesture vector of the prompt to be diffused corresponding to the step number t, that is, the speech gesture vector of the prompt to be diffused that has undergone t-1 denoising processes. t represents the step number, g represents the gesture description guidance vector, and A represents the speech guidance vector.

[0250] The maximum number of steps in step 1330 is a pre-set value, such as 1000. It can be understood that by using the diffusion noise ∈ corresponding to the step number t t to subtract the speech gesture vector Z0 of the prompt to be diffused, it is actually a process of adding detailed information to the speech gesture vector of the prompt to be diffused and updating the speech gesture vector of the prompt to be diffused. This update process can be expressed as: Z n+1 ~p(Z n+1 |Z n , U), where U is the output of the encoder, that is, the gesture description guidance vector, Z n+1 represents the result of the (n + 1)-th step of denoising, and Z n represents the result of the n-th step of denoising. This equation shows that in the (n + 1)-th step of denoising, the diffusion model predicts the noise according to the gesture description guidance vector U and denoises the denoising result Z n of the n-th step to obtain the denoising result Z n+1 of the (n + 1)-th step. The specific denoising process can be expressed as:

[0251]

[0252] where is the speech gesture vector of the prompt to be diffused corresponding to the step number t, and ∈ t is the diffusion noise corresponding to the step number t, is the speech gesture vector of the prompt to be diffused obtained by subtracting the speech gesture vector Z t of the prompt to be diffused corresponding to the step number t using the diffusion noise ∈ corresponding to the step number t, t and is also the speech gesture vector of the prompt to be diffused corresponding to the step number t + 1. Taking the set maximum number of steps as T as an example, when the step number t < T, it means that the denoising process of the speech gesture vector of the prompt to be diffused has not been completed. At this time, the step number t is incremented by 1, and then step 1320 is executed again. The vector after subtracting the diffusion noise ∈ t is used as the speech gesture vector Z0 of the prompt to be diffused and concatenated with the step number t + 1, and then input into the diffusion model together with the gesture description guidance vector and the speech guidance vector to predict the diffusion noise ∈ corresponding to the step number t + 1, and ∈ t+1 is subtracted from Z0 t+1 , Repeat the above steps until the step number t reaches the set maximum number of steps T. At this time, it can be considered that the denoising process of the to-be-diffused prompt speech pose vector is completed, and the target prompt speech pose vector corresponding to the pose description guidance vector is obtained.

[0253] In this embodiment, through the initialization of the to-be-diffused prompt speech pose vector and the step number, and then inputting it together with the pose description guidance vector and the speech guidance vector into the diffusion model, the diffusion noise corresponding to each step number is predicted in turn by the diffusion model, and the diffusion noise of each step number is used to subtract the to-be-diffused prompt speech pose vector. Through this subtraction process, relevant details of the target prompt speech pose vector are gradually added to the to-be-diffused prompt speech pose vector until the step number reaches the maximum number of steps. At this time, it can be considered that enough details have been added to the to-be-diffused prompt speech pose vector. At this time, it can be considered that the to-be-diffused prompt speech pose vector is the target prompt speech pose vector.

[0254] In one embodiment, referring to Figure 4 , before the diffusion model, there are also connected a multi-head attention model, a first feed-forward neural network, an adaptive instance normalization layer, a second feed-forward neural network, and multiple stacked normalization layers. Then, referring to Figure 19 , step 1320 may include:

[0255] Step 1910, input the concatenated vector of the to-be-diffused prompt speech pose vector and the step number into the multi-head attention model to obtain the multi-head attention vector;

[0256] Step 1920, stack and normalize the concatenated vector and the multi-head attention vector to obtain the first normalized vector;

[0257] Step 1930, input the first normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain the diffusion noise corresponding to the step number.

[0258] In step 1910, concatenating the to-be-diffused prompt speech pose vector and the step number results in a concatenated vector. Inputting this concatenated vector into the multi-head attention model, the multi-head attention model can capture the important features in the input concatenated vector. Specifically, it splits the input concatenated vector into multiple segments, calculates the attention weights of each segment in each attention head, and finally performs weighted summation on these segments based on the attention weights, thereby reconstructing the input concatenated vector, highlighting the important features and weakening the unimportant features. It can be understood that the concatenated vector is obtained by concatenating the to-be-diffused prompt speech vector and the step number, and the features in it are mainly the features in the to-be-diffused prompt speech pose vector. That is, the multi-head attention vector can actually be regarded as the to-be-diffused prompt speech pose vector that highlights the important features and weakens the unimportant features.

[0259] In step 1920, after the multi-head attention vector is obtained by inputting the cascade vector into the multi-head attention model, the multi-head attention vector highlights the important feature parts in the cascade vector and weakens some unimportant features, which may cause the multi-head attention vector to lose some details compared to the cascade vector. Therefore, referring to Figure 4 In the diffusion model, the cascade vector and the multi-head attention vector are superimposed through a skip connection structure, and the superimposed vector is normalized to obtain the first normalized vector. Compared with the cascade vector, the first normalized vector better highlights the important features, and compared with the multi-head attention vector, the first normalized vector better retains the details in the cascade vector. In this way, the target prompt speech posture vector finally generated can not only better remove the influence of redundant information, but also retain some detailed information, thereby improving the accuracy of the target prompt speech posture vector finally generated.

[0260] In step 1930, the first normalized vector, the posture description guide vector and the voice guide vector are input into the diffusion model. The posture description guide vector and the voice guide vector are used to change the mean center of the diffusion model when sampling Gaussian noise, so that the diffusion model directionally samples the noise that matches the posture description guide vector and the voice guide vector as diffusion noise. In this way, the diffusion process of the diffusion model is guided by the posture description guide vector and the voice guide vector, and the target prompt voice posture vector is directionally generated.

[0261] In this embodiment, a multi-head attention model is set in front of the diffusion model to extract important features in the gesture vector of the prompt speech to be diffused, and the multi-head attention vector and the cascade vector are superimposed and normalized through a skip connection structure. In this way, the important features in the gesture vector are highlighted and the interference of redundant information is reduced while the detailed information therein can be effectively retained, thereby improving the accuracy of the generated target prompt speech gesture vector.

[0262] Reference Figure 20 In one embodiment, step 1930 includes:

[0263] Step 2010: input the first normalized vector into a first feedforward neural network to obtain a first feedforward vector;

[0264] Step 2020: superimpose the first feedforward vector and the first normalized vector and normalize them to obtain a second normalized vector;

[0265] Step 2030: Input the second normalized vector, the posture description guide vector, and the voice guide vector into a diffusion model to obtain a diffused vector.

[0266] Step 2040: Input the diffused vector into the second feed-forward neural network to obtain the second feed-forward vector;

[0267] Step 2050: Superimpose and normalize the second feed-forward vector and the diffused vector to obtain diffused noise.

[0268] The first feed-forward neural network in Step 2010 can further extract deeper features in the first normalized vector, further highlighting the important features in the first normalized vector. After the first normalized vector passes through the first feed-forward neural network, the important features therein can be further highlighted.

[0269] In Step 2020, with reference to Figure 4 , the first feed-forward vector output by the first feed-forward neural network and the first normalized vector input into the first feed-forward neural network are superimposed through the first skip connection structure and normalized to obtain the second normalized vector. The important features of the to-be-diffused prompt speech pose vector are further highlighted in the second normalized vector, and at the same time, the details therein are well retained. At the same time, this skip connection structure also constitutes a residual structure, which can effectively reduce the problem of network degradation and improve the convergence speed of the model during training

[0270] In Step 2030, the second normalized vector, the pose description guidance vector, and the speech guidance vector are input into the diffusion model, and the diffusion model will predict the diffused state of the to-be-diffused prompt speech pose vector, thus obtaining the diffused vector.

[0271] [[ID=1 / 17]]The second feed-forward neural network in Step 2040 is similar in function to the first feed-forward neural network and will not be elaborated here. The diffused vector is input into the second feed-forward neural network to extract the important features in the second feed-forward neural network and obtain the second feed-forward vector.

[0272] In Step 2050, the second feed-forward vector and the diffused vector are superimposed. This skip connection structure constitutes the second residual structure, and through this residual structure, the convergence speed of the diffusion model during training can be further improved. In the case of a relatively deep network layer, the model can also converge quickly.

[0273] In this embodiment, by respectively setting the first feed-forward neural network and the second feed-forward neural network before and after the diffusion model, and forming a residual structure by skip connecting the input and output of the first feed-forward neural network, and forming a residual structure by skip connecting the input and output of the second feed-forward neural network, through the connection of this residual structure, the problem of network degradation is reduced, and the convergence speed during the training of the diffusion model is improved.

[0274] Detailed description of the training process of the diffusion model

[0275] A diffusion model is a generative model that generates a specific target vector by gradually predicting noise and then denoising the input initial vector based on the predicted noise. It can be understood that the initial vector input to the diffusion model is a vector completely composed of random Gaussian noise, which can be regarded as obtained by adding random Gaussian noise to any given target vector multiple times. Based on this, by performing specific denoising processing on this initial vector, the initial vector can be restored to any target vector. In this embodiment, the diffusion model is used to generate a target prompt speech pose vector according to the input pose description guidance vector. Then, it is necessary to train the diffusion model to learn the process of generating the corresponding prompt speech pose vector according to the pose description guidance vector.

[0276] Referring to Figure 15 , in one embodiment, the diffusion model is trained in the following manner:

[0277] Step 1510, obtain a second sample set, where each second sample in the second sample set includes a sample reference prompt speech pose vector, a sample pose description vector, and a sample speech guidance vector;

[0278] Step 1520, perform multi-step noise addition processing on the sample reference prompt speech pose vector, and record the noise added at each step number as the noise label corresponding to the step number;

[0279] Step 1530, generate a prompt speech pose label corresponding to the step number based on the noise-added prompt speech pose vector obtained after each step of noise addition processing;

[0280] Step 1540, obtain a rhythm adjustment pose difference label corresponding to the step number;

[0281] Step 1550, input the sample to-be-diffused prompt speech pose vector obtained after multi-step noise addition processing into the diffusion model, and perform multi-step denoising processing under the guidance of the sample pose description vector and the sample speech guidance vector. Calculate a third loss function based on the denoising result at each step, the noise label corresponding to the step number, the prompt speech pose label, and the rhythm adjustment pose difference label, and train the diffusion model based on the third loss function.

[0282] The second sample set in step 1510 may include multiple second samples, and each second sample includes a sample reference prompt voice pose vector, a sample pose description vector, and a sample voice guidance vector. Specifically, each second sample can be represented as a tuple (Z0, n, g, A, T), where Z0 represents the sample reference prompt voice pose vector, g represents the sample pose description vector, A represents the sample voice guidance vector, n represents the maximum number of diffusion steps, which is also the maximum number of steps for adding noise to the sample reference prompt voice pose vector, and T is the text used to generate the sample pose description vector. Each second sample may not carry the text T. It can be understood that the sample reference prompt voice pose vector Z0, the sample pose description vector g, and the sample voice guidance vector A belonging to the same second sample correspond to each other. For example, if the sample reference prompt voice pose vector corresponds to the pose of "extending the index finger and placing the hand near the eyes", then the sample pose description vector is encoded by the first vector generation model for the pose description of "extending the index finger and placing the hand near the eyes", and the sample voice guidance vector is also encoded for the voice message of the sentence "extending the index finger and placing the hand near the eyes".

[0283] In step 1520, the sample reference prompt voice pose vector can be processed by adding noise by generating random Gaussian noise and adding the generated Gaussian noise to the sample reference prompt voice pose vector. Specifically, referring to Figure 16 , at step number 1, a random Gaussian noise ∈1 is generated. This Gaussian noise ∈1 is a vector with the same dimension as the sample reference prompt voice pose vector Z0. Adding this Gaussian noise ∈1 to the sample reference prompt voice pose vector Z0 completes one noise addition process for the sample reference prompt voice pose vector, and ∈1 is the noise label corresponding to step number 1. After that, the sample reference prompt voice pose vector with added noise is processed for the next noise addition. The above process is repeated until the step number reaches the maximum number of steps n, thus completing the multi-step noise addition process for the sample reference prompt voice pose vector and obtaining the noise label ∈ corresponding to each step number t. t , referring to Figure 16 , when the step number reaches the maximum number of steps n, that is, when the step number t = n, it can be considered that the noise addition process for the sample reference prompt voice pose vector has been completed.

[0284] The prompt voice pose label in step 1530 refers to the sample reference prompt voice pose vector obtained after adding noise at each step number. Referring to Figure 16 , at step number 1, the sample reference prompt speech pose vector Z0 is added to the Gaussian noise ∈1 corresponding to step number 1, obtaining the sample reference prompt speech pose vector Z1 after one round of noise addition. Z1 is the prompt speech pose label corresponding to step number 1. Similarly, at step number t, the sample reference prompt speech pose vector Z after t - 1 rounds of noise addition processing t-1 Add noise ∈ t After that, the sample reference prompt speech pose vector Z is obtained t which is the prompt speech pose label corresponding to step number t.

[0285] The rhythm adjustment pose difference label in step 1540 refers to the difference between the prompt speech pose label of each step number and the mean of all prompt speech pose labels. It can be understood that the prompt speech pose labels are all sample reference prompt speech pose vectors after noise addition processing for a certain number of steps, and they are all vectors with the same dimension. Therefore, the prompt speech pose labels corresponding to all step numbers can be summed and then divided by the maximum number of steps to obtain the mean M of all prompt speech pose labels. After that, the difference between the prompt speech pose label Z corresponding to each step number t t and this mean M is calculated to obtain the rhythm adjustment pose difference label M corresponding to step number t t .

[0286] The sample to-be-diffused prompt speech pose vector in step 1550 is the sample reference prompt speech pose vector after multi-step noise addition processing. Specifically, referring to Figure 16 , when the step number t reaches the maximum number of steps n, it is considered that the multi-step noise addition processing of the sample reference prompt speech pose vector is completed at this time, and the prompt speech pose label Z corresponding to step number t = n n is the sample to-be-diffused prompt speech pose vector. It can be understood that n is a sufficiently large value, such as 1000. When step number t = n, that is, after the sample reference prompt speech pose vector has been subjected to n rounds of noise addition processing, it can be considered that the sample reference prompt speech pose vector at this time is already a vector completely composed of random noise, that is, the sample to-be-diffused prompt speech pose vector. Inputting this sample to-be-diffused prompt speech pose vector into the diffusion model, the diffusion model predicts the noise of each step number under the guidance of the sample pose description vector and the sample speech guidance vector, and performs multi-step denoising processing on this sample to-be-diffused prompt speech pose vector, enabling the diffusion model to learn how to predict the corresponding noise according to the given pose description guidance vector and speech guidance vector. Specifically, referring to Figure 16 , the step number can be directly passed from the noise addition process to the denoising process, and the step number is continuously decreased by 1 during the denoising process. When the step number is decreased to 1, it is considered that the denoising process is completed. By reusing the step number, on the one hand, it can ensure that the number of steps in noise addition and denoising is equal, making the multi-step denoising process formally equivalent to the reverse process of multi-step noise addition. On the other hand, it can better correspond each step of the multi-step noise addition process and the multi-step denoising process, and input the sample to-be-diffused prompt speech pose vector Z n and the step number into the diffusion model. Under the guidance of the sample pose description vector and the sample speech guidance vector, the diffusion model will predict a noise ∈′ t . It can be understood that the noise predicted by the diffusion model is a vector with the same dimension as the input sample to-be-diffused prompt speech pose vector. Then, through the predicted noise ∈′ t denoise the sample to-be-diffused prompt speech pose vector Z n , that is, subtract the noise ∈′ n from Z t to obtain Z′ t-1 , and at the same time decrease the step number t by 1. Then, use Z′ t-1 as the input of the diffusion model to predict the noise ∈′ t-1 at step number t - 1, and denoise Z′ t-1 through this noise ∈′ t-1 . Repeat the above process until the step number t = 1. At this time, it is considered that the diffusion model has completed the multi-step denoising process of the sample to-be-diffused prompt speech pose vector Z n . Among them, ∈′ t is the predicted noise corresponding to the step number t, and Z′ t is the predicted prompt speech pose corresponding to the step number t. Referring to step 1540, according to the predicted prompt speech poses Z′ t of all step numbers t, the corresponding predicted rhythm adjustment pose difference M′ ′ can also be calculated. At this time, according to the predicted diffusion noise ∈′ t corresponding to each step number t, the noise label ∈ t corresponding to each step number t, the predicted prompt speech pose Z′ t corresponding to each step number t, the prompt speech pose label Z t corresponding to each step number t, the predicted rhythm adjustment pose difference M′ t corresponding to each step number t, and the rhythm adjustment pose difference label M t The third loss function can be calculated, which can characterize the difference between the denoising result of the diffusion model and the noise label, the prompt speech pose label, and the rhythm adjustment pose difference label during the noise addition process. The larger this difference is, the larger the difference between the noise predicted by the diffusion model and the true noise label, that is, the lower the quality of the noise predicted by the diffusion model. By minimizing the third loss function, the parameters of the diffusion model are continuously optimized. When the third loss function is less than the set threshold, it indicates that the difference between the noise predicted by the diffusion model and the true noise label is small, and at this time, the quality of the noise predicted by the model is higher.

[0287] In this embodiment, the sample reference prompt speech pose vector is subjected to noise addition processing, and the noise label, the prompt speech pose label, and the rhythm adjustment pose difference label during each step of noise addition are recorded. Then, the sample to-be-diffused prompt speech pose vector obtained after the noise addition processing is subjected to multi-step denoising processing, and the denoising result corresponding to each step number is recorded. The third loss function is calculated according to the denoising result, the noise label, the prompt speech pose label, and the rhythm adjustment pose difference label corresponding to each step number, and the diffusion model is trained based on the third loss function. In this way, the diffusion model can learn how to predict the predicted diffusion noise close to the true noise label during the noise addition process according to the given pose description guidance vector and speech guidance vector.

[0288] Refer to Figure 17 , in one embodiment, step 1550 includes:

[0289] Step 1710, calculate the first loss sub-function based on the predicted diffusion noise corresponding to the step number and the noise label;

[0290] Step 1720, calculate the second loss sub-function based on the predicted prompt speech pose corresponding to the step number and the prompt speech pose label;

[0291] Step 1730, calculate the third loss sub-function based on the predicted rhythm adjustment pose difference corresponding to the step number and the rhythm adjustment pose difference label;

[0292] Step 1740, calculate the third loss function based on the first loss sub-function, the second loss sub-function, and the third loss sub-function.

[0293] The first loss sub - function in step 1710 is used to measure the difference between the predicted diffusion noise corresponding to each step number and the noise label. Both the predicted diffusion noise and the noise label are vectors with the same dimension as the given sample reference prompt speech pose vector. That is, the predicted diffusion noise and the noise label corresponding to step number t are actually two vectors with the same dimension. At this time, the Euclidean norm between these two vectors can be calculated as the first loss function. Specifically, first calculate the first difference between the predicted diffusion noise and the noise label corresponding to the step number, that is, the difference between the two. It can be understood that the predicted diffusion noise and the noise label are vectors with the same dimension. Calculating the difference between the two means subtracting the elements of the two one by one. For example, if the predicted diffusion noise is [1, 5, 3, 4] and the noise label is [2, 1, 3, 6], then the first difference is [1 - 2, 5 - 1, 3 - 3, 4 - 6], that is, [-1, 4, 0, -2]. It can be understood that the first difference between the predicted diffusion noise and the noise label is still a vector. Through this vector, it is not intuitive to determine the difference size between the two vectors. It can be understood that for two relatively close vectors, each element in the two should be relatively close. Then each element in the difference vector obtained by subtracting the two vectors should be close to 0, that is, the norm length of this difference vector should be small. Similarly, for two vectors with a large gap, at least one element will have a large gap, and then the norm length of their difference vector will also be large. Based on this, the square of the norm length of the first difference between the predicted diffusion noise and the noise label can be calculated to determine the difference size between the predicted diffusion noise and the noise label. For example, referring to the above example, the first difference between the predicted diffusion noise and the noise label is [-1, 4, 0, -2], then the corresponding square of the norm length is (-1) 2 +(4) 2 +(0) 2 +(-2) 2 =21. At this time, through the square of the norm length of the first difference, it can be intuitively judged that the difference between the predicted diffusion noise and the noise label at this time is large. Specifically, the first loss sub - function can be expressed as:

[0294]

[0295] where, L noise is the first loss sub - function, ∈ represents the noise label, ∈ θ (Z n ,n,g,A,T) represents that the diffusion model θ according to the input sample to - be - diffused prompt speech pose vector Z n , the predicted diffusion noise generated by the maximum number of steps n, the sample pose description vector g, and the sample voice guidance vector A, where T is used to generate the sample pose description vector g, and ‖‖2 represents calculating the norm. It can be understood that the noise label here refers to the noise label corresponding to any step number. Correspondingly, the predicted diffusion noise ∈ θ (Z n , n, g, A, T) also refers to the predicted diffusion noise corresponding to any step number. During the training process, the first loss sub-function is calculated once at each time step of each step number.

[0296] The second loss sub-function in step 1720 is used to measure the difference between the predicted prompt voice pose and the prompt voice pose label. It can be understood that both the predicted prompt voice pose and the prompt voice pose label are pose vectors. For pose vectors, directionality is a relatively important feature. For example, if the vector representation of the predicted prompt voice pose is [1, 2, 3, 4] and the vector representation of the prompt voice pose label is [2, 4, 6, 8], the elements in these two vectors seem to have a large difference, but the directions of these two vectors are exactly the same. Then, the action poses corresponding to the predicted prompt voice pose and the prompt voice pose label actually partially overlap, that is, these two poses are actually very close. Based on this, in this embodiment, the second loss sub-function can be calculated by the cosine similarity. Specifically, step 1720 includes: calculating the cosine distance between the predicted prompt voice pose corresponding to the step number and the prompt voice pose label; taking the difference between 1 and the cosine distance as the second loss sub-function. Specifically, the second loss sub-function can be expressed as:

[0297]

[0298] where, L semantic is the second loss sub-function, is the prompt voice pose label corresponding to any step number, is the predicted prompt voice pose corresponding to any step number. During the process of training the diffusion model, the second loss sub-function is calculated once for each step number. It can be easily seen from the above formula that when the directions of the prompt voice pose label and the predicted prompt voice pose corresponding to the step number are closer, the cosine distance between the two is closer to 1, and the value of the corresponding second loss sub-function is smaller. On the contrary, when the directions of the prompt voice pose label and the predicted prompt voice pose corresponding to the step number have a larger difference, the cosine distance between the two is closer to 0, and the value of the corresponding second loss sub-function is larger.

[0299] The third loss sub - function in step 1730 is used to measure the difference between the rhythm - adjusted pose difference label corresponding to the step number and the predicted rhythm - adjusted pose difference. It can be understood that the rhythm - adjusted pose difference label corresponding to the step number is actually the difference between the prompt voice pose label corresponding to the step number and the mean of the prompt voice pose labels of all step numbers. That is, the rhythm - adjusted pose label is actually a vector with the same dimension as the prompt voice pose label. Similarly, the predicted rhythm - adjusted pose difference is also a vector with the same dimension as the predicted prompt voice pose. It is not difficult to infer that the rhythm - adjusted pose difference label corresponding to the step number and the predicted rhythm - adjusted pose difference are two vectors with the same dimension. Then, the third loss sub - function can be determined by calculating the difference vector between the rhythm - adjusted pose difference label and the predicted rhythm - adjusted pose difference, and then calculating the norm of this difference vector. The specific process refers to step 1710 and will not be elaborated here. It can be understood that the third loss sub - function can also be calculated by calculating the cosine distance between the rhythm - adjusted pose difference label and the predicted rhythm - adjusted pose difference, and using 1 minus the cosine distance between the two as the third loss sub - function. The specific process refers to step 1720 and will not be elaborated here.

[0300] In step 1740, based on the first loss sub - function, the second loss sub - function, and the third loss sub - function, the corresponding third loss function can be determined. The third loss function is the sum of the diffusion models. Specifically, the third loss function can be obtained by directly adding the first loss sub - function, the second loss sub - function, and the third loss sub - function.

[0301] It can be understood that for the diffusion model, what it most needs to learn is how to predict the diffusion noise according to the pose description guidance vector and the voice guidance vector. Similarly, the importance of the diffusion model learning how to predict the prompt voice pose and how to predict the rhythm - adjusted pose difference is also different. Based on this, corresponding weights can be set for the first loss sub - function, the second loss sub - function, and the third loss sub - function. In one embodiment, step 1740 may include: obtaining the first weight of the first loss sub - function, the second weight of the second loss sub - function, and the third weight of the third loss sub - function; using the first weight, the second weight, and the third weight to perform a weighted sum of the first loss sub - function, the second loss sub - function, and the third loss sub - function to obtain the third loss function. Specifically, the third loss function can be expressed as:

[0302] L3 = αL noise +βL semantic +γL rhythm ;

[0303] where L3 represents the third loss function, α, β, and γ are the first weight, the second weight, and the third weight respectively, and L noise 、L semantic , L rhythm are the first loss sub-function, the second loss sub-function, and the third loss sub-function respectively.

[0304] In this embodiment, by obtaining the first weight, the second weight, and the third weight, the first loss sub-function, the second loss sub-function, and the third loss sub-function are weighted and summed to obtain the third loss function. Based on the importance of performing different types of tasks by the diffusion model, weights are assigned to each loss sub-function. Through the setting of these weights, the diffusion model can not only focus on optimizing the performance of predicting noise during training, but also take into account optimizing the performance of predicting the prompt speech gesture and adjusting the pose difference label of the rhythm.

[0305] In one embodiment, referring to Figure 18 , the predicted diffusion noise can be predicted by the diffusion model in the following manner.

[0306] Step 1810: Input the first-proportion sample to-be-diffused prompt speech gesture vector, step number, sample gesture description vector, and sample speech guidance vector into the diffusion model to obtain the first sub-predicted diffusion noise corresponding to the step number;

[0307] Step 1820: Input the second-proportion sample to-be-diffused prompt speech gesture vector, step number, and sample speech guidance vector into the diffusion model to obtain the second sub-predicted diffusion noise corresponding to the step number, where the sum of the first proportion and the second proportion is 1;

[0308] Step 1830: Weight and sum the first sub-predicted diffusion noise and the second sub-predicted diffusion noise according to the first proportion and the second proportion to obtain the predicted diffusion noise.

[0309] The first sub-predicted diffusion noise in Step 1810 refers to the diffusion noise predicted by the diffusion model under the guidance of the sample gesture description vector. The first proportion can be a pre-set value within the range of (0, 1), such as 0.9. Specifically, in the case of being guided by the sample gesture description vector, the diffusion model will change the sampling center when sampling Gaussian noise under the joint guidance of the sample gesture description vector and the sample speech guidance vector, and then sample the noise with the changed sampling center. At this time, the sampled noise can be regarded as the diffusion noise predicted by the diffusion model, that is, the first sub-predicted diffusion noise.

[0310] The second sub-predicted diffusion noise in Step 1820 is the diffusion noise predicted by the diffusion model without the guidance of the sample gesture description vector. The second proportion is the difference between 1 and the first proportion. Specifically, in the case of lacking the guidance of the sample gesture description vector, the diffusion model will sample a noise as the predicted diffusion noise under the guidance of the sample speech guidance vector. At this time, the sampled noise is the second sub-predicted diffusion noise.

[0311] In step 1830, after obtaining the first sub-predicted diffusion noise and the second sub-predicted diffusion noise, they are weighted and summed according to the first ratio and the second ratio to obtain the predicted diffusion noise. By predicting the diffusion noise in a way that combines guided and unguided methods, the real situation can be better simulated, and the diffusion effect of the diffusion model can be improved. Specifically, the predicted diffusion noise can be obtained by the following formula:

[0312]

[0313] where is the predicted diffusion noise corresponding to the step number n, s represents the first ratio, and correspondingly, (1 - s) represents the second ratio, ∈ θ represents the noise predicted by the diffusion model θ, Z n represents the speech pose vector to be diffused of the sample, n represents the step number, g is the pose description guidance vector, A is the speech guidance vector, and T is the text used to generate the pose description guidance vector g, is an empty set.

[0314] In this embodiment, by using a method that combines the guidance of the sample pose description vector and unguided method during the training of the diffusion model to generate the predicted diffusion noise, the real diffusion task can be better simulated, and the performance of the diffusion model can be improved.

[0315] Detailed description of the training process of adding the rhythm adjustment gesture difference label and the rhythm adjustment gesture difference generation model Description

[0316] Existing prompt voice gesture generation is centered around text data for generating prompt voice gestures, and generally does not use voice data. Even when voice data is used, it is usually recognized as text data and then used to generate prompt voice gestures through the text data. However, text data itself cannot carry information such as speech rate, which may lead to the generated prompt voice gestures being out of sync with the speech of the speaker. For example, in a video scenario, it is possible that the video producer has already started explaining the content of the next part, and the video screen has switched to the relevant screen of the next part, but the generated prompt voice gestures are still presenting the content of the previous part. At this time, the voice of the video producer and the video screen are out of sync with the prompt voice gestures, and the two are respectively describing the content of two different parts of the video. Due to this out-of-sync situation, hearing-impaired people cannot well establish the connection between the video screen and the prompt voice gestures, resulting in the hearing-impaired people being unable to well obtain the content of the video. In another situation, when hearing-impaired people need to obtain the content of the words spoken by multiple people in sequence through prompt voice gestures, for example, when hearing-impaired people watch a video involving conversations between different characters, since the final presentation form of the prompt voice gestures is an image and it cannot carry information such as timbre for distinguishing speakers, at this time, if the prompt voice gestures are out of sync with the voice, it will make it impossible for hearing-impaired people to well distinguish which character's speech content the current prompt voice gesture corresponds to according to the lip movements of each character, which will also cause hearing-impaired people to be unable to well obtain effective information through prompt voice gestures. It is desired that the generated prompt voice gestures can be synchronized with the corresponding voice information. Based on this, referring to Figure 21 After step 350, it further includes:

[0317] Step 2110, input the target voice into the rhythm adjustment gesture difference generation model to obtain the rhythm adjustment gesture difference;

[0318] Step 2120, adjust the target prompt voice gesture with the rhythm adjustment gesture difference to obtain the adjusted prompt voice gesture.

[0319] The rhythm adjustment gesture difference in step 2110 refers to the difference between the prompt voice gesture generated without considering the rhythm of the target voice at all and the prompt voice gesture generated after adding the rhythm information of the target voice. The rhythm adjustment gesture difference generation model is a trained model that can extract the rhythm features of voice information. The training process of the rhythm adjustment gesture difference generation model will be described later and will not be elaborated here. It can be understood that the rhythm adjustment gesture difference is a vector with the same dimension as the vector representation of the generated prompt voice gesture. Referring to Figure 4 , the Mel-scale Frequency Cepstral Coefficients (MFCC) features can be extracted from the target speech first, and then the MFCC features are input into the rhythm adjustment pose difference generation model, and the corresponding rhythm adjustment pose difference is generated from the generation model through the rhythm adjustment pose difference.

[0320] In step 2120, as described in step 2110, the rhythm adjustment pose difference is the difference between the cue speech poses before and after adding the rhythm information of the target speech. The target cue speech pose obtained in step 350 is the cue speech pose generated without considering the rhythm information of the target speech at all. The result of adding the rhythm adjustment pose difference generated according to the target speech to the target cue speech pose is the cue speech pose generated after adding the rhythm information of the target speech. Based on this, by superimposing the target cue speech pose and the rhythm adjustment pose difference, the adjusted cue speech pose is obtained. It can be understood that the adjusted cue speech pose adds the rhythm adjustment pose difference generated according to the rhythm information of the target speech, and compared with the target cue speech pose, it can be better synchronized with the target speech.

[0321] In this embodiment, by inputting the target speech into the rhythm adjustment pose difference generation model to generate the rhythm adjustment pose difference, and then adjusting the generated target cue speech pose through the rhythm adjustment pose difference, injecting the rhythm information of the target speech into the target cue speech pose to obtain the adjusted cue speech pose, which can better ensure that the generated cue speech pose and the target speech are synchronized, so that the hearing-impaired people can better obtain information through the cue speech pose.

[0322] Refer to Figure 22 , in one embodiment, the rhythm adjustment pose difference generation model is trained in the following manner:

[0323] Step 2210, obtain the third sample set, including multiple sample video segments obtained by splitting the target sample video and the sample speech segments corresponding to the sample video segments;

[0324] Step 2220, extract the sample object motion representation from the sample video segment;

[0325] Step 2230, determine the average motion representation of the sample object motion representation;

[0326] Step 2240, determine the difference between the sample object motion representation and the average motion representation as the rhythm adjustment pose difference label;

[0327] Step 2250: Input the sample voice segment corresponding to the sample video segment into the rhythm adjustment pose difference generation model to obtain the predicted rhythm adjustment pose difference;

[0328] Step 2260: Calculate the fourth loss function based on the predicted rhythm adjustment pose difference and the rhythm adjustment pose difference label, and train the rhythm adjustment pose difference generation model based on the fourth loss function.

[0329] The third sample set in Step 2210 is the sample data set for training the rhythm adjustment pose difference generation model. The target sample video can be a video in which the recorder uses a prompt voice pose to express specific content. Refer to Figure 23 , the sample video segment can be obtained by splitting the target sample video. Each sample video segment is a part of the target sample video. The method of splitting the target sample video can be random splitting or splitting based on the movement amplitude of the recorder. For example, at the beginning, the recorder does not move at a certain position A in the target sample video, and then walks from position A to another position B. After walking to position B, the recorder stays at position B and no longer moves. At this time, the target sample video can be split into at least three sample video segments, corresponding to the segment where the recorder is at position A, the segment where the recorder walks from position A to position B, and the segment where the recorder stays at position B. By splitting the target sample video, the influence of the recorder's walking or other behaviors on the coordinates of the key hand bones can be better removed later. The sample voice segment is the voice information corresponding to each sample video segment. For example, in a sample video segment, the content expressed by the recorder using the prompt voice pose is "Hello everyone", then the sample voice segment corresponding to this sample video segment is the voice of the sentence "Hello everyone".

[0330] The sample object motion representation in Step 2220 can be a vector composed of the coordinates of the key human body bones of the recorder in each frame. Specifically, the sample video segment is processed frame by frame to extract the video image of each frame. Then, through human key bone extraction models such as OpenPose and PoseNet to process these video images, the coordinates of the key hand bones of the recorder can be extracted from the sample video segment as the corresponding sample object motion representation. Refer to Figure 23 , for each sample video, multiple frames of sample object motion representations from the first frame to the t-th frame in the segment can be extracted, that is, M1 to M t .

[0331] The average motion representation in Step 2230 refers to the mean value of the sample object motion representations of all frames in the same sample video segment. It can be understood that for each sample video segment, there is a corresponding average motion representation. Refer to Figure 23 , the sample object motion representations of all frames in the same sample video segment, i.e., M1 to M t can be added together, and then the sum of the addition is divided by the number of frames t in the sample video segment, so as to obtain the average motion representation of the sample video segment

[0332] In step 2240, by calculating the sample object motion representation of each frame in the sample video segment and the average motion representation of the sample video segment, the rhythm adjustment pose difference label of each frame is obtained. Refer to Figure 23 , for the first frame in the sample video segment, the corresponding rhythm adjustment pose difference label is The rhythm adjustment pose difference label of the t-th frame in the sample video segment is It can be understood that in the sample video segment, in addition to speaking, the person being recorded may also perform other behaviors. For example, while using the prompt voice gesture to express content, the person being recorded is walking. At this time, the movement of the person being recorded will cause a large movement of the coordinates of each key bone point of the person's hand in the video screen. That is to say, the sample object motion representation M t of the t-th frame includes not only the change in the coordinates of the key hand bone points brought about by the person being recorded using the prompt voice gesture, but also the change in the key hand bone points caused by the movement of the person being recorded. And the average motion representation can reflect the change in the coordinates of the key hand bone points caused by the movement of the person being recorded. At this time, subtracting the average motion representation t from the sample object motion representation M gives the real change in the coordinates of the key hand bone points brought about by the person being recorded using the prompt voice gesture, that is, the rhythm adjustment pose difference label.

[0333] In step 2250, the sample voice segment corresponding to the sample video segment is input into the rhythm adjustment pose difference generation model, and the model will predict the rhythm adjustment pose difference of each frame in the corresponding sample video segment according to the input sample voice segment. Specifically, refer to Figure 23 , after inputting the sample voice segment into the rhythm adjustment pose difference label, the model will output the predicted rhythm adjustment pose difference of the first frame to the rhythm adjustment pose difference of the t-th frame

[0334] The fourth loss function in step 2260 is used to measure the difference between the rhythm adjustment pose difference predicted by the rhythm adjustment pose difference generation model according to the target voice segment and the real rhythm adjustment pose difference label in the sample video segment. It is expected that the rhythm adjustment pose difference predicted by the rhythm adjustment pose difference generation model and the real rhythm adjustment pose difference label As close as possible. Therefore, referring to Figure 23 , the difference between the rhythm adjustment pose difference predicted by the computational model and the true rhythm adjustment pose difference label can be used as the fourth loss function. That is, the fourth loss function can be expressed as:

[0335]

[0336] where L4 is the fourth loss function, is the rhythm adjustment pose difference of any frame in the sample video segment predicted by the model, M is the motion representation of the sample object in any frame of the sample video segment, is the average motion representation of the sample video segment. It can be understood that and M correspond to the same frame in the sample video segment. For any frame in the sample video segment, the fourth loss function needs to be calculated once. When the value of the fourth loss function L4 is smaller, the difference between the rhythm adjustment pose difference predicted by the model and the true rhythm adjustment pose difference label is smaller. Based on this, the rhythm adjustment pose difference generation model can be trained by minimizing the fourth loss function.

[0337] In this embodiment, the true rhythm adjustment pose difference label is calculated through the motion representation of the sample object in the sample video segment. At the same time, the corresponding sample voice segment of the sample video segment is input into the rhythm adjustment pose difference generation model to make the model predict the corresponding rhythm adjustment pose difference. Then, the difference between the rhythm adjustment pose difference label and the rhythm adjustment pose difference predicted by the model is used as the fourth loss function, and the rhythm adjustment pose difference generation model is trained by minimizing the fourth loss function, so that the rhythm adjustment pose difference predicted by the model is as close as possible to the true rhythm adjustment pose difference label. In this way, the trained rhythm adjustment pose difference generation model can generate accurate rhythm adjustment pose differences based on the input voice.

[0338] Detailed description of jointly training the rhythm adjustment gesture difference generation model and the diffusion model

[0339] It can be understood that in this embodiment, the diffusion model and the rhythm adjustment pose difference generation model work together to jointly generate the final prompt voice pose. During the training process, the diffusion model and the rhythm adjustment pose difference generation model are jointly trained, so that after training, the diffusion model and the rhythm adjustment pose difference generation model can work better together to accurately generate prompt voice poses with high fineness.

[0340] In one embodiment, referring to Figure 24 , the process of jointly training the diffusion model and the rhythm adjustment pose difference generation model includes:

[0341] Step 2410: Obtain a fourth sample set, where the fourth sample set includes multiple multimodal samples, and each multimodal sample includes a second sample text, a sample voice, and a sample video;

[0342] Step 2420: Add the second sample text to the guiding text and input it into the large language model to obtain a sample pose description;

[0343] Step 2430: Based on the first guiding vector corresponding to the sample pose description and the second guiding vector corresponding to the sample voice, guide the diffusion model to enable the diffusion model to generate a sample prompt voice pose vector, and based on the sample prompt voice pose vector, generate a sample prompt voice pose;

[0344] Step 2440: Input the sample voice into the rhythm adjustment pose difference generation model to obtain a sample rhythm adjustment pose difference, and use the sample rhythm adjustment pose difference to adjust the sample prompt voice pose to obtain an adjusted sample prompt voice pose;

[0345] Step 2450: Calculate a fifth loss function based on the comparison between the adjusted sample prompt voice pose and the prompt voice pose label extracted from the sample video, and train the rhythm adjustment pose difference generation model and the diffusion model based on the fifth loss function.

[0346] The fourth sample set in Step 2410 is a data set including multiple multimodal samples. Each multimodal sample contains a corresponding second sample text, sample voice, and sample video. Specifically, the second sample text is the text for which a prompt voice pose needs to be generated, the sample voice is the voice reading the second sample text, and the sample video is the video in which the recorder expresses the second sample text through a prompt semantic pose. For example, if the input second sample text is the character "tree", then the sample voice is the voice reading the character "tree", and the sample video is the video expressing the character "tree" through a prompt voice pose, that is, the video of making the gesture of "extending the index finger and middle finger, spreading the two fingers apart, and placing the hand on the neck position". The fourth sample set can be expressed as: D=(T, A, V), where D represents the fourth sample set, T represents all the second sample texts in the fourth sample set, A represents all the sample voices in the fourth sample set, and V represents all the sample videos in the fourth sample set.

[0347] In Step 2420, the step of adding the second sample text to the guiding text and inputting it into the large language model to obtain a sample pose description refers to the relevant description in Step 320 of the above embodiment. Just use the second sample text as the input of the large language model, and details will not be elaborated here. Refer to Figure 25 , the sample pose description generated by the large language model based on the second sample text includes at least the relevant descriptions of hand shape and hand position, and may also include the relevant description of lip shape.

[0348] In step 2430, refer to Figure 25 The first guide vector is a guide vector corresponding to the sample posture description obtained by inputting the sample posture description into the trained first vector generation model, and the second guide vector is a guide vector corresponding to the sample speech object obtained by inputting the sample speech into the trained second vector generation model. Figure 25 , input the first guide vector and the second guide vector into the diffusion model, and change the sampling space of the diffusion model when sampling noise through the first guide vector and the second guide vector, so that the diffusion model predicts noise of multiple step numbers, and then performs multi-step denoising on the given vector to be diffused composed of random Gaussian noise, and obtains the sample prompt speech posture. It can be understood that the diffusion process of the diffusion model is a multi-step denoising process, in which each step of denoising actually generates a corresponding intermediate vector. The process of multi-step denoising of the diffusion model is regarded as a process of continuously generating prompt speech posture vectors of several frames, and each step of denoising is regarded as a time frame. Then the intermediate vector obtained by each denoising process can be regarded as the sample prompt speech posture of the corresponding time frame, which is recorded as represents the sample prompt speech pose of the i-th frame.

[0349] In step 2440, refer to Figure 25 , the sample speech is input into the rhythm adjustment posture difference generation model, so that the rhythm adjustment posture difference generation model predicts the corresponding sample rhythm adjustment posture difference according to the sample speech. It can be understood that the target speech is a continuous speech signal. When the target speech is input into the rhythm adjustment posture difference generation model, the model will actually predict the rhythm adjustment posture difference of several consecutive frames, among which the sample rhythm adjustment posture difference corresponding to the i-th frame in the sample video can be recorded as It can be understood that the sample rhythm adjustment posture difference is a vector with the same dimension as the sample prompt voice posture, and then the sample rhythm adjustment posture difference corresponding to each frame is used. By adjusting the sample prompt speech posture vector of the corresponding time frame, the adjusted sample prompt speech posture corresponding to each time frame is obtained. For details, refer to Figure 25 , the sample prompt speech posture vector and the sample rhythm adjustment posture difference corresponding to the same time frame can be superimposed to obtain the adjusted sample prompt speech posture. Specifically, the superposition process refers to the following formula:

[0350]

[0351] in, That is, the adjusted sample of the i-th frame prompts the speech posture.

[0352] The prompt voice pose label in step 2450 refers to the pose formed by the key human body skeleton points of the person being recorded in the sample video. It can be understood that the sample video is a video with multiple consecutive frames. Referring to Figure 25 , first, perform video frame extraction processing on the sample video to extract the frame images of each time frame in the sample video. Then, methods such as Expose can be used to extract the key human body skeleton points of the person being recorded from the frame images at each time. Based on the coordinates of the key human body skeleton points in each time frame, construct the corresponding prompt voice pose label for that time frame. Among them, the prompt voice pose label of the i-th time frame can be denoted as M i , after that, calculate the norm of the difference between the adjusted sample prompt voice pose and the prompt voice pose label for all time frames as the fifth loss function. Specifically, the fifth loss function can be expressed as:

[0353]

[0354] where N represents the number of time frames. It can be understood that during the training process, it is necessary to set the maximum number of denoising steps and adjust the durations of the sample video and the sample voice so that the sample prompt voice pose, the sample rhythm adjustment pose difference, and the prompt voice pose label have the same number of time frames, that is, the number of time frames of the sample prompt voice pose, the sample rhythm adjustment pose difference, and the prompt voice pose label is all N. Referring to Figure 25 , after obtaining the fifth loss function, simultaneously train the diffusion model and the rhythm adjustment pose difference generation model based on the fifth loss function. Specifically, the diffusion model and the rhythm adjustment pose difference generation model can be trained by minimizing the fifth loss function, so that the adjusted sample prompt voice pose jointly generated by the two is closer to the prompt voice pose label. In this way, the trained diffusion model and the rhythm adjustment pose difference generation model can jointly generate accurate prompt voice poses that are synchronized with the target voice.

[0355] In this embodiment, by jointly training the diffusion model and the rhythm adjustment pose difference generation model, the trained diffusion model and the rhythm adjustment pose difference generation model can work better together to generate accurate prompt voice poses that are synchronized with the input target voice.

[0356] Effect data of embodiments of the present disclosure

[0357] First, Figure 26 shows the parameters of the test set used in the test of the embodiments of the present disclosure. Referring to Figure 26 , the test set includes multiple sentences of different lengths. To enable the trained model to better meet the actual needs of hearing-impaired people, in the embodiments of the present disclosure, the lengths of the sentences in the test set are shortened and the sentence complexity is reduced. The length of the sentences used, that is, the number of words in the sentences, is mainly 7-11, and at the same time, a small number of long and complex sentences are set so that the trained diffusion model also has the ability to handle complex tasks to a certain extent.

[0358] In addition, Figure 27 shows the parameters of some other available open-source test sets. For example, the existing French prompted speech dataset was recorded by one recorder, including 238 sentences, with a total of 12,872 features and 1,190 vocabulary words for all sentences; the existing Spanish prompted speech dataset was recorded by one recorder, including 97 sentences, with a total of 2,741 features and 485 vocabulary words for all sentences; the Chinese prompted speech data MCCS-2023 was recorded by 6 recorders, including 5,000 sentences, with a total of 131,581 features and 42,248 vocabulary words for all sentences.

[0359] In the embodiments of the present disclosure, the maximum number of diffusion steps is set to 1,000. The second sample set used for training the diffusion model includes 128 groups of sample reference prompt speech pose vectors, sample pose description guidance vectors, and sample speech guidance vectors. When training the diffusion model, the first weight α is taken as 1, the second weight β is taken as 0.2, and the third weight γ is taken as 0.1. The diffusion model is trained in an end-to-end pipeline manner.

[0360] Referring to Figure 28 , Figure 28 Shows the performance parameters of the method proposed in the embodiments of the present disclosure and four existing methods, namely Speech2Gesture, GTC, HA2G, and DiffGesture, when using the MCCS-2023 dataset as the test set. Among them, the percentage of correct key points (PCK) determines the accuracy of the key skeletal point positions in the generated prompt speech gestures by judging whether the distance between the key skeletal point positions in the generated prompt speech gestures and those in the actual prompt speech gestures remains within a predetermined limit range. It is not difficult to see that the higher the percentage of correct key points, the more accurate the key skeletal point positions in the prompt speech gestures generated by the model; the pose-audio difference is used to measure the synchronization degree between the generated prompt speech gestures and the rhythm of the audio. It is not difficult to see that the higher the pose-audio difference, the more synchronous the generated prompt speech gestures and the audio are; the mean absolute joint error (MAJE) is used to measure the average change in the key skeletal point positions in the prompt speech gestures generated by the model compared to the key skeletal points of the true prompt speech gestures, thereby measuring the matching degree between the generated pose and the true pose. It is not difficult to see that the smaller the mean absolute joint error, the closer the prompt speech gestures generated by the model are to the true prompt speech gestures; the mean acceleration difference (MAD) is used to measure the L2 norm of the difference between the acceleration of the coordinate changes of each key skeletal point in the prompt speech gestures generated by the model and the actual data. It can reflect the matching degree between the predicted key skeletal point movement by the model and the true key skeletal point movement. The smaller this value, the closer the generated prompt speech gestures are to the true prompt speech gestures; the Fréchet pose distance (FGD) is used to measure the difference between the features of the generated pose and the features of the true pose. The smaller the difference in features, the closer the prompt speech gestures generated by the model are to the actual prompt speech gestures. Refer to Figure 28 , compared with the four existing methods of Speech2Gesture, GTC, HA2G, and DiffGesture, the percentage of correct key points (PCK) and the pose-audio difference (GAD) of the embodiments of the present disclosure are significantly higher, indicating that the prompt speech gestures generated by the embodiments of the present disclosure are closer to the true prompt speech gestures and have better synchronization with the audio. At the same time, from ​ It can be seen that the Fréchet pose distance, mean absolute joint error, and mean acceleration difference of the prompted speech gestures generated by the embodiments of the present disclosure are lower than those of the prompted speech gestures generated by four existing methods, namely Speech2Gesture, GTC, HA2G, and DiffGesture. This indicates that the difference between the features of the prompted speech gestures generated by the embodiments of the present disclosure and the features of the true prompted speech gestures is smaller. At the same time, the positions of the key skeletal points in the generated prompted speech gestures are closer to those in the true prompted speech gestures, and the speeds of movement of each key skeletal point are also closer to those of the key skeletal points in the true prompted speech gestures. It is not difficult to know from ​ that, compared with the four existing methods such as Speech2Gesture, GTC, HA2G, and DiffGesture, the prompted speech gestures generated by the embodiments of the present disclosure are more realistic and the rhythm of the gesture movement is closer to the true prompted speech gestures, that is, the accuracy and fineness of the prompted speech gestures generated by the embodiments of the present disclosure are higher.

[0361] In addition, ​ also gives the performance parameters of generating prompted speech gestures in different variants of the embodiments of the present disclosure. It can be known from ​ that in the embodiments of the present disclosure, compared with only adding the rhythm adjustment pose difference label, adding the pose description guidance vector can further improve the percentage of correct key points (PCK) and the pose audio difference (GAD), and reduce the Fréchet pose distance, mean absolute joint error, and mean acceleration difference. This indicates that in the embodiments of the present disclosure, converting the input target text into a pose description through a large language model and then generating a pose description guidance vector to guide the diffusion process of the diffusion model can effectively improve the accuracy and fineness of the model in generating prompted speech gestures.

[0362] In addition, ​ also gives the performance parameters of generating prompted speech gestures using the prompted speech gesture generation method of the embodiments of the present disclosure after pre-training the first vector generation model by the method of contrastive learning. It can be known from ​ It can be seen that after the contrastive pre-training of the first vector generation model, the percentage of correct key points (PCK) and gesture-audio difference (GAD) of the prompt speech gestures generated by the prompt speech gesture generation method proposed in the embodiments of the present disclosure are higher than those without the contrastive pre-training of the first vector generation model. At the same time, the three parameters of Fréchet pose distance, mean absolute joint error, and mean acceleration difference are lower than those without the contrastive pre-training of the first vector generation model. This indicates that in the embodiments of the present disclosure, by performing contrastive pre-training on the first vector generation model, then using the first vector generation model to encode the gesture descriptions output by the large language model into gesture description guidance vectors, and then guiding the diffusion process of the diffusion model, the accuracy and fineness of the prompt speech gestures finally generated by the diffusion model can be effectively improved.

[0363] In addition ​ The performance parameters of the diffusion model for generating prompt speech gestures when the method of the embodiments of the present disclosure is tested on a mixed test set of data recorded by hearing-impaired and non-hearing-impaired persons are also given. It is not difficult to see that even when tested on a test set obtained by mixing the data recorded by hearing-impaired persons and the input of non-hearing-impaired persons, the embodiments of the present disclosure can still achieve better performance compared to the four existing methods of Speech2Gesture, GTC, HA2G, and DiffGesture.

[0364] Referring to ​ , ​ shows the performance parameters of the diffusion model for generating prompt speech gestures when only using the data recorded by non-hearing-impaired persons to train the diffusion model, and, when using the data recorded by hearing-impaired persons and the data recorded by non-hearing-impaired persons to train the diffusion model in combination, the performance parameters of the diffusion model for generating prompt speech gestures. It can be seen from ​ that using the data recorded by hearing-impaired persons and the data recorded by non-hearing-impaired persons to train the diffusion model in combination can enable the trained diffusion model to achieve better performance in the task of generating prompt speech gestures.

[0365] Referring to ​ , ​ shows the performance parameters of the diffusion model for generating prompt speech gestures in two cases: using the method of mel cepstral coefficients to extract audio features from the target speech and using the WavLM model to extract audio features from the target speech. It can be seen from ​ that compared with extracting audio features through the WavLM model, using mel cepstral coefficients to extract audio features from the target speech only achieves better performance in one parameter, namely the mean absolute joint error (MAJE). That is, compared with using mel cepstral coefficients to extract audio features from the target speech, using the WavLM model to extract audio features can further improve the accuracy of the diffusion model for generating prompt speech gestures.

[0366] ​

[0367] The following refers to ​ to describe in detail a specific usage process of the prompt voice gesture generation method according to the embodiments of the present disclosure. This process includes, but is not limited to, the following steps 3101 to 3121. In this process, it will be described that the terminal 110 and the server 140 each undertake a part of the tasks in the prompt voice gesture generation. However, those skilled in the art should understand that this process can also be independently completed only by the terminal 110 or by the server 140 independently.

[0368] 3101. Obtain the target input entered by the user through the terminal 110 and transmit it to the server 140.

[0369] 3102. The server 140 expands the target input into a target text and a target voice.

[0370] 3103. The server 140 obtains the overall rules for converting the target text into a gesture description.

[0371] 3104. The server 140 inputs the target text, the guiding language, and the overall rules into the large language model to generate a gesture description corresponding to the target text.

[0372] 3105. The server 140 inputs the gesture description into the first vector generation model to encode the gesture description into a gesture description guiding vector.

[0373] 3106. The server 140 inputs the target voice into the second vector generation model to encode the target voice into a voice guiding vector.

[0374] 3107. The server 140 initializes the prompt voice gesture vector to be diffused and initializes the step number to 1.

[0375] 3108. The server 140 inputs the concatenated vector of the prompt voice gesture vector to be diffused and the step number into the multi-head attention model to obtain a multi-head attention vector.

[0376] 3109. The server 140 superimposes and normalizes the concatenated vector and the multi-head attention vector to obtain a first normalized vector.

[0377] 3110. The server 140 inputs the first normalized vector into the first feed-forward neural network to obtain a first feed-forward vector.

[0378] 3111. The server 140 superimposes and normalizes the first feed-forward vector and the first normalized vector to obtain a second normalized vector.

[0379] 3112. The server 140 inputs the second normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain the diffused vector.

[0380] 3113. The server 140 inputs the diffused vector into the second feed-forward neural network to obtain the second feed-forward vector.

[0381] 3114. The server 140 superimposes and normalizes the second feed-forward vector and the diffused vector to obtain the diffusion noise corresponding to the step number.

[0382] 3-115. The server 140 subtracts the diffusion noise corresponding to the step number from the pose vector of the prompt speech to be diffused, and increments the step number by 1.

[0383] 3116. The server 140 determines whether the step number has reached the maximum preset number of steps.

[0384] 3117. If the step number has not reached the maximum preset number of steps, the server 140 uses the subtracted pose vector of the prompt speech to be diffused as the new pose vector of the prompt speech to be diffused, and returns to execute step 3108 until the step number reaches the maximum preset number of steps. At this time, the subtracted pose vector of the prompt speech to be diffused is used as the target pose vector of the prompt speech.

[0385] 3118. The server 140 generates a target prompt speech pose based on the target pose vector of the prompt speech.

[0386] 3119. The server 140 inputs the target speech into the rhythm adjustment pose difference generation model to obtain the rhythm adjustment pose difference.

[0387] 3120. The server 140 adjusts the target prompt speech pose with the rhythm adjustment pose difference to obtain the adjusted prompt speech pose.

[0388] 3121. The server 140 returns the target pose vector of the prompt speech to the terminal 110.

[0389] ​

[0390] It can be understood that although the steps in each of the above flowcharts are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0391] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or attribute information set, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the target object attribute information, it will obtain the separate permission or separate consent of the target object through methods such as pop-up windows or jumping to the confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.

[0392] Referring to ​ , ​ FIG. 3200 is a schematic structural diagram of a prompt voice gesture generation device provided by an embodiment of the present disclosure. The prompt voice gesture generation device 3200 includes:

[0393] According to one aspect of the present disclosure, there is provided a prompt voice gesture generation device, including:

[0394] A first acquisition unit 3210, configured to acquire a target input, where the target input includes at least one of a target text and a target voice;

[0395] A first generation unit 3220, configured to add the target input to a guiding language input large language model to obtain a gesture description describing a target prompt voice gesture;

[0396] A second generation unit 3230, configured to generate a gesture description guiding vector corresponding to the gesture description;

[0397] A diffusion unit 3240, configured to guide a diffusion model by using the gesture description guiding vector to enable the diffusion model to generate a target prompt voice gesture vector;

[0398] A third generation unit 3250, configured to generate a target prompt voice gesture based on the target prompt voice gesture vector.

[0399] Optionally, the first generation unit is specifically configured to:

[0400] Obtain the overall rule for converting the target input to the pose description;

[0401] Add the target input to the guiding language and the overall rule, and input them into the large language model to obtain the pose description.

[0402] Optionally, the prompt voice pose generation device further includes a first training unit (not shown), which is used to jointly train the guiding language and the overall rule in the following manner:

[0403] Obtain a first sample set, where the first sample set has first sample texts and pose description labels corresponding to the first sample texts;

[0404] Add each first sample text in the first sample set to the guiding language and the overall rule, and input them into the large language model to obtain a prediction result corresponding to the first sample text;

[0405] Calculate a first loss function based on the prediction result corresponding to the first sample text and the pose description label;

[0406] If the first loss function is less than a first threshold, stop training the guiding language and the overall rule; otherwise, adjust the guiding language and the overall rule, and return to the step of adding each first sample text in the first sample set to the guiding language and the overall rule, and inputting them into the large language model to obtain a prediction result corresponding to the sample text.

[0407] Optionally, in one embodiment, the target input includes a target text and a target voice, and the first generation unit is specifically configured to: add the target text to the guiding language and input it into the large language model;

[0408] The prompt voice pose generation device further includes a fourth generation unit (not shown), which is used to generate a voice guiding vector corresponding to the target voice;

[0409] The diffusion unit is specifically configured to: use the pose description guiding vector and the voice guiding vector to guide the diffusion model, so that the diffusion model generates a target prompt voice pose vector;

[0410] In one embodiment, the second generation unit is specifically configured to: input the pose description into a first vector generation model to obtain the pose description guiding vector;

[0411] The fourth generation unit is specifically configured to: input the target voice into a second vector generation model to obtain the voice guiding vector;

[0412] Optionally, the prompt voice gesture generation device further includes a second training unit (not shown) for training the first vector generation model in the following manner:

[0413] Obtain a set of sample pairs, where the sample pairs in the set of sample pairs include a plurality of first sample pairs and a plurality of second sample pairs. The first sample pairs include matching sample postures and sample posture descriptions, and the second sample pairs include the non-matching sample postures and the sample posture descriptions;

[0414] Convert the sample posture description into a sample posture description vector through the first vector generation model, and convert the sample posture into a sample posture vector through the third vector generation model;

[0415] Jointly train the first vector generation model and the third vector generation model to make the distance between the sample posture description vector and the sample posture vector of the first sample pair smaller, and the distance between the sample posture description vector and the sample posture vector of the second sample pair larger.

[0416] In one embodiment, the second training unit specifically is configured to:

[0417] Set the matching label of the first sample pair to 1 and the matching label of the second sample pair to 0;

[0418] Through the first vector generation model, based on the sample posture description vector and the sample posture vector of the sample pair, predict a first probability that the sample posture description and the sample posture in the sample pair match;

[0419] Through the third vector generation model, based on the sample posture description vector and the sample posture vector of the sample pair, predict a second probability that the sample posture and the sample posture description in the sample pair match;

[0420] Based on the matching label, the first probability, and the second probability of each sample pair, calculate a second loss function, and jointly train the first vector generation model and the third vector generation model based on the second loss function.

[0421] In one embodiment, the diffusion unit specifically is configured to:

[0422] In some embodiments, the guiding the diffusion model with the posture description guiding vector and the voice guiding vector to enable the diffusion model to generate a target prompt voice gesture vector includes:

[0423] Initialize the to-be-diffused prompt voice gesture vector as an initial prompt voice gesture vector, and initialize the step number as 1;

[0424] Inputting the gesture vector of the prompt voice to be diffused, the step number, the gesture description guide vector and the voice guide vector into the diffusion model to obtain the diffusion noise corresponding to the step number;

[0425] Use the diffusion noise corresponding to the step number to offset the voice gesture vector to be diffused, increase the step number by 1, and return to the step of inputting the voice gesture vector to be diffused, the step number, the gesture description guide vector and the voice guide vector into the diffusion model until the step number increases to the preset maximum number of steps.

[0426] Optionally, the prompt voice gesture generation device further includes a third training unit (not shown) for training the diffusion model in the following manner:

[0427] Acquire a second sample set, wherein each second sample in the second sample set includes a sample reference prompt speech gesture vector, a sample gesture description vector, and a sample speech guidance vector;

[0428] Performing a multi-step noise addition process on the sample reference prompt speech gesture vector, and recording the noise added at each step number as a noise label corresponding to the step number;

[0429] Based on the noisy prompt speech gesture vector obtained after the noisy processing of each step, generating a prompt speech gesture label corresponding to the step number;

[0430] Obtaining a rhythm adjustment posture difference label corresponding to the step sequence number;

[0431] The sample prompt speech posture vector to be diffused obtained after multiple steps of the noise addition processing is input into the diffusion model, and multiple steps of denoising processing are performed under the guidance of the sample posture description vector and the sample speech guide vector. Based on the denoising result of each step, the noise label corresponding to the step number, the prompt speech posture label, and the rhythm adjustment posture difference label, a third loss function is calculated, and the diffusion model is trained based on the third loss function.

[0432] In one embodiment, the third training unit is specifically configured to:

[0433] The third loss function is calculated based on the denoising result of each step, the noise label corresponding to the step number, the prompt voice posture label, and the rhythm adjustment posture difference label, including:

[0434] Calculating a first loss sub-function based on the predicted diffusion noise corresponding to the step number and the noise label;

[0435] Calculate a second loss sub - function based on the predicted prompt voice gesture corresponding to the step number and the prompt voice gesture label;

[0436] Calculate a third loss sub - function based on the predicted rhythm adjustment pose difference corresponding to the step number and the rhythm adjustment pose difference label;

[0437] Calculate the third loss function based on the first loss sub - function, the second loss sub - function, and the third loss sub - function.

[0438] In one embodiment, the third training unit is specifically configured to:

[0439] Calculate the first difference between the predicted diffusion noise corresponding to the step number and the noise label;

[0440] Calculate the square of the norm of the first difference as the first loss sub - function.

[0441] In one embodiment, the third training unit is specifically configured to:

[0442] Calculate the cosine distance between the predicted prompt voice gesture corresponding to the step number and the prompt voice gesture label;

[0443] Take the difference between 1 and the cosine distance as the second loss sub - function.

[0444] In one embodiment, the third training unit is specifically configured to:

[0445] Obtain the first weight of the first loss sub - function, the second weight of the second loss sub - function, and the third weight of the third loss sub - function;

[0446] Use the first weight, the second weight, and the third weight to perform a weighted sum of the first loss sub - function, the second loss sub - function, and the third loss sub - function to obtain the third loss function.

[0447] In one embodiment, the diffusion model is used to predict diffusion noise in the following manner:

[0448] Input a first proportion of the sample prompt voice gesture vectors to be diffused, the step number, the sample pose description vector, and the sample voice guidance vector into the diffusion model to obtain the first sub - predicted diffusion noise corresponding to the step number;

[0449] Input a second proportion of the sample prompt voice gesture vectors to be diffused, the step number, and the sample voice guidance vector into the diffusion model to obtain the second sub - predicted diffusion noise corresponding to the step number, where the sum of the first proportion and the second proportion is 1;

[0450] Weighted sum of the first sub-predicted diffusion noise and the second sub-predicted diffusion noise according to the first ratio and the second ratio to obtain the predicted diffusion noise.

[0451] In one embodiment, the diffusion unit is specifically configured to:

[0452] Input the concatenated vector of the prompt voice pose vector to be diffused and the step number into a multi-head attention model to obtain a multi-head attention vector;

[0453] Superimpose and normalize the concatenated vector and the multi-head attention vector to obtain a first normalized vector;

[0454] Input the first normalized vector, the pose description guidance vector, and the voice guidance vector into the diffusion model to obtain the diffusion noise corresponding to the step number.

[0455] In one embodiment, the diffusion unit is specifically configured to:

[0456] Input the first normalized vector into a first feed-forward neural network to obtain a first feed-forward vector;

[0457] Superimpose and normalize the first feed-forward vector and the first normalized vector to obtain a second normalized vector;

[0458] Input the second normalized vector, the pose description guidance vector, and the voice guidance vector into the diffusion model to obtain a diffused vector;

[0459] Input the diffused vector into a second feed-forward neural network to obtain a second feed-forward vector;

[0460] Superimpose and normalize the second feed-forward vector and the diffused vector to obtain the diffusion noise.

[0461] Optionally, the prompt voice pose generation device further includes a rhythm adjustment unit (not shown), and the rhythm adjustment unit is specifically configured to:

[0462] Input the target voice into a rhythm adjustment pose difference generation model to obtain a rhythm adjustment pose difference;

[0463] Adjust the target prompt voice pose with the rhythm adjustment pose difference to obtain an adjusted prompt voice pose.

[0464] Optionally, the prompt voice pose generation device further includes a fourth training unit (not shown), and the fourth training unit is used to train the rhythm adjustment pose difference generation model in the following manner:

[0465] Obtain a third sample set, including multiple sample video segments obtained by splitting a target sample video and sample voice segments corresponding to the sample video segments;

[0466] Extract sample object motion representations from the sample video segments;

[0467] Determine the average motion representation of the sample object motion representations;

[0468] Determine the difference between the sample object motion representation and the average motion representation as the rhythm adjustment pose difference label;

[0469] Input the sample voice segment corresponding to the sample video segment into the rhythm adjustment pose difference generation model to obtain the predicted rhythm adjustment pose difference;

[0470] Calculate a fourth loss function based on the predicted rhythm adjustment pose difference and the rhythm adjustment pose difference label, and train the rhythm adjustment pose difference generation model based on the fourth loss function.

[0471] In one embodiment, the third training unit and the fourth training unit are specifically configured to jointly train the diffusion model and the rhythm adjustment pose difference generation model in the following manner:

[0472] Obtain a fourth sample set, where the fourth sample set includes multiple multimodal samples, and each multimodal sample includes a second sample text, a sample voice, and a sample video;

[0473] Input the second sample text into the large language model with the guiding text to obtain a sample pose description;

[0474] Guide the diffusion model based on the first guiding vector corresponding to the sample pose description and the second guiding vector corresponding to the sample voice, so that the diffusion model generates a sample prompt voice pose vector, and generate a sample prompt voice pose based on the sample prompt voice pose vector;

[0475] Input the sample voice into the rhythm adjustment pose difference generation model to obtain a sample rhythm adjustment pose difference, and adjust the sample prompt voice pose with the sample rhythm adjustment pose difference to obtain an adjusted sample prompt voice pose;

[0476] Calculate a fifth loss function based on the comparison between the adjusted sample prompt voice pose and the prompt voice pose label extracted from the sample video, and train the rhythm adjustment pose difference generation model and the diffusion model based on the fifth loss function.

[0477] Refer to ​ , ​ A partial structural block diagram of the object terminal 110 for implementing the prompt voice gesture generation method according to an embodiment of the present disclosure. The terminal includes components such as a Radio Frequency (RF) circuit 3310, a memory 3315, an input unit 3330, a display unit 3340, a sensor 3350, an audio circuit 3360, a wireless fidelity (WiFi) module 3370, a processor 3380, and a power supply 3390. Those skilled in the art can understand that ​ The structure of the object terminal 110 shown does not limit a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0478] The RF circuit 3310 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 3380 for processing; in addition, the designed uplink data is sent to the base station.

[0479] The memory 3315 can be used to store software programs and modules. The processor 3380 executes various functional applications and data processing of the content terminal by running the software programs and modules stored in the memory 3315.

[0480] The input unit 3330 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the content terminal. Specifically, the input unit 3330 can include a touch panel 3331 and other input devices 3332.

[0481] The display unit 3340 can be used to display input information or provided information and various menus of the content terminal. The display unit 3340 can include a display panel 2941.

[0482] The audio circuit 3360, the speaker 3361, and the microphone 3362 can provide an audio interface.

[0483] In this embodiment, the processor 3380 included in the object terminal 110 can execute the prompt voice gesture generation method of the previous embodiment.

[0484] The object terminal 110 according to an embodiment of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc. Embodiments of the present invention can be applied to various scenarios, including but not limited to content recommendation, data screening, etc.

[0485] ​ Block diagram of a part of server 140 for implementing the method for generating prompt voice gestures according to an embodiment of the present disclosure. Server 140 may vary significantly due to configuration or performance differences, and may include one or more central processing units (CPUs) 3422 (for example, one or more processors) and a memory 3432, and one or more storage media 3430 (for example, one or more mass storage devices) for storing application programs 3442 or data 3444. Among them, the memory 3432 and the storage media 3430 may be transient storage or persistent storage. The program stored in the storage media 3430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 3422 may be configured to communicate with the storage media 3430 and execute a series of instruction operations in the storage media 3430 on the server.

[0486] Server 140 may further include one or more power supplies 3433, one or more wired or wireless network interfaces 3450, one or more input / output interfaces 3458, and / or one or more operating systems 3441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0487] The central processing unit 3422 in server 140 may be used to execute the method for generating prompt voice gestures according to an embodiment of the present disclosure.

[0488] The embodiment of the present disclosure further provides a computer-readable storage medium for storing program codes for executing the method for generating prompt voice gestures in the foregoing respective embodiments.

[0489] The embodiment of the present disclosure further provides a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes the method for generating prompt voice gestures as described above.

[0490] The terms "first", "second", "third", "fourth", etc. (if any) in the description of the present disclosure and the above-mentioned drawings are used to distinguish similar contents and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0491] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated contents and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated contents before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0492] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.

[0493] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0494] The unit described as a separate component may or may not be physically separated. The component presented as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0495] In addition, in the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0496] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0497] It should also be understood that the various embodiments provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0498] The above is a specific description of the embodiments of this disclosure, but this disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of this disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this disclosure.< / s> < / s> < / s> < / s> < / s> < / s>

Claims

1. A method for generating a prompt voice gesture, characterized in that, Including: Obtain a target input, where the target input includes at least one of a target text and a target voice; Add the target input to a guiding language input into a large language model to obtain a posture description describing a target prompt voice posture; Generate a posture description guiding vector corresponding to the posture description; Use the posture description guiding vector to guide a diffusion model to enable the diffusion model to generate a target prompt voice posture vector; Generate a target prompt voice posture based on the target prompt voice posture vector.

2. The prompting voice gesture generation method according to claim 1, wherein The step of adding the target input to a guiding language input into a large language model to obtain a posture description describing a target prompt voice posture includes: Obtain an overall rule for converting the target input to the posture description; Add the target input to the guiding language and the overall rule, and input them into the large language model to obtain the posture description.

3. The prompting voice gesture generation method according to claim 2, wherein The guiding language and the overall rule are jointly trained in the following manner: Obtain a first sample set, where the first sample set has a first sample text and a posture label corresponding to the first sample text; Add each first sample text of the first sample set to the guiding language and the overall rule, and input them into the large language model to obtain a prediction result corresponding to the first sample text; Calculate a first loss function based on the prediction result corresponding to the first sample text and the posture label; If the first loss function is less than a first threshold, stop training the guiding language and the overall rule, otherwise adjust the guiding language and the overall rule, and return to the step of adding each first sample text of the first sample set to the guiding language and the overall rule, and inputting them into the large language model to obtain a prediction result corresponding to the sample text.

4. The prompting voice gesture generation method according to claim 1, wherein The target input includes a target text and a target voice; The step of adding the target input to a guiding language input into a large language model includes: adding the target text to the guiding language and inputting it into the large language model; The step of using the posture description guiding vector to guide a diffusion model to enable the diffusion model to generate a target prompt voice posture vector includes: generating a voice guiding vector corresponding to the target voice: using the posture description guiding vector and the voice guiding vector to guide the diffusion model to enable the diffusion model to generate a target prompt voice posture vector.

5. The prompting voice gesture generation method according to claim 4, wherein, The step of generating a posture description guiding vector corresponding to the posture description includes: Input the posture description into a first vector generation model to obtain the posture description guiding vector; The step of generating a voice guiding vector corresponding to the target voice includes: Input the target voice into a second vector generation model to obtain the voice guiding vector.

6. The method for generating a prompt voice gesture according to claim 5, wherein The first vector generation model is pre-trained in the following manner: Obtain a sample pair set, where the sample pairs in the sample pair set include a plurality of first sample pairs and a plurality of second sample pairs, the first sample pair includes a matching sample posture and a sample posture description, and the second sample pair includes a non-matching sample posture and the sample posture description; Convert the sample pose description into a sample pose description vector through the first vector generation model, and convert the sample pose into a sample pose vector through the third vector generation model; Jointly train the first vector generation model and the third vector generation model to reduce the distance between the sample pose description vector and the sample pose vector of the first sample pair, and increase the distance between the sample pose description vector and the sample pose vector of the second sample pair.

7. The prompting voice gesture generation method according to claim 6, wherein The jointly training the first vector generation model and the third vector generation model to reduce the distance between the sample pose description vector and the sample pose vector of the first sample pair, and increase the distance between the sample pose description vector and the sample pose vector of the second sample pair includes: Set the matching label of the first sample pair to 1 and the matching label of the second sample pair to 0; Through the first vector generation model, based on the sample pose description vector and the sample pose vector of the sample pair, predict the first probability that the sample pose description and the sample pose in the sample pair match; Through the third vector generation model, based on the sample pose description vector and the sample pose vector of the sample pair, predict the second probability that the sample pose and the sample pose description in the sample pair match; Based on the matching label, the first probability, and the second probability of each sample pair, calculate a second loss function, and jointly train the first vector generation model and the third vector generation model based on the second loss function.

8. The prompting voice gesture generation method according to claim 4, wherein The guiding the diffusion model using the pose description guiding vector and the speech guiding vector to enable the diffusion model to generate a target prompt speech pose vector includes: Initialize the to-be-diffused prompt speech pose vector as an initial prompt speech pose vector, and initialize the step number as 1; Input the to-be-diffused prompt speech pose vector, the step number, the pose description guiding vector, and the speech guiding vector into the diffusion model to obtain the diffusion noise corresponding to the step number; Subtract the diffusion noise corresponding to the step number from the to-be-diffused prompt speech pose vector, increment the step number by 1, and return to the step of inputting the to-be-diffused prompt speech pose vector, the step number, the pose description guiding vector, and the speech guiding vector into the diffusion model until the step number increases to a preset maximum number of steps.

9. The prompting voice gesture generation method according to claim 8, wherein The diffusion model is trained in the following manner: Obtain a second sample set, where each second sample in the second sample set includes a sample reference prompt speech pose vector, a sample pose description vector, and a sample speech guiding vector; Perform multi-step noise addition processing on the sample reference prompt speech pose vector, and record the noise added at each step number as the noise label corresponding to the step number; Generate a prompt speech pose label corresponding to the step number based on the noise-added prompt speech pose vector obtained after each step of the noise addition processing; Obtain a rhythm adjustment pose difference label corresponding to the step number; Input the sample prompt speech pose vector obtained after the multi-step noise addition process into the diffusion model, and perform multi-step denoising under the guidance of the sample pose description vector and the sample speech guidance vector. Calculate the third loss function based on the denoising result of each step, the noise label corresponding to the step number, the prompt speech pose label, and the rhythm adjustment pose difference label, and train the diffusion model based on the third loss function.

10. The prompting voice gesture generation method according to claim 9, wherein The denoising result of each step includes: the predicted diffusion noise predicted by the diffusion model corresponding to the step number, the predicted prompt speech pose generated based on the predicted diffusion noise corresponding to the step number, and the predicted rhythm adjustment pose difference corresponding to the step number; Calculating the third loss function based on the denoising result of each step, the noise label corresponding to the step number, the prompt speech pose label, and the rhythm adjustment pose difference label includes: Calculate the first loss sub-function based on the predicted diffusion noise corresponding to the step number and the noise label; Calculate the second loss sub-function based on the predicted prompt speech pose corresponding to the step number and the prompt speech pose label; Calculate the third loss sub-function based on the predicted rhythm adjustment pose difference corresponding to the step number and the rhythm adjustment pose difference label; Calculate the third loss function based on the first loss sub-function, the second loss sub-function, and the third loss sub-function.

11. The prompting voice gesture generation method according to claim 10, wherein Calculating the third loss function based on the first loss sub-function, the second loss sub-function, and the third loss sub-function includes: Obtain the first weight of the first loss sub-function, the second weight of the second loss sub-function, and the third weight of the third loss sub-function; Use the first weight, the second weight, and the third weight to perform a weighted sum of the first loss sub-function, the second loss sub-function, and the third loss sub-function to obtain the third loss function.

12. The prompting voice gesture generation method according to claim 8, wherein, Inputting the to-be-diffused prompt speech pose vector, the step number, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain the diffusion noise corresponding to the step number includes: Input the concatenated vector of the to-be-diffused prompt speech pose vector and the step number into the multi-head attention model to obtain the multi-head attention vector; Superimpose and normalize the concatenated vector and the multi-head attention vector to obtain the first normalized vector; Input the first normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain the diffusion noise corresponding to the step number.

13. The prompting voice gesture generation method according to claim 12, wherein Inputting the first normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain the diffusion noise corresponding to the step number includes: Input the first normalized vector into the first feed-forward neural network to obtain the first feed-forward vector; Superimpose and normalize the first feed-forward vector and the first normalized vector to obtain the second normalized vector; Input the second normalized vector, the pose description guidance vector, and the speech guidance vector into the diffusion model to obtain a diffused vector; Input the diffused vector into a second feed-forward neural network to obtain a second feed-forward vector; Superimpose and normalize the second feed-forward vector and the diffused vector to obtain the diffusion noise.

14. The prompting voice gesture generation method according to claim 1, wherein After generating a target prompt speech pose based on the target prompt speech pose vector, the prompt speech pose generation method further includes: Input the target speech into a rhythm adjustment pose difference generation model to obtain a rhythm adjustment pose difference; Adjust the target prompt speech pose with the rhythm adjustment pose difference to obtain an adjusted prompt speech pose.

15. The method for generating a prompt voice gesture according to claim 14, wherein The rhythm adjustment pose difference generation model is trained in the following manner: Obtain a third sample set, including a plurality of sample video segments obtained by splitting a target sample video, and sample speech segments corresponding to the sample video segments; Extract sample object motion representations from the sample video segments; Determine the average motion representation of the sample object motion representations; Determine the difference between the sample object motion representation and the average motion representation as the rhythm adjustment pose difference label; Input the sample speech segments corresponding to the sample video segments into the rhythm adjustment pose difference generation model to obtain the predicted rhythm adjustment pose difference; Calculate a fourth loss function based on the predicted rhythm adjustment pose difference and the rhythm adjustment pose difference label, and train the rhythm adjustment pose difference generation model based on the fourth loss function.

16. The prompting voice gesture generation method according to claim 14, wherein The rhythm adjustment pose difference generation model and the diffusion model are jointly trained in the following manner: Obtain a fourth sample set, which includes a plurality of multimodal samples, and each multimodal sample includes a second sample text, a sample speech, and a sample video; Add the second sample text to the guiding text and input it into the large language model to obtain a sample pose description; Based on the first guiding vector corresponding to the sample pose description and the second guiding vector corresponding to the sample speech, guide the diffusion model to enable the diffusion model to generate a sample prompt speech pose vector, and generate a sample prompt speech pose based on the sample prompt speech pose vector; Input the sample speech into the rhythm adjustment pose difference generation model to obtain a sample rhythm adjustment pose difference, and adjust the sample prompt speech pose with the sample rhythm adjustment pose difference to obtain an adjusted sample prompt speech pose; Calculate a fifth loss function based on the comparison between the adjusted sample prompt speech pose and the prompt speech pose label extracted from the sample video, and train the rhythm adjustment pose difference generation model and the diffusion model based on the fifth loss function.

17. A prompting voice gesture generation device, characterized in that, Including: A first acquisition unit for acquiring a target input, where the target input includes at least one of a target text and a target speech; A first generation unit for adding the target input to the guiding text and inputting it into the large language model to obtain a pose description describing the target prompt speech pose; A second generation unit for generating a pose description guidance vector corresponding to the pose description; A diffusion unit for guiding a diffusion model using the pose description guidance vector to cause the diffusion model to generate a target prompt speech pose vector; A third generation unit for generating a target prompt speech pose based on the target prompt speech pose vector.

18. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the prompt speech pose generation method according to any one of claims 1 to 16.

19. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the prompt speech pose generation method according to any one of claims 1 to 16.

20. A computer program product, the computer program product includes a computer program, the computer program is read and executed by a processor of a computer device, so that the computer device executes the prompt speech pose generation method according to any one of claims 1 to 16.