Method and device for generating posture of humanoid robot, robot, medium and product

By generating the gestures of humanoid robots through a diffusion model, the problem of lack of autonomy and naturalness of humanoid robots is solved, and a more natural human-computer interaction experience is achieved.

CN120755870APending Publication Date: 2025-10-10GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510984486.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing humanoid robots lack autonomy in human-computer interaction, have fixed gestures, low naturalness, and poor interactive experience.

Method used

A trained diffusion model is used to generate the posture of a humanoid robot. By obtaining inputs such as speech data, Mel spectrum, deep semantic features, and Gaussian noise, gestures that are closer to human performance are generated. Special action markers are combined to ensure reasonable performance when there is no speech output.

Benefits of technology

It improves the naturalness and intimacy of human-computer interaction, increases the diversity and randomness of actions, reduces the amount of calculation, and improves the efficiency of action generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120755870A_ABST
    Figure CN120755870A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robots, and discloses a posture generation method and device of a humanoid robot, the robot, a medium and a product, and the method comprises the steps: obtaining to-be-output first voice data, a first special action mark, Gaussian noise and a time step; according to the first voice data, a first Mel spectrum and a first deep semantic feature are determined, and the first deep semantic feature is the first voice data processed through a self-supervised voice representation learning model; the first Mel spectrum, the first depth semantic feature, the first special action mark, the Gaussian noise and the time step are input into a posture generation model, first action data are determined according to output of the posture generation model, and the posture generation model is a trained diffusion model; and when the first voice data is output, synchronously controlling the humanoid robot to generate a gesture action represented by the first action data. According to the method, more diversified actions can be generated in the distribution, and the naturalness and the closeness of man-machine interaction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to a posture generation method, device, robot, medium and product of a humanoid robot. Background Art

[0002] Humanoid robots are robots that mimic human appearance and behavior and possess a high level of intelligence. Compared to traditional industrial and service robots, humanoid robots share similar perception, limb structure, and movement patterns, playing a vital role in applications such as manufacturing, social services, and specialized operations.

[0003] Humanoid robots typically possess the ability to interact with humans in real time, quickly understanding and effectively executing commands given through language, gestures, and other means. To enhance the interactive experience, humanoid robots often employ gestures during interactions with humans. However, these gestures are typically recorded and played back or remotely controlled, resulting in fixed gestures and a lack of autonomy, making human-robot interaction less natural. Summary of the Invention

[0004] In view of this, the present invention provides a method, device, robot, medium and product for generating posture of a humanoid robot to solve the problem that existing humanoid robots lack autonomy and low naturalness during human-machine interaction.

[0005] In a first aspect, the present invention provides a posture generation method for a humanoid robot, the method comprising: obtaining first voice data to be output, a first special action marker, Gaussian noise and a time step, wherein the first special action marker is used to characterize the style of the generated gesture action or a preset special action; determining a first Mel spectrum and a first deep semantic feature based on the first voice data, wherein the first deep semantic feature is the first voice data processed by a self-supervised speech representation learning model; inputting the first Mel spectrum, the first deep semantic feature, the first special action marker, Gaussian noise and the time step into a posture generation model, determining the first action data based on the output of the posture generation model, wherein the posture generation model is a diffusion model that has completed training; when outputting the first voice data, synchronously controlling the humanoid robot to generate the gesture action represented by the first action data.

[0006] The gesture generation model in this embodiment is a trained diffusion model. Due to the randomness of the model's initial noise, it can enhance the randomness and diversity of generated gestures, making the humanoid robot behave more like a human and improving the naturalness and intimacy of human-machine interaction. Furthermore, adding a first special action marker to the input enables the output of a special action even when the first voice data is empty, ensuring that gesture generation maintains reasonable performance even when there is no voice output. Furthermore, converting the first voice data into a first mel-spectrogram and first deep semantic features before input not only enables the gesture generation model to more accurately generate reasonable actions and improve generalization, but also reduces the model's computational complexity and improves gesture generation efficiency.

[0007] In an optional embodiment, before inputting the first Mel spectrum, the first deep semantic feature, the first special action marker, the Gaussian noise and the time step into the posture generation model, the method further includes: constructing a posture generation model based on a sample database, wherein the sample database includes a plurality of training samples, the training samples include second voice data, second action data that changes with the second voice data and a second special action marker, and the duration of the second voice data is the same as the duration of the first voice data.

[0008] In this embodiment, before the first Mel spectrum, the first deep semantic feature, the first special action marker, the Gaussian noise and the time step are input into the posture generation model, the posture generation model is constructed through a sample database. The sample database contains a variety of training samples, which allows the diffusion model to capture the mapping relationship between the posture and inputs such as the Mel spectrum, semantic features, special action markers, etc., thereby improving the accuracy of the generated posture.

[0009] In an optional embodiment, a gesture generation model is constructed based on a sample database, including: extracting a target training sample from the sample database, wherein the target training sample is one of a plurality of training samples, and the target training sample includes target speech data, target action data that changes with the target speech data, and a target special action tag; determining a target Mel spectrum and a target deep semantic feature based on the target speech data; performing noise processing on the target action data based on the time step to obtain target noise data; and training a diffusion model based on the target Mel spectrum, the target deep semantic feature, the target noise data, the time step, the target special action tag, and the target action data to obtain a gesture generation model.

[0010] In an optional embodiment, before training the diffusion model, the method also includes: determining whether the target Mel spectrum includes 0; when the target Mel spectrum includes 0, mapping the area where the target Mel spectrum is 0 and the part of the deep semantic feature corresponding to the area where the target Mel spectrum is 0 into learnable parameters to obtain an updated target Mel spectrum and an updated target deep semantic feature; training the diffusion model according to the target Mel spectrum, the target deep semantic feature, the target noise data, the time step, the target special action tag and the target action data, including: training the diffusion model according to the updated target Mel spectrum, the updated target deep semantic feature, the target noise data, the time step, the target special action tag and the target action data.

[0011] In this embodiment, the area where the Mel spectrum is 0 is replaced with a learnable parameter, which enables the diffusion model to fully associate the empty audio with the non-speaking state action. In this way, even if only a part of a segment of length L is empty audio, the diffusion model can generate the corresponding non-speaking action.

[0012] In an optional embodiment, before training the diffusion model, the method also includes: determining whether the target Mel spectrum is all 0; when the target Mel spectrum is all 0, mapping the target Mel spectrum and the target deep semantic features into learnable parameters to obtain first model parameters and second model parameters; training the diffusion model according to the target Mel spectrum, target deep semantic features, target noise data, time step, target special action mark and target action data, including: training the diffusion model according to the first model parameters, the second model parameters, the target noise data, the time step, the target special action mark and the target action data.

[0013] In this embodiment, the mel-spectrograms with all zeros are replaced with learnable parameters, so that the diffusion model can learn more coherent non-speech state actions.

[0014] In an optional embodiment, before extracting the target training sample from the sample database, the method also includes: obtaining at least one third voice data and third action data corresponding to the third voice data, the duration of the third voice data being greater than the duration of the second voice data; slicing the at least one third voice data to obtain multiple training samples.

[0015] In an optional embodiment, the third action data includes an action start stage, an action hold stage, and an action end stage. Before extracting the target training sample from the sample database, the method also includes: setting the weight of the first training sample to a first preset value, setting the weight of the second training sample to a second preset value, and setting the weight of the third training sample to a third preset value, wherein the first training sample is a training sample corresponding to the second motion data including the action start stage, the second training sample is a training sample corresponding to the second motion data including the action hold stage, and the third training sample is a training sample corresponding to the second motion data including the action end stage, the first preset value is greater than the second preset value, and the second preset value is greater than the third preset value.

[0016] In this embodiment, by setting different weights for training samples and increasing the weight of training samples at the beginning of the movement, the posture generation model can have a greater probability of performing special movements in idle scenes and ensure the continuity of the movements.

[0017] In an optional implementation, the first voice data is empty audio, and the first special action mark is determined according to a preset frequency of the special action.

[0018] In a second aspect, the present invention provides a posture generation device for a humanoid robot, the device comprising: an acquisition module for acquiring first voice data, a first special action marker, Gaussian noise and a time step to be output, wherein the first special action marker is used to characterize the style of the generated gesture action or a preset special action; a feature determination module for determining a first Mel spectrum and a first deep semantic feature based on the first voice data, wherein the first deep semantic feature is the first voice data processed by a self-supervised speech representation learning model; an action determination module for inputting the first Mel spectrum, the first deep semantic feature, the first special action marker, Gaussian noise and the time step into a posture generation model, and determining the first action data according to the output of the posture generation model; an output module for synchronously controlling the humanoid robot to generate the gesture action represented by the first action data when outputting the first voice data.

[0019] In an optional embodiment, the device also includes: a construction module for constructing a posture generation model based on a sample database, wherein the sample database includes multiple training samples, the training samples include second voice data, second action data that changes with the second voice data, and a second special action mark, and the duration of the second voice data is the same as the duration of the first voice data.

[0020] In an optional embodiment, the construction module includes: an extraction unit for extracting target training samples from a sample database, wherein the target training sample is one of multiple training samples, and the target training sample includes target speech data, target action data that changes with the target speech data, and a target special action mark; a first determination unit for determining the target Mel spectrum and target deep semantic features based on the target speech data; a noise addition unit for performing noise processing on the target action data according to the time step to obtain target noise data; a training unit for training the diffusion model according to the target Mel spectrum, target deep semantic features, target noise data, time step, target special action mark and target action data to obtain a posture generation model.

[0021] In an optional embodiment, the device also includes: a first determination module, used to determine whether the target Mel spectrum includes 0; a first mapping module, used to, when the target Mel spectrum includes 0, map the area where the target Mel spectrum is 0 and the part of the deep semantic feature corresponding to the area where the target Mel spectrum is 0 into a learnable parameter, to obtain an updated target Mel spectrum and an updated target deep semantic feature; the training unit includes: a first sub-training unit, used to train the diffusion model based on the updated target Mel spectrum, the updated target deep semantic feature, the target noise data, the time step, the target special action mark and the target action data.

[0022] In an optional embodiment, the device also includes: a second determination module, used to determine whether the target Mel spectrum is all 0; a second mapping module, used to map the target Mel spectrum and the target deep semantic features into learnable parameters when the target Mel spectrum is all 0, to obtain first model parameters and second model parameters; the training unit includes: a second sub-training unit, used to train the diffusion model according to the first model parameters, the second model parameters, the target noise data, the time step, the target special action mark and the target action data.

[0023] In an optional embodiment, the device also includes: a data acquisition module, used to obtain at least one third voice data and third action data corresponding to the third voice data, the duration of the third voice data is greater than the duration of the second voice data; a slicing module, used to slice the at least one third voice data to obtain multiple training samples.

[0024] In an optional embodiment, the third action data includes an action start stage, an action hold stage, and an action end stage, and the device also includes: a processing module, used to set the weight of the first training sample to a first preset value, the weight of the second training sample to a second preset value, and the weight of the third training sample to a third preset value, wherein the first training sample is a training sample corresponding to the second motion data including the action start stage, the second training sample is a training sample corresponding to the second motion data including the action hold stage, and the third training sample is a training sample corresponding to the second motion data including the action end stage, the first preset value is greater than the second preset value, and the second preset value is greater than the third preset value.

[0025] In an optional implementation, the first voice data is empty audio, and the first special action mark is determined according to a preset frequency of the special action.

[0026] In a third aspect, the present invention provides a humanoid robot comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the humanoid robot posture generation method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0027] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a humanoid robot to execute the humanoid robot posture generation method of the first aspect or any corresponding embodiment thereof.

[0028] In a fifth aspect, the present invention provides a computer program product comprising computer instructions for causing a humanoid robot to execute the method for generating a posture of a humanoid robot according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 is a flow chart of a method for generating a posture of a humanoid robot according to an embodiment of the present invention;

[0031] Figure 2 is a schematic diagram of a humanoid robot outputting gesture actions in an idle scenario according to an embodiment of the present invention;

[0032] Figure 3 is a schematic diagram of a humanoid robot outputting gesture actions in an idle scenario according to an embodiment of the present invention;

[0033] Figure 4 is a flow chart of another method for generating a posture of a humanoid robot according to an embodiment of the present invention;

[0034] Figure 5 is a schematic diagram of a process of constructing a gesture generation model according to an embodiment of the present invention;

[0035] Figure 6 is a schematic structural diagram of a diffusion model according to an embodiment of the present invention;

[0036] Figure 7 is a schematic diagram of third action data according to an embodiment of the present invention;

[0037] Figure 8 1 is a schematic diagram of partial replacement of learnable parameters whose Mel spectrum is 0 according to an embodiment of the present invention;

[0038] Figure 9 is a schematic diagram of a mel spectrum processing process according to an embodiment of the present invention;

[0039] Figure 10 is a schematic diagram of replacing learnable parameters when the Mel spectrum is all 0 according to an embodiment of the present invention;

[0040] Figure 11 is a structural block diagram of a posture generating device for a humanoid robot according to an embodiment of the present invention;

[0041] Figure 12 4 is a schematic diagram of the hardware structure of the humanoid robot according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. According to the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of the present invention.

[0043] As mentioned in the background art, the gestures output by humanoid robots when interacting with users are limited to pre-recorded movements or operator manipulation. They lack autonomy, have a strong mechanical feel, and the interactive experience appears stiff and not intimate enough.

[0044] In view of this, the present invention provides a posture generation method, device, robot, medium and product for a humanoid robot. A trained diffusion model (posture generation model) is configured on the humanoid robot, and motion data is obtained using the posture generation model. Due to the randomness of the initial noise of the model, more diverse movements can be generated within the distribution compared to recorded broadcasts, thereby improving the naturalness and intimacy of human-computer interaction.

[0045] The posture generation method of the humanoid robot provided by the present invention is described in detail below with reference to the accompanying drawings. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a humanoid robot such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0046] In this embodiment, a method for generating a posture of a humanoid robot is provided, which can be used for a humanoid robot. Figure 1 FIG. 1 is a flow chart of a method for generating a posture of a humanoid robot according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0047] Step S101 : Acquire first speech data to be output, a first special action marker, Gaussian noise, and a time step.

[0048] The first special action tag is used to characterize the style of the generated gesture or a preset special action. The preset special action can be standing naturally with a slight shake, waving, giving a thumbs up, moving fingers, or looking down at hands.

[0049] Specifically, in a dialogue scenario, the first special action mark is used to characterize the style of the gesture action generated by the humanoid robot when outputting the first voice data. The style reflects the personality characteristics of the humanoid robot (such as lively or rigorous). The first special action mark can be pre-configured in the humanoid robot by the staff; in an idle (static) scenario, the first special action mark is used to characterize a special action generated by the humanoid robot when outputting the first voice data. Multiple special actions are pre-configured in the humanoid robot, and multiple special actions appear at a preset frequency. The first special action mark can be randomly generated according to the preset special action frequency.

[0050] For example, a humanoid robot collects audio from an environment in real time. If no voice data from a user is detected within a long period (a first time period), the robot is determined to be in an idle state. If the robot detects voice data from a user within a long period (a first time period), the robot is determined to be in a conversation state. The first time period is a preset value that is greater than the duration of the first voice data to be output. For example, the duration of the first voice data to be output may be 5 seconds or 10 seconds, and the duration of the first time period may be 5 minutes or 10 minutes.

[0051] In other embodiments, the time periods of the dialogue scene and the idle scene can also be configured by the staff. For example, in the second time period (such as 10:00-12:00, 16:00-20:00, etc.), the humanoid robot is in the dialogue scene; in the third time period (such as 13:00-15:30), the humanoid robot is in the idle scene.

[0052] The first voice data is the voice data of a preset length output by the humanoid robot to the user at the moment. In an idle scenario, the first voice data may be empty audio of a preset length. In a conversation scenario, if the humanoid robot recognizes the user's voice data within a preset length (such as "Hello, Xiao A, how do I get to xx shop from my current location"), the recognized user's voice data is converted into text, and the text is input into a large voice model (machine learning model) to obtain a reply text (such as go straight for 200m from the current location, turn left, and then go forward 300m to see xx shop). After obtaining the reply text, the reply text is converted into voice data based on text-to-speech (TTS) technology, and the voice data is the first voice data. If the user's voice data is not recognized within the preset length, the first voice data may also be empty audio of a preset length.

[0053] The preset duration can be configured by the staff, for example, the preset duration can be 5s, 8s or 10s, etc. It should be understood that if the duration of the voice data converted from the reply text is less than the preset duration, the remaining duration will be padded with 0; if the duration of the voice data converted from the reply text is greater than the preset duration, it will be output in batches.

[0054] Gaussian noise and time step are both inputs to the subsequent posture generation model. Gaussian noise is a random noise that obeys Gaussian distribution (normal distribution). The time step t is a positive integer, corresponding to the number of denoising steps of randomly sampled Gaussian noise.

[0055] Specifically, the posture generation model is a diffusion model that completes the training. The diffusion model is a deep learning method based on probabilistic generative modeling. The core idea is to gradually add noise to the data to convert it into random noise (usually Gaussian white noise), and then gradually remove the noise through an inverse process to generate new data samples. Usually the input is noise x t and time step t, the output is the noise x corresponding to the previous time step t-1 Therefore, for the T-step diffusion model, a Gaussian noise x is randomly generated when applied T , through T-step denoising of the diffusion model, the unnoised action sequence x0 can be obtained.

[0056] The diffusion model uses multi-level denoising to preserve more motion characteristics, making movements closer to real human motion patterns. Furthermore, the generation process is guided by the randomness of the noise, which more comprehensively captures the natural variations in movement, ensuring a higher degree of naturalness and fluidity in the generated movements. Furthermore, because the inverse process starts from random noise, different yet plausible movements can be generated under the same conditions, ensuring diversity in the generated movements.

[0057] Step S102: Determine a first mel spectrum and a first deep semantic feature according to the first speech data.

[0058] The first deep semantic feature is the first speech data processed by a self-supervised speech representation learning (Hidden unit BERT, HuBERT) model. Specifically, the first speech data is input into a trained HuBERT model, and the output of the trained HuBERT model is the first deep semantic (HuBERT) feature.

[0059] Specifically, the HuBERT model is a self-supervised speech representation learning model based on the Transformer architecture. It performs self-supervised speech representation learning through masked predictions from hidden units. It uses an offline clustering step to provide aligned target labels for a BERT-like prediction loss. It first extracts features from unlabeled speech data and then divides it into different categories through clustering. For example, it uses the same strategy as SpanBERT and wav2vec 2.0 to generate masks, applying a prediction loss to the masked regions. This forces the model to learn acoustic and language models from continuous input to obtain a representation of the sound.

[0060] The first Mel spectrum is the Mel spectrum (basic acoustic features) corresponding to the first speech data. The Mel spectrum is a method of converting the spectrum of an audio signal into a Mel scale representation, which is mainly used for speech and audio signal processing. The principle of the Mel spectrum is to perform nonlinear conversion of the frequency through a Mel filter group to better reflect the human auditory system's perception of the frequency. The calculation of the Mel spectrum usually involves fast Fourier transform (FFT) and overlapping windowing of the frequency. The final Mel spectrum can be used for feature extraction and pattern recognition. The calculation method from the first speech data to the first Mel spectrum is a conventional calculation method in this field and is not described in detail here.

[0061] Step S103: input the first Mel spectrum, the first depth semantic feature, the first special action marker, the Gaussian noise and the time step into a posture generation model, and determine the first action data according to the output of the posture generation model.

[0062] Specifically, the output of the posture generation model is the first action data, which is used to describe the limb movement information of the humanoid robot when outputting the first voice data. The first motion data may include information such as the angle (or coordinate), speed, moment and torque of the joints of the humanoid robot at each moment corresponding to the first voice data.

[0063] Mel-spectrograms are more robust to noise and environmental changes, enabling the gesture generation model to maintain stable performance in different audio environments. Deep semantic features allow the gesture generation model to learn the general semantic information behind the speech, rather than being limited to the surface features of a specific speech sample. Therefore, compared to directly inputting the first speech data, converting the first speech data into the first Mel-spectrogram and the first deep semantic features and then inputting them, the gesture generation model can more accurately generate reasonable actions based on the Mel-spectrogram and deep semantic features, thereby improving generalization capabilities.

[0064] Furthermore, the Mel-spectrogram compresses the original linear spectrum into a smaller dimensional space, reducing the amount of data. This reduces the model's computational complexity, improves operational efficiency, and accelerates action generation. Deep semantic features provide a richer source of information for gesture generation. Based on semantic understanding, the model can combine different Mel-spectrogram features to generate a variety of actions that align with the semantics and speech rhythm, meeting diverse needs.

[0065] Step S104: When outputting the first voice data, synchronously controlling the humanoid robot to generate a gesture action represented by the first action data.

[0066] Specifically, after determining the first motion data, the humanoid robot is controlled to output the first voice data while the limbs of the humanoid robot are controlled based on the first motion data to make the humanoid robot output the gesture action corresponding to the first action data synchronously.

[0067] The posture generation method of the humanoid robot provided in this embodiment, after obtaining the first voice data to be output, the first special action mark, the Gaussian noise and the time step, first processes the first voice data to obtain the first mel spectrum and the first deep semantic feature, then inputs the first mel spectrum, the first deep semantic feature, the first special action mark, the Gaussian noise and the time step into the posture generation model, determines the first action data according to the output of the posture generation model, and finally controls the humanoid robot to output the gesture action based on the first action data when the first voice data is output.

[0068] The posture generation model in this embodiment is a trained diffusion model. Due to the randomness of the initial noise of the model, the randomness and diversity of the generated action can be improved, so that the humanoid robot has a more human-like performance, and the naturalness and closeness of human-computer interaction are improved. Moreover, by inputting the first special action mark, special actions can be output when the first voice data is empty audio, ensuring that the gesture generation can still maintain reasonable performance when there is no voice output. In addition, after the first voice data is converted into the first mel spectrum and the first deep semantic feature, it is inputted, which not only enables the posture generation model to generate reasonable actions more accurately and improve the generalization ability, but also reduces the model calculation amount and improves the action generation efficiency.

[0069] The processes of the humanoid robot outputting gesture actions in the idle scene and the dialogue scene will be described below with reference to the accompanying drawings.

[0070] As shown in FIG. 1, Figure 2 In the idle scene, the humanoid robot collects voice data in the environment for a preset time length to obtain an empty audio of a preset time length, and converts the empty audio of the preset time length into a mel spectrum and a deep semantic feature. At the same time, the humanoid robot generates a special action mark based on the frequency of a preset special action. Then, the mel spectrum, the deep semantic feature, the randomly generated Gaussian noise x T , the preset time step T and the special action mark are inputted into the posture generation model, and the action data can be obtained according to the output of the posture generation model, and finally the humanoid robot is controlled to output a gesture action of a preset time length.

[0071] As shown in FIG. 2, Figure 3As shown in the figure, in a conversation scenario, the humanoid robot collects voice data in an environment of a preset duration, obtains the user's voice data (or empty audio) of the preset duration, and based on the collected user's voice data (or empty audio) of the preset duration, obtains TTS audio (or empty audio) of the preset duration, and converts the TTS audio (or empty audio) of the preset duration into Mel spectrum and deep semantic features. At the same time, the humanoid robot generates special action tags based on the preset speaking style. Then, the Mel spectrum, deep semantic features, and randomly generated Gaussian noise x are combined into a single word. T , the preset time step T and special action mark are input into the posture generation model, and the action data can be obtained according to the output of the posture generation model. Finally, the humanoid robot is controlled to synchronously output TTS audio (or empty audio) of the preset duration and the gesture action corresponding to the action data.

[0072] In this embodiment, another method for generating a posture of a humanoid robot is provided, which can be used for the above-mentioned humanoid robot. Figure 4 FIG. 1 is a flow chart of another method for generating a posture of a humanoid robot according to an embodiment of the present invention. Figure 4 As shown, the method includes the following steps:

[0073] Step S401: construct a posture generation model based on a sample database.

[0074] The sample database includes multiple training samples, which include second voice data, second action data that changes with the second voice data, and a second special action mark. The duration of the second voice data is the same as that of the first voice data.

[0075] Specifically, multiple motion capture actors can record action data based on TTS voice, each motion capture actor represents a style, and then multiple training samples are obtained based on the TTS voice, the action data corresponding to the TTS voice, and the motion capture actor corresponding to the TTS voice.

[0076] Specifically, the above step S401 includes:

[0077] Step S4011: extract target training samples from the sample database.

[0078] The target training sample is one of multiple training samples, and the target training sample includes target speech data, target action data that changes with the target speech data, and a target special action tag.

[0079] Specifically, a training sample can be randomly extracted from the sample database as the target training sample. At this time, the second voice data in the extracted training sample is the target voice data, the second motion data corresponding to the second voice data in the extracted training sample is the target motion data, and the second special action mark in the extracted training sample is the target special action mark.

[0080] Step S4012: Determine target Mel spectrum and target deep semantic features based on target speech data.

[0081] Specifically, after obtaining the target speech data, it is fed into the trained HuBERT model. The output of the trained HuBERT model is the target deep semantic (HuBERT) feature. The calculation method from the target speech data to the target mel-spectrogram is the same as the calculation method from the first speech data to the first mel-spectrogram, and is not explained here.

[0082] Step S4013: performing noise processing on the target motion data according to the time step to obtain target noise data.

[0083] Specifically, the diffusion model includes a forward process and a reverse process. The forward process gradually adds noise to the original data, while the reverse process is the denoising process. For example, if the time step is T, after determining the target motion data, the target motion data after adding T steps of noise is the target noise data. The process of adding noise to the target motion data is a conventional method in the field and will not be explained in detail here.

[0084] Step S4014 , training the diffusion model based on the target Mel spectrum, target depth semantic features, target noise data, time step, target special action tag and target action data to obtain a posture generation model.

[0085] Specifically, after determining the target Mel-spectrogram, target deep semantic features, target noise data, time step, target special action tag, and target action data, the target Mel-spectrogram, target deep semantic features, target noise data, time step, and target special action tag are input into the diffusion model. The diffusion model denoises the target noise data to obtain predicted action data. The loss between the predicted action data and the target action data is then calculated. Based on the calculated loss function, the parameters of the diffusion model are updated using an optimization algorithm, such as stochastic gradient descent (SGD) or SGD variants such as Adagrad and Adadelta. After the update, the target training sample is reacquired and the above steps are repeated, continuously adding noise to the target action data. The updated diffusion model is then used to denoise the predictions, calculate the loss, and optimize the model parameters until the loss function converges, i.e., the gap between the predictions of the updated diffusion model and the original target action data no longer decreases significantly. At this point, training is completed, and the updated diffusion model is determined as the pose generation model.

[0086] The loss function may be mean squared error (MSE).

[0087] For example, the process of constructing a gesture generation model can be as follows: Figure 5 As shown, first, the motion data of the motion capture actor recorded according to the TTS voice is obtained, wherein the time of the TTS voice and the motion data is aligned; then, the TTS voice is sliced ​​by a fixed length to obtain multiple training samples, wherein the training sample includes voice data of length L, motion data time-aligned with the voice data of length L, and the style of the voice data (special motion marker); then, a training sample (i.e., the target training sample) is randomly sampled from the multiple training samples, and then the target motion data is subjected to noise processing at the first time step t1 to obtain the first noise data x t1 , and perform noise processing on the target motion data at the second time step t2 to obtain the second noise data x t2 , and simultaneously process the target speech data to obtain the target Mel spectrum and target deep semantic features.

[0088] The time step is greater than the first time step, the first time step is greater than the second time step, and the difference between the first time step and the second time step is a preset time step. For example, the time step may be 50, 100, or 500, the first time step may be 30 or 20, and the second time step may be 5 or 10.

[0089] After getting the first noise data x t1 , the second noise data x t2 , target Mel spectrum and target deep semantic features, the first noise data x t1, target mel-spectrogram and target deep semantic feature input diffusion model, the diffusion model performs denoising processing on the first noise data x t1 for a preset time step to output predicted noise data, then calculates a loss between the predicted noise data and the second noise data x t2 , and then updates parameters of the diffusion model using an optimization algorithm according to the calculated loss. After the update, the target training sample is reacquired, and the above steps are repeated until the loss converges.

[0090] Step S402, acquiring first speech data to be output, a first special action label, Gaussian noise and a time step.

[0091] For details, please refer to step S101 of the embodiment shown in Figure 1 , which will not be repeated here.

[0092] Step S403, determining a first mel-spectrogram and a first deep semantic feature according to the first speech data.

[0093] For details, please refer to step S102 of the embodiment shown in Figure 1 , which will not be repeated here.

[0094] Step S404, inputting the first mel-spectrogram, the first deep semantic feature, the first special action label, the Gaussian noise and the time step into a pose generation model, and determining first action data according to an output of the pose generation model.

[0095] For details, please refer to step S103 of the embodiment shown in Figure 1 , which will not be repeated here.

[0096] Step S405, synchronously controlling a humanoid robot to generate a gesture action represented by the first action data when the first speech data is output.

[0097] For details, please refer to step S104 of the embodiment shown in Figure 1 , which will not be repeated here.

[0098] In this embodiment, before inputting the first mel-spectrogram, the first deep semantic feature, the first special action label, the Gaussian noise and the time step into the pose generation model, the pose generation model is constructed through a sample database, and the sample database contains diverse training samples, so that the diffusion model can capture the mapping relationship between the pose and the input of the mel-spectrogram, the semantic feature and the special action label, thereby improving the accuracy of the generated pose.

[0099] For example, Figure 6As shown in FIG, the diffusion model may include an input layer, a hidden layer, and an output layer. The hidden layer adopts a Transformer network architecture and includes layer normalization, a multilayer perceptron (MLP), a self-attention layer, and a stylization layer.

[0100] Specifically, the Mel spectrum explicitly interacts with the time step through a Transformer encoder, allowing it to focus on conditional features at different time steps (i.e., different noise levels) (for example, focusing on global rhythm in the early high-noise stage and focusing on detailed synchronization in the later low-noise stage), resulting in a matrix of size L×256. The HuBERT feature passes through a linear projection layer to obtain a matrix of size L×128. The two are concatenated in the second dimension to obtain a matrix of size L×384, which is input to the input layer of the diffusion model. Among them, L represents the preset time length mentioned above. The linear projection layer is a common method for deep neural networks to process input features. The purpose is to map them into a unified latent space while maintaining their continuity.

[0101] The noise (target noise data, Gaussian noise, or first noise data) can be an L×70 matrix. After a linear projection layer, it is converted into an L×512 matrix, which is then fed into the input layer of the diffusion model. 70 represents the number of free joints of the humanoid robot. The time step is a positive integer corresponding to the number of noise steps added. After a linear projection layer, it is converted into a vector of length 512, which is then fed into the input layer of the diffusion model. Special action markers are added to the time step features to guide the model to perform specific actions or control the style of the actions generated by the diffusion model.

[0102] The Transformer network architecture consists of N Transformer blocks. In one iteration, the audio and time step inputs are fixed (i.e., the two inputs are the same for each block). The noise component is calculated through a series of layers within the block and becomes the output of the block. It also serves as the input of the next block. After passing through N blocks, it becomes the final output of the diffusion model.

[0103] Figure 6 in This represents concatenation, specifically concatenating audio features and noise. Both have a time axis and are temporally aligned, so they can be concatenated in the feature dimension while preserving temporal consistency, allowing the hidden layer to uniformly process both components. Layer normalization is a technique used in neural networks to accelerate training and improve model stability. Its core idea is to standardize the intermediate results of a sample.

[0104] Figure 6 in The word "sum" represents addition, which is used to implement residual connections. By adding the intermediate results of several layers ago to the current intermediate result, the model can learn the difference between the two, which makes it easier for the model to converge than directly learning a single function. The self-attention layer enables the model to dynamically focus on information at different positions in the input sequence, thereby better capturing long-range dependencies and contextual semantics. In the present invention, this layer enables the model to learn coherent actions in the clip.

[0105] In the diffusion model of generative action, the stylization layer often considers "corresponding to audio" as a form of stylization. This layer achieves stylization by projecting the conditional input (i.e., audio, time step, special action marker) into the parameters of an affine transformation, and then applying this transformation to the noise input, adjusting the distribution of the noise (because different styles have different feature distributions). The addition, layer normalization, MLP, and stylization layers appear multiple times to increase the depth of the network, allowing it to learn more content. The output layer is the output of the Transformer block, which is also the output of the entire diffusion model.

[0106] In some embodiments, before extracting target training samples from the sample database, the method for generating postures of a humanoid robot further includes steps a1 and a2:

[0107] Step a1: Acquire at least one third voice data and third action data corresponding to the third voice data.

[0108] The duration of the third voice data is greater than the duration of the second voice data.

[0109] Specifically, the third voice data may be a TTS voice, which may be determined from the historical TTS voice of the humanoid robot, and the third motion data may be motion data recorded by a motion capture actor based on the TTS voice. The third motion data may include multiple motions.

[0110] Step a2: Slice at least one third speech data to obtain multiple training samples.

[0111] Specifically, each third voice data may be sliced ​​by sliding a data window of a preset duration to divide each third voice data into at least one second voice data of a preset duration, thereby obtaining a plurality of training samples.

[0112] In this embodiment, the data window of the preset time length cuts the third voice data into segments of fixed time length, which can improve the stability of the motion data generated by the diffusion model.

[0113] The present invention targets two human-machine interaction scenarios of a humanoid robot, namely, idle and dialogue. The training processes for the idle and dialogue scenarios are slightly different, which will be described in detail below with reference to the accompanying drawings.

[0114] Specifically, in idle scenarios, the humanoid robot is required to stand naturally with slight swaying, interspersed with special actions such as waving, giving a thumbs-up, moving fingers, and looking down at hands. This invention introduces special action numbers (i.e., special action tags) as additional conditional information in the diffusion model to control the diffusion model's generation of corresponding preset actions.

[0115] In an idle scenario, in order to make the posture generation model more likely to perform special actions and ensure the continuity of the actions, before extracting the target training sample from the sample database, the posture generation method of the humanoid robot also includes: setting the weight of the first training sample to a first preset value, setting the weight of the second training sample to a second preset value, and setting the weight of the third training sample to a third preset value.

[0116] Among them, the first training sample is a training sample corresponding to the second motion data including the start stage of the action, the second training sample is a training sample corresponding to the second motion data including the maintenance stage of the action, and the third training sample is a training sample corresponding to the second motion data including the end stage of the action.

[0117] The first preset value is greater than the second preset value, and the second preset value is greater than the third preset value. For example, the first preset value can be any value within 10 to 100, for example, the first preset value can be 10, 20, 40, 50, 60, 80 or 100, etc. The second preset value can be greater than or equal to 1, and the second preset value is less than 10, for example, the second preset value can be 1, 2, 5, 6 or 9, etc. The third preset value can be greater than or equal to 0, and the third preset value is less than 1, for example, the third preset value can be 0, 0.1, 0.5, 0.7 or 0.9, etc.

[0118] Specifically, the schematic diagram of the third action data can be as follows: Figure 7 As shown, Figure 7 The horizontal axis is the time axis, and the vertical axis is the arm height. The arm height can refer to the rotation angle of a joint of the humanoid robot arm. The larger the value, the higher the arm is raised ( Figure 7 This is just an example, not the actual value). The straight segment indicates that the arm has basically no movement (i.e., standing naturally), and the protruding segment indicates that the arm is raised (i.e., a special action). The action label corresponds to the height of the arm on the timeline, 0 indicates basically no movement, and 1 indicates the action numbered 1 (the number is the special action mark mentioned above). Figure 7 As shown, the execution of a special action includes an action start phase, an action hold phase, and an action end phase.

[0119] During training, in order to make the diffusion model have a greater probability to make the corresponding action, the sampling probability of the training sample in the action start stage is increased by setting the first preset value. At the same time, since the action corresponding to the action end stage is usually the process of lowering the arm (corresponding to the decrease of the arm height value), the data in this part will interfere with the action distribution during model training, so that the model is easy to generate the action of lowering the hand immediately after lifting the hand when generating continuous and coherent actions. Therefore, the data in this part is reduced (or even discarded) by setting the third preset value.

[0120] For example, if the first preset value is 10-100 and the second preset value is 1, the sampling probability of the training sample in the action start stage can be 10-100 times the sampling probability of the training sample in the action maintenance stage.

[0121] Specifically, in the dialogue scene, the humanoid robot is required to make actions matching the rhythm and content of the output voice while outputting the voice, and to maintain natural standing actions in the non-speaking state. In order to ensure that the same model can process different actions in dialogue and idle states, and to optimize the synchronization of speaking actions and voice rhythm, in some embodiments, before training the diffusion model, the posture generation method of the humanoid robot further includes steps b1 and b2:

[0122] Step b1, determining whether the target mel spectrum includes 0.

[0123] Specifically, after obtaining the target mel spectrum, each element in the mel spectrum matrix can be checked to determine whether there is a value equal to 0.

[0124] Step b2, when the target mel spectrum includes 0, mapping the region of the target mel spectrum with 0 and the part of the deep semantic feature corresponding to the region of the target mel spectrum with 0 to a learnable parameter to obtain an updated target mel spectrum and an updated target deep semantic feature.

[0125] The learnable parameter is a model parameter that can learn good audio representations in the encoding space, and is used to replace the mel spectrum and deep semantic feature of the audio to improve the effect of the posture generation model generating action data. The learnable parameter can capture additional information in the audio and be associated with the corresponding action through the gradient backpropagation of deep learning.

[0126] Specifically, when the target mel spectrum includes 0, a parameterized function (such as a neural network layer) can be used to map the region of the target mel spectrum with 0 to a new feature value (i.e. a learnable parameter) to obtain an updated target mel spectrum, and map the part of the deep semantic feature corresponding to the region of the target mel spectrum with 0 to a learnable parameter to obtain an updated target deep semantic feature.

[0127] For example, as Figure 8 As shown, the blank areas in the three Mel spectra are represented as 0 areas, and the blank areas in the three Mel spectra are replaced by learnable parameters.

[0128] In one example, if Figure 9 As shown in the figure, after obtaining the target Mel spectrum, we first determine whether the target Mel spectrum includes 0. If it does, the corresponding parts of the target Mel spectrum and deep semantic features are replaced with learnable parameters before being input into the diffusion model. If it does not include 0, the target Mel spectrum and deep semantic features are directly input into the diffusion model.

[0129] At this time, the above step S4014 is specifically: training the diffusion model according to the updated target Mel spectrum, the updated target deep semantic features, the target noise data, the time step, the target special action tag and the target action data.

[0130] Specifically, for frames with a Mel spectrum of 0 (including empty audio in idle scenes and pauses in TTS speech), the audio encoding is mapped to learnable parameters, and the best empty audio representation is learned in the encoding space, enhancing the model's ability to distinguish between speaking and idle states.

[0131] In this embodiment, the process of replacing the area with a Mel-spectrum spectrum of 0 with a learnable parameter is recorded as a frame-by-frame mode. In the frame-by-frame mode, the diffusion model can completely associate the empty audio with the non-speaking state action. In this way, even if only a part of a segment of length L is empty audio, the diffusion model can generate a corresponding non-speaking action.

[0132] In some other embodiments, before training the diffusion model, the method for generating a posture of a humanoid robot further includes steps c1 and c2:

[0133] Step c1: Determine whether the target Mel spectrum is all zero.

[0134] Specifically, after obtaining the target mel spectrum, each element in the mel spectrum matrix may be checked to determine whether all elements are 0.

[0135] Step c2: When the target Mel spectrum is all zero, the target Mel spectrum and the target deep semantic features are mapped into learnable parameters to obtain first model parameters and second model parameters.

[0136] Specifically, when the target Mel spectrum is all 0, a parameterized function (such as a neural network layer) can be used to map the target Mel spectrum to a new feature value (i.e., a learnable parameter) to obtain the first model parameter, and the deep semantic feature can be mapped to a learnable parameter to obtain the second model parameter.

[0137] For example, Figure 10As shown, the blank areas in the three Mel-spectra represent areas of 0. At this time, only the second Mel-spectra is entirely 0. In this embodiment, only the second Mel-spectra needs to be mapped to a learnable parameter.

[0138] Among them, when the target Mel spectrum is not all 0, the target Mel spectrum and deep semantic features are directly input into the diffusion model.

[0139] At this time, the above step S4014 is specifically: training the diffusion model according to the first model parameters, the second model parameters, the target noise data, the time step, the target special action label and the target action data.

[0140] In this embodiment, the process of replacing all 0 Mel spectra with learnable parameters is recorded as a sample mode. In the sample mode, the diffusion model can learn more coherent non-speech state actions.

[0141] This embodiment also provides a posture generation device for a humanoid robot, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0142] This embodiment provides a posture generating device for a humanoid robot, such as Figure 11 Shown, including:

[0143] An acquisition module 1101 is configured to acquire first speech data to be output, a first special action marker, Gaussian noise, and a time step, wherein the first special action marker is used to characterize the style of a generated gesture or a preset special action;

[0144] A feature determination module 1102 is configured to determine a first mel-spectrogram and a first deep semantic feature based on the first speech data, wherein the first deep semantic feature is the first speech data processed by the self-supervised speech representation learning model;

[0145] an action determination module 1103 , configured to input the first mel spectrum, the first deep semantic feature, the first special action tag, the Gaussian noise, and the time step into a gesture generation model, and determine first action data according to an output of the gesture generation model;

[0146] The output module 1104 is configured to synchronously control the humanoid robot to generate a gesture action represented by the first action data when outputting the first voice data.

[0147] In some optional embodiments, the device further comprises:

[0148] A construction module is used to construct a gesture generation model based on a sample database, wherein the sample database includes multiple training samples, the training samples include second voice data, second action data that changes with the second voice data, and a second special action mark, and the duration of the second voice data is the same as the duration of the first voice data.

[0149] In some optional embodiments, the building blocks include:

[0150] An extraction unit is configured to extract a target training sample from a sample database, wherein the target training sample is one of a plurality of training samples, and the target training sample includes target speech data, target action data that changes with the target speech data, and a target special action tag;

[0151] A first determining unit is used to determine a target Mel spectrum and a target deep semantic feature according to the target speech data;

[0152] A denoising unit, configured to perform denoising processing on the target motion data according to the time step to obtain target noise data;

[0153] A training unit is used to train the diffusion model according to the target Mel spectrum, the target deep semantic features, the target noise data, the time step, the target special action label and the target action data to obtain the posture generation model.

[0154] In some optional embodiments, the device further comprises:

[0155] A first determination module is used to determine whether the target Mel spectrum includes 0;

[0156] A first mapping module is configured to, when the target Mel spectrum includes 0, map the region where the target Mel spectrum is 0 and the portion of the deep semantic feature corresponding to the region where the target Mel spectrum is 0 into a learnable parameter, thereby obtaining an updated target Mel spectrum and an updated target deep semantic feature;

[0157] The training unit includes: a first sub-training unit, for training the diffusion model according to the updated target Mel spectrum, the updated target deep semantic features, the target noise data, the time step, the target special action tag and the target action data.

[0158] In some optional embodiments, the device further comprises:

[0159] The second determination module is used to determine whether the target Mel spectrum is all 0;

[0160] A second mapping module is used to map the target Mel spectrum and the target deep semantic features into learnable parameters when the target Mel spectrum is all 0, so as to obtain the first model parameters and the second model parameters;

[0161] The training unit includes: a second sub-training unit, which is used to train the diffusion model according to the first model parameter, the second model parameter, the target noise data, the time step, the target special action label and the target action data.

[0162] In some optional embodiments, the device further comprises:

[0163] a data acquisition module, configured to acquire at least one third voice data and third action data corresponding to the third voice data, wherein the duration of the third voice data is greater than the duration of the second voice data;

[0164] The slicing module is used to perform slicing processing on at least one third voice data to obtain multiple training samples.

[0165] In some optional implementations, the third action data includes an action start phase, an action hold phase, and an action end phase, and the apparatus further includes:

[0166] A processing module is used to set the weight of a first training sample to a first preset value, set the weight of a second training sample to a second preset value, and set the weight of a third training sample to a third preset value, wherein the first training sample is a training sample corresponding to the second motion data including the start stage of an action, the second training sample is a training sample corresponding to the second motion data including the maintenance stage of an action, and the third training sample is a training sample corresponding to the second motion data including the end stage of an action, the first preset value is greater than the second preset value, and the second preset value is greater than the third preset value.

[0167] In some optional implementations, the first voice data is empty audio, and the first special action mark is determined according to a preset frequency of the special action.

[0168] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0169] The posture generation device of the humanoid robot in this embodiment is presented in the form of a functional unit, where the unit refers to an application-specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0170] The embodiment of the present invention also provides a humanoid robot, such as Figure 12As shown, the humanoid robot includes: one or more processors 1210, a memory 1220, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the humanoid robot, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple humanoid robots can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 12 A processor 1210 is taken as an example.

[0171] Processor 1210 may be a central processing unit (CPU), a network processor (NPU), or a combination thereof. Processor 1210 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a general purpose array logic (GAL), or any combination thereof.

[0172] The memory 1220 stores instructions that can be executed by at least one processor 1210, so as to enable at least one processor 1210 to execute the method shown in the above embodiment.

[0173] The memory 1220 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the humanoid robot, etc. Furthermore, the memory 1220 may include high-speed random access memory and may also include non-transient memory, such as at least one disk storage device, flash memory device, or other non-transient solid-state memory device. In some optional embodiments, the memory 1220 may optionally include a memory remotely located relative to the processor 10, and such remote memory may be connected to the humanoid robot via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0174] The memory 1220 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 1220 may also include a combination of the above types of memory.

[0175] The humanoid robot further includes an input device 1230 and an output device 1240. The processor 1210, the memory 1220, the input device 1230 and the output device 1240 may be connected via a bus or other means. Figure 12 The bus connection is taken as an example.

[0176] The input device 1230 can receive input digital or character information and generate key signal input related to user settings and function control of the humanoid robot, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, a pointer, one or more mouse buttons, a trackball, a joystick, etc. The output device 1240 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0177] The embodiments of the present invention also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded on a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded via a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a humanoid robot, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that the humanoid robot, processor, microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the humanoid robot, processor or hardware, the method shown in the above embodiment is implemented.

[0178] A portion of the present invention may be applied as a computer program product, such as computer program instructions, which, when executed by a humanoid robot, can call or provide the method and / or technical solution according to the present invention through the operation of the humanoid robot. Those skilled in the art will appreciate that the computer program instructions may exist in a computer-readable medium in forms including, but not limited to, source files, executable files, installation package files, and the like. Accordingly, the manner in which the computer program instructions are executed by a humanoid robot includes, but is not limited to: the humanoid robot directly executing the instruction, or the humanoid robot compiling the instruction and then executing the corresponding compiled program, or the humanoid robot reading and executing the instruction, or the humanoid robot reading and installing the instruction and then executing the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the humanoid robot.

[0179] In the description of this specification, the reference terms "this embodiment", "one embodiment", "some embodiments", "example", "specific example" or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0180] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0181] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations shall all fall within the scope defined by the present invention.

Claims

1. A method for generating a posture of a humanoid robot, characterized in that: The method comprises: Obtaining first voice data to be output, a first special action marker, Gaussian noise, and a time step, wherein the first special action marker is used to characterize the style of the generated gesture action or the preset special action; Determining a first mel spectrum and a first deep semantic feature based on the first speech data, wherein the first deep semantic feature is the first speech data processed by a self-supervised speech representation learning model; inputting the first Mel spectrum, the first deep semantic feature, the first special action marker, the Gaussian noise, and the time step into a posture generation model, and determining first action data according to an output of the posture generation model, wherein the posture generation model is a trained diffusion model; When outputting the first voice data, the humanoid robot is synchronously controlled to generate a gesture action represented by the first action data.

2. The method according to claim 1, characterized in that Before inputting the first mel-spectrogram, the first deep semantic feature, the first special action marker, the Gaussian noise, and the time step into a gesture generation model, the method further includes: The gesture generation model is constructed based on a sample database, wherein the sample database includes multiple training samples, and the training samples include second voice data, second action data that changes with the second voice data, and a second special action mark, and the duration of the second voice data is the same as the duration of the first voice data.

3. The method according to claim 2, characterized in that The step of constructing the posture generation model based on the sample database includes: Extracting a target training sample from the sample database, wherein the target training sample is one of the multiple training samples, and the target training sample includes target speech data, target action data that changes with the target speech data, and a target special action tag; Determining target Mel spectrum and target deep semantic features according to the target speech data; performing noise processing on the target motion data according to the time step to obtain target noise data; The diffusion model is trained according to the target Mel spectrum, the target deep semantic features, the target noise data, the time step, the target special action tag and the target action data to obtain a posture generation model.

4. The method according to claim 3, characterized in that Before training the diffusion model, the method further includes: Determine whether the target Mel spectrum includes 0; When the target Mel spectrum includes 0, mapping the area where the target Mel spectrum is 0 and the portion of the deep semantic feature corresponding to the area where the target Mel spectrum is 0 into learnable parameters to obtain an updated target Mel spectrum and an updated target deep semantic feature; The step of training the diffusion model according to the target Mel spectrum, the target deep semantic features, the target noise data, the time step, the target special action tag, and the target action data includes: The diffusion model is trained according to the updated target Mel spectrum, the updated target deep semantic features, the target noise data, the time step, the target special action tag and the target action data.

5. The method according to claim 3, characterized in that Before training the diffusion model, the method further includes: Determine whether the target Mel spectrum is all 0; When the target Mel spectrum is all 0, mapping the target Mel spectrum and the target deep semantic feature into learnable parameters to obtain first model parameters and second model parameters; The step of training the diffusion model according to the target Mel spectrum, the target deep semantic features, the target noise data, the time step, the target special action tag, and the target action data includes: The diffusion model is trained based on the first model parameters, the second model parameters, the target noise data, the time step, the target special action label and the target action data.

6. The method according to any one of claims 3 to 5, characterized in that Before extracting target training samples from the sample database, the method further includes: Acquire at least one third voice data and third action data corresponding to the third voice data, wherein the duration of the third voice data is greater than the duration of the second voice data; Slice the at least one third voice data to obtain the multiple training samples.

7. The method according to claim 6, characterized in that The third action data includes an action start phase, an action hold phase, and an action end phase. Before extracting a target training sample from the sample database, the method further includes: The weight of the first training sample is set to a first preset value, the weight of the second training sample is set to a second preset value, and the weight of the third training sample is set to a third preset value, wherein the first training sample is a training sample corresponding to the second motion data including the start stage of the action, the second training sample is a training sample corresponding to the second motion data including the maintenance stage of the action, and the third training sample is a training sample corresponding to the second motion data including the end stage of the action, the first preset value is greater than the second preset value, and the second preset value is greater than the third preset value.

8. The method according to any one of claims 1 to 5, characterized in that The first voice data is empty audio, and the first special action mark is determined according to a preset frequency of the special action.

9. A posture generating device for a humanoid robot, characterized in that: The device comprises: an acquisition module, configured to acquire first voice data to be output, a first special action marker, Gaussian noise, and a time step, wherein the first special action marker is used to characterize the style of the generated gesture action or the preset special action; a feature determination module, configured to determine a first mel-spectrogram and a first deep semantic feature based on the first speech data, wherein the first deep semantic feature is the first speech data processed by a self-supervised speech representation learning model; an action determination module, configured to input the first mel spectrum, the first deep semantic feature, the first special action marker, the Gaussian noise, and the time step into a posture generation model, and determine first action data according to an output of the posture generation model; The output module is used to synchronously control the humanoid robot to generate a gesture action represented by the first action data when outputting the first voice data.

10. A humanoid robot, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the gesture generation method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a humanoid robot to execute the gesture generation method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions for causing a humanoid robot to execute the gesture generation method according to any one of claims 1 to 8.