Gesture generation system and gesture generation method
The gesture generation system addresses the challenge of generating synchronized and semantically relevant gestures by using a de-diffusion process with inter-patch and intra-patch attention, enhancing the naturalness and accuracy of avatar and robot gestures.
Patent Information
- Application Number
- JP2024095522
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-12-25
AI Technical Summary
Existing technologies struggle to generate gestures for avatars and robots that accurately synchronize with speech, exhibit naturalness, and maintain semantic relevance, requiring expertise and being resource-intensive.
A gesture generation system utilizing a de-diffusion process with a pre-trained machine learning algorithm that employs inter-patch and intra-patch attention mechanisms to generate gesture data from audio signals, incorporating a diffusion model trained on large-scale datasets to produce synchronized and semantically relevant gestures.
The system generates gestures that more naturally and accurately synchronize with speech, reducing the need for expert intervention and resource intensity.
Smart Images

Figure 2025187051000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a gesture generation technique, and more particularly to a technique for generating gesture data that defines gestures accompanying speech, such as those of robots and so-called avatars that have a human-like appearance. [Background technology]
[0002] Avatars with human-like appearance are widely used in film and game production. Like humans, avatars can interact with listeners, such as interlocutors or audiences, not only through speech but also through non-speech activities. The most important of these activities is gesture. A gesture can be described as a postural movement of the avatar, i.e., a sequence of positional information of each part of the avatar's body.
[0003] Gestures can supplement or complement spoken information to better convey the thoughts or feelings of the speaker behind the avatar to the listener. The listener can use the avatar's gestures to better understand what the speaker is saying compared to when only the speech is provided. Therefore, it is very important for the avatar to perform appropriate gestures in conjunction with its speech.
[0004] However, although the importance of gestures is known, it is not easy to actually make an avatar perform appropriate gestures. Appropriate gesture data is essential for making an avatar perform appropriate gestures. Generating gesture data requires an expert with extensive knowledge and skills to design appropriate gestures to accompany speech. Even if such knowledge and skills are available, generating gesture data takes a long time. This is because gesture data is generally represented as multidimensional information consisting of time-series information specifying the rotation angles of each of multiple joints of the skeleton associated with the avatar, or time-series information specifying three-dimensional position coordinates.
[0005] Another method for generating gesture data is motion capture, but motion capture is known to be expensive and must be re-run for each new scenario in which the gestures are to be applied.
[0006] An economical method is to use a simple model of human skeletal data. The model's posture is changed and photographed using a normal camera or a distance-measuring camera. This processing allows the identification of 2D and 3D information for each human joint. However, the gesture data obtained using this method contains a lot of noise. Therefore, this gesture data cannot be used as is.
[0007] To solve these problems, one approach is to generate gesture data based on speech audio signals. Recently, large-scale datasets of speech and gesture data have been proposed. Using such data, researchers can use deep learning to train a model that encodes the relationship between the speech audio signal, which is a time-series signal, and the corresponding gesture data. Such a model receives input segments of the audio signal and generates and outputs gesture data that define gesture segments. Such a model can generate multiple gesture data for the same input audio signal. Users of this model can select their preferred gesture data from among these (Non-Patent Document 1). Using such a model can significantly reduce the workload of generating gesture data. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] “Controlling the impression of robots via GAN-based gesture generation,” in 2022 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 20, 2022 . 2022, pp. 101-1 9288-9
Outdoor Tool2
Outdoor Tools3
Outdoor Tools 4
[0009] As described above, technologies related to models that generate gestures based on speech signals have made remarkable progress. However, there are still areas for improvement between the gestures of avatars generated by these technologies and those of actual humans in terms of naturalness, accuracy of synchronization with speech, and semantic relevance to speech. These problems are not limited to avatars, but can also occur in robots. Furthermore, the same problems can occur not only in avatars or robots that have nearly human-like appearances, but also in avatars or robots that have the appearance of other things, including anthropomorphic ones.
[0010] Therefore, an object of the present invention is to provide a gesture generation system and a gesture generation method that can generate gestures from voice signals more naturally than conventional techniques. [Means for solving the problem]
[0011] A gesture generation system according to a first aspect of the present invention is a gesture data generation device including a memory for storing a group of instructions and a processor connected to the memory for executing the group of instructions to perform predetermined processing, the processing including a gesture data generation process for generating gesture data adapted to audio data from noise conditioned by the audio data by a de-diffusion process using a pre-trained machine learning algorithm, the noise including a plurality of frames conditioned by the audio data, the plurality of frames being classified into a plurality of patches, and the machine learning algorithm generating gesture data by a de-diffusion process using an inter-patch attention mechanism that uses inter-patch attention, which is attention between the plurality of patches, and an intra-patch attention mechanism that uses intra-patch attention, which is attention between the plurality of frames included in each of the plurality of patches.
[0012] Preferably, each of the plurality of patches includes two or more frames that are consecutive to one another.
[0013] More preferably, each of the multiple patches includes an equal number of frames.
[0014] More preferably, the machine learning algorithm runs the inter-patch attention mechanism and the intra-patch attention mechanism sequentially on multiple frames.
[0015] Preferably, the machine learning algorithm performs an inter-patch attention mechanism before an intra-patch attention mechanism over multiple frames.
[0016] More preferably, each of the plurality of frames includes a numerical value for controlling the three-dimensional position of each joint of a predetermined skeleton, the de-diffusion process executes a noise removal process on the plurality of frames to generate a plurality of output frames, and further inputs the plurality of output frames into the noise removal process to perform the noise removal process, repeating this process until a predetermined termination condition is met, and the de-diffusion process further includes a restriction process for restricting the numerical values included in the plurality of output frames obtained in each of the noise removal processes of the de-diffusion process within a predetermined range.
[0017] More preferably, the memory includes an area for storing a restriction value for defining the range of motion of the skeleton, and the restriction process includes a process for restricting, in each repetition of the de-diffusion process, the numerical values contained in the multiple output frames obtained by the de-diffusion process to within a range defined by the restriction value.
[0018] Preferably, the memory includes an area for storing a restriction value for defining the range of motion of the skeleton, and the restriction process includes a process for restricting, in each repetition of the de-diffusion process, the numerical values contained in the multiple output frames obtained by the de-diffusion process to fall outside the range defined by the restriction value.
[0019] More preferably, each of the plurality of frames includes a vector whose elements are the rotation angles of each joint of a predetermined skeleton.
[0020] More preferably, each of the plurality of frames includes a vector having elements each representing a three-dimensional coordinate of each joint of a predetermined skeleton.
[0021] Preferably, the predetermined skeleton comprises a human skeleton.
[0022] More preferably, the predetermined skeleton includes at least the upper body skeleton of a human.
[0023] A gesture generation system according to a second aspect of the present invention is a gesture data generation system including: a memory for storing at least a group of instructions; and a processor connected to the memory for executing the group of instructions to perform predetermined processing, wherein the processing is a gesture generation process for generating gesture data adapted to voice data from noise conditioned by voice data by a de-diffusion process using a machine learning algorithm that has been trained in advance by a diffusion process, wherein the gesture data includes a time series of numerical values for controlling the three-dimensional position of each joint of the skeleton, and the memory stores restriction values for defining the range of each of the numerical values generated by the gesture generation process in each repetition of the de-diffusion process.
[0024] A gesture generation method according to a third aspect of the present invention is a gesture generation method that generates gesture data that matches voice data from noise conditioned by voice data by a de-diffusion process using a machine learning algorithm that has previously been trained by a diffusion process, wherein the noise includes a plurality of frames conditioned by the voice data, and the plurality of frames are classified into a plurality of patches, and the machine learning algorithm includes a step in which a computer modifies the plurality of frames using inter-patch attention, which is attention between the plurality of patches, and a step in which a computer modifies the plurality of frames in each of the plurality of patches using intra-patch attention, which is attention between the plurality of frames included in the patch.
[0025] A gesture generation system according to a fourth aspect of the present invention includes: initial noise sampling means for sampling initial noise from a predetermined distribution; and a gesture generation unit including a machine learning model that has been trained in advance by a diffusion process so as to estimate the noise based on input data to which the noise has been added, wherein the input data includes a plurality of frames, and the plurality of frames are classified into a plurality of patches; the gesture generation unit generates gesture data by a de-diffusion process consisting of a plurality of steps for the diffusion process; a selection unit that selects either an input to the gesture generation unit or an input from the machine learning model; a conditioning unit that conditions an output of the selection unit using a sigma level that specifies a noise schedule determined in association with the plurality of steps, and input voice data; The system includes a sampling means that estimates a new distribution of gesture data in response to the machine learning model outputting gesture data candidates in accordance with input conditioned by a conditioning unit, and samples a new sample from the estimated distribution; and a branching unit that selects either a process of outputting the new sample as gesture data or a process of providing the new sample to a selection unit as input from the machine learning model, depending on whether multiple steps have ended.The selection unit selects an input to the gesture generation unit or an output from the branching unit, depending on whether multiple steps have started.The machine learning model includes an inter-patch attention mechanism that corrects multiple frames using inter-patch attention, which is attention between multiple patches, and a mechanism that corrects multiple frames using intra-patch attention, which is attention within multiple patches.
[0026] The above and other objects, features, aspects and advantages of the present invention will become apparent from the following detailed description of the invention taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0027] [Figure 1] FIG. 1 is a diagram showing an example of an avatar making a gesture in response to speech. [Figure 2]FIG. 2 is a diagram for explaining the diffusion process and the de-diffusion process of the diffusion model. [Figure 3] FIG. 3 is a functional block diagram of a system for training a diffusion model according to one embodiment of the present invention. [Figure 4] FIG. 4 is a functional block diagram of a gesture generation system that generates gesture data using the diffusion model shown in FIG. [Figure 5] FIG. 5 is a functional block diagram of the diffusion model shown in FIGS. [Figure 6] FIG. 6 is a functional block diagram of FiLM (Feature-wise Linear Modulation) shown in FIG. [Figure 7] FIG. 7 is a functional block diagram of the audio signal subsampling unit and point-wise FiLM shown in FIG. [Figure 8] FIG. 8 is a functional block diagram of the first decoder block shown in FIG. [Figure 9] FIG. 9 is a functional block diagram of the inter-patch attention mechanism shown in FIG. [Figure 10] FIG. 10 is a functional block diagram of the depth-wise convolutional layer shown in FIG. 9. [Figure 11] FIG. 11 is a functional block diagram of the transposed convolution layer shown in FIG. [Figure 12] FIG. 12 is a block diagram showing the functional configuration of the intra-patch attention mechanism shown in FIG. [Figure 13] FIG. 13 is a flowchart showing a control structure of a computer program that causes a computer to function as the training system shown in FIG. [Figure 14] FIG. 14 is a flowchart showing a control structure of a computer program causing a computer to function as the gesture generation system shown in FIG. [Figure 15] FIG. 15 is a graph for explaining the effect of the diffusion model according to the first embodiment of the present invention. [Figure 16]FIG. 16 is a schematic diagram showing skeletal movements to explain points that should be improved in gesture data. [Figure 17] FIG. 17 is a schematic diagram showing skeletal movements to explain what should be improved in gesture data. [Figure 18] FIG. 18 is a schematic diagram showing the movement of a skeleton to explain one method of correcting gesture data. [Figure 19] FIG. 19 is a schematic diagram for explaining the problem with gesture data obtained by conventional techniques. [Figure 20] FIG. 20 is a block diagram showing the functional configuration of a diffusion model according to the second embodiment of the present invention. [Figure 21] FIG. 21 is a flowchart showing a control structure of a program for realizing the diffusion model according to the second embodiment of the present invention shown in FIG. [Figure 22] FIG. 22 is a schematic diagram showing an outline of a gesture generated by a diffusion model according to the second embodiment of the present invention. [Figure 23] FIG. 23 is a diagram showing, in a table format, the evaluation results of gestures obtained by the diffusion model (FIG. 20) according to the second embodiment of the present invention. [Figure 24] 24 is a block diagram showing the functional configuration of a diffusion model according to the third embodiment of the present invention. [Figure 25] FIG. 25 is a flowchart showing a control structure of a program realizing the diffusion model according to the third embodiment of the present invention. [Figure 26] FIG. 26 is an external view of a computer system that executes a program having a control structure shown in FIG. 13, 14, 21, or 25. In FIG. [Figure 27] FIG. 27 is a block diagram showing the hardware configuration of the computer system shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0028] In the following description and drawings, the same parts are designated by the same reference numerals, and therefore detailed description thereof will not be repeated.
[0029] First embodiment 1.1 Configuration 1.1.1. About Gestures 1 shows a mechanical avatar robot 50, an avatar 52 on an image, and a robot-shaped robot avatar 54 on an image, which are used by a user at a remote location to communicate with a listener. With reference to FIG. 1, appropriate gestures are generated based on the user's speech and the avatar robot 50, the avatar 52, and the robot avatar 54 are operated, thereby making it easier for the listener to understand what the user is saying. The following embodiment relates to a gesture generation system for this purpose.
[0030] Terminology Gestures In this specification, "gesture" refers to movements of the upper body (so-called body movements, hand movements, and postures) of human-like avatar images, robots, etc. Gestures play an important role in supplementing linguistic information in communication with humans.
[0031] Gesture Data Gesture data refers to a collection of data that defines the gestures of an avatar. A gesture is defined by the movement of each joint of the avatar. Therefore, the gesture data is a sequence of information that defines the position of each joint as a function of time. In this embodiment, the gesture data includes data indicating the state of each joint at each step, which is defined in advance at a predetermined interval. This data indicating the state of each joint at a certain step is called one frame. In other words, the gesture data includes a sequence of multiple frames. In this embodiment, the state of each joint at each step is represented by a numerical value representing the rotation angle of that joint around three axes at that step, or a numerical value representing the coordinate in two-dimensional or three-dimensional space. Therefore, if the number of joints in a skeleton is Np, each frame includes at least 3Np data items. Note that in the following embodiments, for simplicity, generating gesture data is referred to as gesture generation.
[0032] patch A patch refers to a collection of a predetermined number of consecutive frames included in gesture data. For example, if gesture data includes M frames and these frames are classified into P patches, each patch includes M / P frames (where M is divisible by P). In this embodiment, each patch includes an equal number of consecutive frames, but the present invention is not limited to such an embodiment. Gesture data can also be divided into multiple patches so that each patch includes a different number of frames. Furthermore, the frames classified into one patch do not have to be consecutive.
[0033] Transformer A transformer is a basic component of neural networks, which are currently the mainstream of machine learning models, and was proposed in Non-Patent Document 2. The basic configuration of a transformer proposed in Non-Patent Document 2 includes an encoder and a decoder. A transformer is an architecture originally proposed for natural language processing, but is now applied to a variety of fields due to its high performance and learning efficiency. In addition, the encoder and decoder are also independently applied in many fields.
[0034] Diffusion Model In this embodiment, the "diffusion model" is a type of generative model that includes a machine learning algorithm that generates data from noise in a domain that is the target of training. The diffusion model consists only of a decoder made of a neural network.
[0035] 2, the learning of the diffusion model and the generation of data by the diffusion model are performed by processes called a diffusion process 60 and a de-diffusion process 62, respectively. The diffusion process 60 performs steps t0, t1, ..., t T In multiple steps, we sequentially add noise to the samples x0, x1, x2, ... x T The process is to give the diffusion model a value of x and optimize the parameters of the diffusion model to predict the added noise. This process is repeated for the entire training data until the termination condition is met. At each step, the amount of noise (proportion of noise) added to the sample increases. The final sample x T In this embodiment, the ratio of sample data to noise in the data given to the diffusion model is determined as α for each step t=t0, t1, ..., T. t The noise ratio is determined as 1-α t α t and 1-α t is called the noise schedule. The variable t represents the number of steps and also represents the noise level.
[0036] The de-diffusion process 62 starts with the noise at step T and gradually removes the noise contained in the data at each step by reversing the process of the diffusion process 60. The de-diffusion process 62 estimates the distribution of the data represented by the training data, and generates data at step t1 by sampling the final sample from that distribution.
[0037] An outline of these processes in the present invention will be described later with reference to FIGS.
[0038] Sigma Level In this embodiment, the sigma level refers to the number of each step in the diffusion process. If the total number of steps in the diffusion process is T, the sigma level takes on values from 1 to T in order. In the de-diffusion process, the sigma level takes on values from T to 1 in order. In the detailed description of the invention below, the sigma level is represented by the variable t. In the diffusion process, the noise (proportion) added to the sample increases with each step. Therefore, the sigma level t also represents the magnitude of the noise (proportion) added to the sample. Note that hereafter, when the variable t is simply mentioned, it may also represent the noise level or the number of steps in the diffusion process or de-diffusion process.
[0039] 1.1.3. Training System 70 3 shows a block diagram of a training system 70 for a diffusion model 90 according to this embodiment. A large number of training samples, each consisting of a pair of gesture data and corresponding audio data, are used for this training. The training system 70 described below is designed to generate upper-body gesture data for an avatar.
[0040] Referring to FIG. 3, the training system 70 includes a training data storage unit 80 that stores in advance multiple sets of training data for training the diffusion model 90. Each set of training data includes a predetermined length of human speech and information (gesture data) that identifies the movement of each joint of the speaker, obtained from a gesture image of the speaker at that time. In the following embodiment, this gesture data identifies gestures of the speaker's upper body. The lengths of the sets of training data do not need to be the same.
[0041] In this embodiment, the gesture data 92 includes multiple frames acquired at a certain time interval. Each frame includes three-dimensional coordinate data of each joint of the speaker at the corresponding step. If the number of joints is N, the data for each frame includes at least 3N elements. That is, each frame is represented as a vector having at least 3N elements. If the number of frames included in the gesture data 92 is M, the gesture data 92 is represented as a matrix with M rows and 3N columns.
[0042] The training system 70 further includes a learning control unit 82 for controlling each unit of the training system 70 in order to train the diffusion model 90, a learning data reading unit 86 for reading each set of gesture data and voice data from the learning data storage unit 80 and outputting the data separately as gesture data 92 and voice data 94, and a hyperparameter storage unit 84 for storing hyperparameters referenced by the learning control unit 82 in order to set the behavior of the training system 70. During learning of the diffusion model 90, learning data to which noise has been added is provided to the diffusion model 90, and the parameters of the diffusion model 90 are optimized so that the diffusion model 90 can estimate the noise.
[0043] The training system 70 further includes a noise sampling unit 88 that samples noise 96 required for learning from a standard normal distribution, and a noise mixing unit 100 that receives a parameter α t (the symbol " t" is shown immediately above α in the figures and formulas) called a noise schedule from the learning control unit 82 and noise from the noise sampling unit 88, mixes the gesture data 92 with the noise according to the noise schedule, and inputs the result as noise-containing gesture data 102 to the diffusion model 90. In this embodiment, the proportion of gesture data indicated as input to the diffusion model 90 is determined by the parameter α t. t The noise ratio is expressed by the parameter 1-α t These values are called the noise schedule. The values of these parameters are functions of the value of the variable t. During learning, as the value of the variable t increases, the parameter α t The value of gradually decreases from 1, and instead the parameter 1-α t The value of gradually increases from 0. Then, at the end of training, the input to the diffusion model 90 becomes pure noise. However, as will be described later, in this embodiment, during training, the variable t is repeatedly sampled from a uniform distribution [1, 2, ..., T].
[0044] The training system 70 further includes a parameter update unit 98 for updating the parameters of the diffusion model 90 by backpropagation so as to minimize the squared error between the output of the diffusion model 90 during training and the noise added to the sample by the noise sampling unit 88 and the noise mixing unit 100, and a learning control unit 82 for controlling each unit of the training system 70 to perform training of the diffusion model 90. Each unit of the training system 70 is realized by computer hardware and a computer program (hereinafter simply referred to as a "program") running on the computer hardware. Of these programs, the configuration of the program that realizes the learning control unit 82 will be described later with reference to FIG. 13.
[0045] 1.1.4. Gesture Generation System 110 4, the gesture generation system 110 has a function of receiving voice data 112 as input and generating appropriate gesture data 114 for the voice data 112. To this end, the gesture generation system 110 includes: a diffusion model 90 trained by the training system 70 shown in FIG. 3; a hyperparameter storage unit 84 that stores hyperparameters; a noise sampling unit 122 that samples noise 124 from a standard normal distribution as initial noise at the beginning of gesture generation and provides the noise 124 to the diffusion model 90; and a noise sampling unit 126 that samples noise 128 from a standard normal distribution and provides the noise 128 to the diffusion model 90 at each step of the process of generating the gesture data 114 using the diffusion model 90. The gesture generation system 110 further includes a generation control unit 120 that controls each unit of the gesture generation system 110 to generate the gesture data 114 according to the hyperparameters stored in the hyperparameter storage unit 84. Each unit of the gesture generation system 110 is realized by computer hardware and a program executed by the computer hardware.
[0046] The operations of the gesture generation system 110 and the diffusion model 90 when generating the gesture data 114 will be described later. The programs that realize the various parts of the gesture generation system 110 will be described later in detail with reference to FIG.
[0047] 1.1.5. Diffusion Model 90 5 shows an outline of the configuration of the diffusion model 90. Referring to FIG. 5, the diffusion model 90 selectively receives as input 130 (gesture data 102 including noise, which is the output of the noise mixer 100 in FIG. 3, or noise 124 in FIG. 4) or the output of the previous step of the de-diffusion process of the diffusion model 90 received via a loop 158. The diffusion model 90 includes an FiLM 140 for conditioning each element of the matrix included in the input by performing a transformation (modulation) determined by a variable t indicating the number of steps and sigma level specified in the diffusion process during learning, or by a variable t during the de-diffusion process, and normalizing the transformation so that the diffusion model 90 generates a gesture reflecting the value of the variable t, and a residual connection unit 144 for adding a branch 142 from the input of the FiLM 140 to the output of the FiLM 140. FiLM is a technique proposed in Non-Patent Document 4.
[0048] The diffusion model 90 further includes an audio signal subsampling unit 154 that subsamples the audio signal 132 (the training audio data 94 and the generation audio data 112) and converts it into a frame sequence that matches the number of frames in the input 130, and outputs it as an audio signal 156, and a point-wise FiLM 146 that performs an affine transformation on each element of the data output by the residual connection unit 144, as determined by the corresponding element of the audio signal 156, normalizes it, and outputs it as an output 160. Using the point-wise FiLM 146, the parameters of the diffusion model 90 can be trained so as to condition the gesture data by modulating it with the audio data, and generate gestures that match the input audio.
[0049] The diffusion model 90 further includes a transformer-based decoder 148 and a fully connected layer 150 for predicting the noise component contained in the output 160 and outputting it as output data 162.
[0050] As shown in FIG. 5, the decoder 148 includes a first decoder block 170, a second decoder block 172, a third decoder block 174, and a fourth decoder block 176. These blocks all have the same configuration and are stacked so that the output of the preceding decoder block is input to the succeeding decoder block. The input of the first decoder block 170 is provided with the output 160 of the pointwise FiLM 146. The output of the fourth decoder block 176 is provided to the input of the fully connected layer 150. Note that in this embodiment, four decoder blocks are used because the best performance was obtained when four decoder blocks were used in experiments based on the configuration of this embodiment. The number of decoder blocks may be three or less, or five or more.
[0051] FIG. 6 shows the functional configuration of FiLM 140 in FIG. 5. In the following drawings and descriptions, the number of dimensions of each frame of gesture data is represented by C, and the total number of steps in the diffusion process is represented by T. Referring to FIG. 6, FiLM 140 includes a fully connected layer 180, a ReLU (Rectified Linear Unit) layer 182, and a fully connected layer 184, which are connected to each other so as to receive an input of sigma level t with a dimensionality of C and output two C-dimensional output vectors h1 and h2, respectively. The outputs of the fully connected layer 184 are the C-dimensional vectors h1 and h2, respectively, as described above.
[0052] FiLM 140 further includes repeat units 186 and 188 that output matrices h1 and h2, each of which has T rows and C columns, by copying vectors h1 and h2 that are the output of fully connected layer 184.
[0053] FiLM140 further includes a layer normalization layer 190 that receives the noisy input 130 with T rows and C columns and performs layer normalization on the input 130, a multiplication unit 192 that multiplies the output of the layer normalization layer 190 element by element of a matrix h1, and an addition unit 194 that adds the corresponding element of a matrix h2 to each element of the output of the multiplication unit 192. The output of the addition unit 194 is input to the residual connection unit 144.
[0054] Fig. 7 shows a functional block diagram of the audio signal subsampling unit 154 and point-wise FiLM 146 shown in Fig. 5. Referring to Fig. 7, the audio signal subsampling unit 154 includes an embedding unit 210 using Wav2Vec 2.0 to convert the audio signal 132 into an embedding vector, and a downsampling unit 212 that downsamples the vector output by the embedding unit 210 to output the audio signal as an audio signal 156 in a matrix with T rows and C columns.
[0055] Point-wise FiLM 146 receives input of speech signal 156 and includes a fully connected layer 230, a ReLU layer 232, and a fully connected layer 234, each connected to output matrices h1 and h2, each of which has T rows and C columns. Point-wise FiLM 146 further includes a layer normalization unit 236 that performs layer normalization on the T rows and C columns of data output by the residual connection unit 144, a multiplication unit 238 that multiplies each element of the output of the layer normalization unit 236 by the corresponding element of matrix h1, and an addition unit 240 that adds each element of the output of the multiplication unit 238 by the corresponding element of matrix h2. Output 160 of the addition unit 240 is a matrix with T rows and C columns, and is provided to the input of a first decoder block 170 shown in FIG. 8.
[0056] 8 shows the configuration of the first decoder block 170. The configurations of the second decoder block 172, third decoder block 174, and fourth decoder block 176 are the same as that of the first decoder block 170. However, the parameters that define these functions are learned individually during learning.
[0057] 8, the first decoder block 170 has a configuration similar to the encoder part of a transformer. Specifically, the first decoder block 170 includes an attention processing unit 260 that performs predetermined attention calculations between elements of input data and outputs the results, a feedforward layer 262 that receives an output 296 from the attention processing unit 260 as an input, and an adder and normalizer unit 264 that performs residual connection and normalization on an output 298 of the feedforward layer 262 and outputs the result.
[0058] The attention processing unit 260 includes an inter-patch attention mechanism 270 that classifies frame data included in the input gesture data into multiple patches, performs attention calculation between the patches, and outputs the data in the same format as the original input data, and an addition and normalization unit 272 that performs residual connection and normalization on the output of the inter-patch attention mechanism 270. In this embodiment, each patch contains data for a predetermined number of consecutive frames. Each patch contains the same number of frame data.
[0059] The attention processing unit 260 further includes an intra-patch attention mechanism 274 for performing attention calculation within each patch on the output 292 of the adder and normalizer 272, and an adder and normalizer 278 for performing residual connection and normalization on the frame sequence 294 output by the intra-patch attention mechanism 274. The intra-patch attention calculation performed by the intra-patch attention mechanism 274 is an attention calculation between the data of frames included in each patch classified by the inter-patch attention mechanism 270.
[0060] 9 shows the configuration of the inter-patch attention mechanism 270. Referring to Fig. 9, the inter-patch attention mechanism 270 includes a depth-wise convolutional layer 332 that classifies the output 160 into a plurality of patches each consisting of a predetermined number of consecutive frames and performs depth-wise convolution on each patch to output patch data 334 equal to the number of patches, a multi-head self-attention mechanism 336 that calculates self-attention between the plurality of patch data 334 output by the depth-wise convolutional layer 332 and outputs a patch sequence 338 equal to the number of patches, and a transposed convolutional layer 340 that performs transposed convolution on the patch sequence 338, which is the inverse of the convolution performed by the depth-wise convolutional layer 332, to output an output 290 consisting of the same number of frames as the output 160.
[0061] Fig. 10 shows the functional configuration of the depthwise convolutional layer 332 shown in Fig. 9. In the following figures and explanation, the configuration of the depthwise convolutional layer 332 and the function of each part will be explained assuming T=310 and the number of patches=60, in order to make it easier to understand the specific content of the convolution.
[0062] Referring to FIG. 10, the depth-wise convolutional layer 332 includes a one-dimensional convolutional layer 360 with a kernel size of 10, a kernel stride of 5, and C as the number of groups, for the output 160 (310 rows and C columns) of the point-wise FiLM 146 shown in FIG. 7. The layer 360 treats each of the C elements of the input data as a separate channel and performs one-dimensional convolution in the T direction using C kernels for each channel. Because the kernel size and stride of the kernel applied to each channel are 10 and 5, respectively, (310-10) / 5=60 elements are calculated for each channel per kernel. Therefore, the output of the one-dimensional convolutional layer 360 includes C vectors, each of which has 60 convolution results as elements for each of the C channels.
[0063] To exchange information between channels, the depth-wise convolutional layer 332 further includes a one-dimensional convolutional layer 362 that performs channel-wise convolution on the output of the one-dimensional convolutional layer 360 using a kernel with a kernel width of 1, thereby outputting patch data 334 consisting of data in 60 rows and C columns. The patch data 334 includes 60 patches.
[0064] FIG. 11 shows the functional configuration of the transposed convolutional layer 340 shown in FIG. 11. Referring to FIG. 11, the transposed convolutional layer 340 includes a one-dimensional transposed convolutional layer 380 that receives a patch sequence 338 output by the multi-head self-attention mechanism 336 shown in FIG. 9 and performs transposed convolution on 60 patches included in the patch sequence 338 to output a frame sequence consisting of 310 frames. The transposed convolution by the one-dimensional transposed convolutional layer 380 is performed with the following settings: kernel width = 10, stride = 5, and number of groups = C. According to the transposed convolution processing based on these settings, the values of 300 frames in the time axis direction among the outputs of the one-dimensional transposed convolutional layer 380 are about to have the convolutional values added twice.
[0065] The transposed convolutional layer 340 further includes a division layer 382 that divides the element values of the central 300 frames, which are output from the one-dimensional transposed convolutional layer 380 and to which values obtained by two convolutions have been added, by two, and a one-dimensional convolutional layer 384 that performs one-dimensional convolution in the channel direction on the output of the division layer 382 using C kernels with a kernel width of 1, thereby outputting a frame sequence with 310 rows and C columns.
[0066] Fig. 12 shows a functional block diagram of the intra-patch attention mechanism 274 shown in Fig. 8. Referring to Fig. 12, the intra-patch attention mechanism 274 includes a multi-head self-attention mechanism 402 that receives the output 292 from the adder and normalizer 272, classifies (divides) the output into a patch sequence 400 such as patch 420, patch 422, ..., patch 424, calculates self-attention between frames included in each patch for each patch, and outputs a frame sequence 294 including patch 430, patch 432, ..., patch 434, etc., modified by intra-patch attention.
[0067] The noise removal unit 152 shown in FIG. 5 removes the sample x input to the diffusion model 90 at step t when generating data through the de-diffusion process. t From the estimated noise ε estimated by the diffusion model 90 θ (x t , t) 164 and the noise 128 sampled by the noise sampling unit 88, the noise-removed sample x at the next step t-1 is calculated by the following equation: t-1 and inputs it to the diffusion model 90 via a loop 158.
[0068]
number
[0069] 1.2. Program Structure 13 shows an outline of the control structure of a program for training the diffusion model 90 (diffusion process 60 shown in FIG. 2). Referring to FIG. 13, this program includes step 500 for repeating process 502 until the parameters of the diffusion model 90 converge. In this embodiment, process 502 is repeated until the parameters converge, but the present invention is not limited to such an embodiment. For example, execution of the program may be terminated when process 502 has been executed a predetermined number of times.
[0070] The process 502 includes a step 550 of sampling a training data sample x0 from the training data set. The sampling is done according to a uniform distribution, so that each sample is used for training approximately the same number of times.
[0071] Process 502 further includes step 552 of sampling the value of step t from [1, ..., T] according to a uniform distribution, step 554 of sampling a random number, which is an initial noise, from a standard normal distribution, and step 556 of providing the gesture data and the voice data of the training data as input 130 and voice signal 132 shown in FIG. 5 to diffusion model 90, respectively, and updating the parameters of diffusion model 90 by gradient descent according to the following equations:
[0072]
number
[0073] 14 shows the control structure of a program for generating data for the gesture generation system 110 shown in FIG. 4 using the diffusion model 90. Referring to FIG. 14, the program first extracts noise x TThe program includes step 600 of sampling x0, step 602 of repeating process 604 while varying the value of t from T to 1, and step 606 of outputting the sample x0 obtained in step 602 after the process of step 602 is completed and terminating the execution of the program.
[0074] The process 604 includes step 650, which branches the flow of control depending on whether the value of step t is greater than 1; step 652, which samples noise z from a standard normal distribution (by the noise sampling unit 88 shown in FIG. 4) when the determination in step 650 is affirmative; and step 654, which assigns 0 to noise z when the determination in step 650 is negative.
[0075] The process 604 further includes, after steps 652 and 654, determining the next sample x for the despreading process according to equation (1) above. t-1 After step 602 has been performed for all values of t from 1, . . . , T, the processing of step 602 ends and the next step 606 is performed.
[0076] 1.3 Operation 1.3.1 Training the Diffusion Model 90 3 and 13, the training system 70 operates as follows. A large number of training data are stored in advance in the training data storage unit 80. Each training data includes a speech signal and gesture data obtained from an image of the speaker while making the speech. The training control unit 82 reads the values of hyperparameters stored in the hyperparameter storage unit 84 and stores them in each variable as appropriate.
[0077] The learning control unit 82 samples the value of the variable t from a uniform distribution of 1, ..., T (step 552 in FIG. 13). That is, one of 1, ..., T is selected as the value of the variable t. The learning control unit 82 provides this value of the variable t to FiLM 140 (FIG. 5) of the diffusion model 90.
[0078] The learning data reading unit 86 reads out learning data from the learning data storage unit 80 under the control of the learning control unit 82, and separates the data into gesture data 92 and voice data 94. Referring to FIG. 5, the learning data reading unit 86 provides the gesture data 92 to the noise mixing unit 100. The noise sampling unit 88 samples noise from a standard normal distribution (step 554 in FIG. 13), and provides the noise data 92 and the noise mixing unit 100. The learning control unit 82 calculates a noise schedule α t and α ̄ t and provides the noise mixer 100. The noise mixer 100 mixes the gesture data 92 with the noise from the noise sampler 88 according to the noise schedule α t and provides the resulting mixture as input 130 (FIG. 5) to the FiLM 140 of the diffusion model 90. The training data reader 86 provides the voice data 94 as a voice signal 132 to the voice signal subsampling unit 154.
[0079] 5 and 6 and branch 142 shown in Fig. 5, input 130 conditioned by modulation with the value of variable t is provided to point-wise FiLM 146. Point-wise FiLM 146 converts the gesture data of this input 130 element by element using audio signal 156 from audio signal subsampling unit 154, and inputs it as output 160 to first decoder block 170. By this process, the gesture data input to first decoder block 170 is conditioned by the audio signal by being modulated by the audio signal.
[0080] The first decoder block 170, the second decoder block 172, the third decoder block 174, and the fourth decoder block 176 perform processing on the input data in this order, and provide the results to the fully connected layer 150 as output data 162.
[0081] 8 to 12, for example, in the first decoder block 170, the inter-patch attention mechanism 270 calculates the inter-patch attention for the output 160. The method for calculating the inter-patch attention is as shown in FIG. 9. As shown in FIG. 10, the frame sequence to the depth-wise convolutional layer 332 is classified into multiple patches by the one-dimensional convolutional layer 360 and one-dimensional convolutional layer 362. In this embodiment, each patch contains the same number of consecutive frames. The attention between these multiple patches is calculated by the multi-head self-attention mechanism 336 shown in FIG. 9 and output as a patch sequence 338. The patch sequence 338 is converted back to the original frame sequence by the one-dimensional transposed convolutional layer 380, division layer 382, and one-dimensional convolutional layer 384 of the transposed convolutional layer 340 shown in FIG. 11, and output as the output 290 from the inter-patch attention mechanism 270.
[0082] The frame sequence of output 290 is input to intra-patch attention mechanism 274 as output 292 of adder and normalizer 272. The method of calculating intra-patch attention in intra-patch attention mechanism 274 is as shown in Figure 12. The frame data included in output 292 is classified by patch. For each patch, attention between frames belonging to that patch is calculated by multi-head self-attention mechanism 402 and output from intra-patch attention mechanism 274 as frame sequence 294.
[0083] Returning to FIG. 8, the frame sequence 294 from the intra-patch attention mechanism 274 passes through the adder and normalizer 278, the feedforward layer 262 and the adder and normalizer 264 and is input as frame sequence 200 to the first decoder block 170 and the second decoder block 172.
[0084] Similar processing is repeated in the second decoder block 172, the third decoder block 174, and the fourth decoder block 176, and the data is input as output data 162 from the decoder 148 shown in Fig. 5 to the fully connected layer 150. The estimated noise 164 output from the fully connected layer 150 is input to the parameter update unit 98 shown in Fig. 3.
[0085] As shown in steps 554 and 556 of FIG. 13, the parameter update unit 98 uses the estimated noise 164, which is the output of the fully connected layer 150, the current noise, the noise schedule, and the value of the variable t indicating the sigma level, and updates the parameters of the diffusion model 90 by the gradient descent method in accordance with the equation shown in equation (2) above.
[0086] In this embodiment, the above process is repeated until the parameters of the diffusion model 90 converge, and training of the diffusion model 90 is terminated when the parameters of the diffusion model 90 converge. Note that the parameters converge does not necessarily mean that all parameter values have stopped changing or that the change is below (or less than) a certain threshold. For example, the parameters may be considered to have converged when the values of a predetermined number or a predetermined percentage of the parameters have stopped changing, or when the number of parameters that have changed is below a threshold.
[0087] The condition for ending training is not limited to parameter convergence. For example, training may be ended when the number of repetitions of training using learning data reaches a predetermined number.
[0088] 1.3.2 Gesture Data Generation Once the training of the diffusion model 90 is complete, gesture data can be generated by the gesture generation system 110 shown in Fig. 4. The operation of the gesture generation system 110 when generating gestures will be described with particular reference to Fig. 14.
[0089] To generate gesture data, speech data of the speech to be used for gesture generation is required. This speech data is input to the speech signal subsampling unit 154 as the speech signal 132 in Fig. 5, and is then input to the pointwise FiLM 146 as a sequence of speech data frames with the same number of frames as the gesture signal.
[0090] Referring to FIG. 14, in generating gesture data, first, a sample x T is sampled from the standard normal distribution (step 600 in FIG. 14). This sampling is performed by the noise sampling unit 122 shown in FIG. 4. When the noise 124 in FIG. 4 is sampled from the initial sample x T The initial sample x T is given to the diffusion model 90.
[0091] 14 is repeated while the value of the variable t is decremented by 1 from T to 1. In each repetition, the diffusion model 90 performs the same process as during learning on the input noise to estimate the estimated noise 164 (FIG. 5) and provides it to the noise removal unit 152.
[0092] The noise removal unit 152 determines whether the value of the variable t is greater than 1 (step 650 in FIG. 14). At the beginning of the repetition, this determination is negative. Control proceeds to step 652. In step 652, the noise removal unit 152 samples the noise z from the standard normal distribution. This processing corresponds to the processing by the noise sampling unit 126 shown in FIG. 4.
[0093] The noise removal unit 152 further calculates x in step 656 of FIG. 14 according to the above-described formula (1). T-1 Since the value of variable t is T, subtract 1 from the value of variable t and calculate x T-1 is again input to the diffusion model 90. This process is repeated until the value of the variable t becomes 1. That is, in step 652, the noise z is sampled, and in step 656, the sample x is calculated by equation (1). t-1 , sample xt and noise z, and the process of calculating the value of variable t by 1 is repeated until the value of variable t becomes 1.
[0094] When the value of variable t becomes 1, the determination in step 650 in Figure 14 becomes negative, and step 654 is executed. That is, noise z is set to 0, and in step 656, sample x0 is generated based on sample x1 with noise z = 0. The iterative process of process 604 ends, and x0 is output as the sample generated by diffusion model 90.
[0095] 1.3 Effects In this embodiment, during both training and generation, a sequence of gesture data frames is classified into multiple patches, and attention between patches and attention within patches is calculated to generate gesture data. Gesture data is sequence data that follows time. Conventionally, attention has been calculated using all frames as input. However, as described above, gesture data is sequence data, and if all frames are input, the characteristics of the sequence data may not be reflected in the generated gesture data.
[0096] In contrast, in this embodiment, we group consecutive frames into patches and use both inter-patch and intra-patch frame attention to train the diffusion model and generate gesture data, resulting in gestures that are more natural and synchronized with speech, closer to human gestures.
[0097] Furthermore, in the above embodiment, the computational complexity of attention calculation is proportional to the sum of the square of the number of patches and the product of the square of the patch size (the number of frames in a patch) and the number of patches. When the computational complexity is calculated based on a formula, a graph such as that shown in FIG. 15 is obtained. In FIG. 15, graph 680 assumes T=100, graph 682 assumes T=300, and graph 684 assumes T=500. From these graphs, it can be seen that in all cases, the computational complexity is minimized around patch size=10, as indicated by line 670. Furthermore, from the graph in FIG. 15, it can be seen that the patch size with the lowest computational complexity gradually increases as the value of T increases. Considering that the time required for calculation varies depending on the type of calculation and the amount of data, which depend on conditions such as the hardware used, the computational complexity required to obtain natural gesture data can be reduced by selecting a patch size in a certain range indicated by lines 672 and 674, centered on line 670, which indicates the theoretically minimum computational complexity, as shown in FIG. 15.
[0098] Second Embodiment FIG. 16 shows an example of gesture data 700 obtained using a diffusion model such as that of the first embodiment. The gesture data 700 is a drawing in which the joint positions at each time of the gesture are superimposed. The gesture data 700 realizes natural movement, but using the gesture data 700 as is can sometimes cause problems. For example, when a robot is made to make a gesture according to this gesture data, if there is an object near the body, the robot's arm may hit the object, causing problems. This is not limited to robots, but is also true for avatars. In the case of an avatar, the avatar will not actually come into contact with anything in the vicinity, but may overlap with other objects in the image. Therefore, it is necessary to avoid such a situation. The second embodiment is an embodiment that avoids such problems.
[0099] One way to avoid such problems is to prepare a virtual tolerance frame 710 that defines the joint positions of a robot or the like that moves based on gesture data 700, as shown in Fig. 17. In the example shown in Fig. 17, the joints of the robot are outside the tolerance frame 710 in areas 712 and 714.
[0100] A simple method that can be considered for dealing with this situation is shown in Figure 18. Referring to Figure 18, if any joints are outside tolerance frame 710, as shown in areas 712 and 714 in Figure 17, they are moved to the boundary of tolerance frame 710, as shown by areas 720 and 722. If any joints are inside tolerance frame 710, there is at least no need to worry about the robot's arm hitting an object outside it.
[0101] However, using such a method has the problem that the gestures become unnatural. That is, referring to Fig. 19, according to gesture data 730 corrected by the method shown in Fig. 18, the movement trajectory of the robot's joints may become linear, as shown by areas 732 and 734. Normally, humans do not perform gestures that include such linear movements. Therefore, the robot's movements appear unnatural. The second embodiment solves this problem.
[0102] FIG. 20 shows a schematic configuration of a diffusion model 750 according to the second embodiment. The diffusion model 750 differs from the diffusion model 90 according to the first embodiment in that it includes, after the noise removal unit 152, an allowable range memory unit 766 for storing data defining an allowable range for the positions of the robot's joints, and a clipping unit 760 connected to the allowable range memory unit 766. The clipping unit 760 receives gesture data 762 output from the noise removal unit 152. If, at each step of the de-diffusion process performed by the diffusion model 750, any joints of the robot or the like are outside an allowable range 768 defined by the data stored in the allowable range memory unit 766, the clipping unit 760 clips the joints so that they are positioned on the nearest boundary. Furthermore, when the value of a variable t representing the sigma level is greater than 1, the clipping unit 760 re-inputs the clipped gesture data to the FiLM 140 via a loop 764. When the value of variable t is equal to 1, clipping unit 760 outputs clipped gesture data 752 as the generated gesture data.
[0103] It should be noted that the clipping unit 760 is not used when training the diffusion model 750. Therefore, training of the diffusion model 750 can be performed in the same way as the diffusion model 90 according to the first embodiment.
[0104] Fig. 21 shows a control structure of a program for generating gesture data using diffusion model 750 shown in Fig. 20. The flowchart shown in Fig. 21 differs from that shown in Fig. 15 in that it includes step 780 for repeating process 782 instead of step 602 for repeating process 604 shown in Fig. 15.
[0105] The structure of the first half of process 782 is similar to that of process 604. Process 782 differs from process 604 in that, after step 656, it includes steps 790, 792, 794, and 796 for moving the position of any joint outside the allowable movement frame to the boundary of the allowable movement frame. Note that this embodiment is intended for calculating robot gestures. When controlling the movement of a robot, the rotation angle of each joint is often used as gesture data. When specifying a gesture using a rotation angle and restricting the position of each joint within an allowable range, it is first necessary to know the position of each joint. To do this, it is necessary to calculate the three-dimensional coordinates of each joint from the rotation angle of each joint using so-called forward kinematics.
[0106] Step 790 is a process of calculating the three-dimensional coordinates of each joint by forward kinematics based on the rotation angle of each joint of the gesture data 762 generated by the processes up to the noise removal unit 152. Step 792 is a process of calculating the three-dimensional coordinates of each joint by forward kinematics based on the three-dimensional coordinates of each joint calculated in step 790. t-1 If the result of this determination is negative, the sample x t-1 If the determination in step 792 is affirmative, then in step 794, the joints outside the tolerance frame are moved to the closest position within the tolerance frame. Then, in step 796, the rotation angle of each joint is calculated by applying so-called inverse kinematics to the three-dimensional coordinates of each joint corrected in step 794, and the sample x is calculated using the calculated value. t-1 Replace.
[0107] This process is performed each time step 782 is executed. Therefore, even if the joint position moves to the boundary of the tolerance frame by the process of step 794, there is a high possibility that the position will change again due to new noise in the next iteration. As a result, it is expected that few joints will ultimately be positioned on the tolerance frame, and even if any, they will quickly move away from the tolerance frame over time. As a result, the positions of each joint defined by the gesture data obtained by this embodiment will be natural, without following a linear trajectory.
[0108] 22 shows an example of gesture data 800 in which joint positions have been corrected according to this embodiment, along with the joint position tolerance range frame 710. As shown in areas 810 and 812, linear movements have been removed from the movements of each joint, achieving natural movements.
[0109] Table 820 shown in Fig. 23 shows the results of calculating the evaluation FGD (Frechet Gesture Distance) proposed in Non-Patent Document 2 for the results of clipping (Clip Generation) according to this embodiment and the results of clipping (Clip) alone, which simply moves joints along the boundary. As can be seen from Fig. 23, the FGD of the gesture according to this embodiment shows a significant improvement over the FGD of the gesture according to clipping alone.
[0110] The above description of the embodiment is based on two-dimensional gesture data. However, in the case of a robot, natural three-dimensional processing is naturally required, and natural data can be easily realized by following the above description. In the above embodiment, the joint positions are corrected in each step of the de-diffusion process. However, the present invention is not limited to such an embodiment. The joint positions may be corrected in only some of the steps of the de-diffusion process. In this case, it is desirable to perform processing to correct the joint positions at least in the final step of the de-diffusion process.
[0111] In the above embodiment, the allowable ranges of the robot's joint positions are specified. It goes without saying that a storage area for specifying such allowable ranges is required. Furthermore, considering that the environment changes as the robot moves, it is desirable that the data specifying the allowable ranges be dynamically updated.
[0112] Furthermore, in the above embodiment, an allowable range for the robot's joint positions is specified. However, the present invention is not limited to such an embodiment. Conversely, a prohibited range for the robot's joint positions may be specified. In a situation where the robot and surrounding objects are moving, it is simpler to set an independent prohibited range for each object in the environment and prevent the robot's joints from entering that prohibited range.
[0113] To more reliably eliminate the risk of a joint going outside the tolerance range, if there is such a joint, the position of the joint can be set slightly inside the tolerance range rather than on the boundary of the tolerance range. In this case, the gesture is more likely to be slightly smaller, but the possibility of the joint remaining on the boundary of the tolerance range is reduced.
[0114] Third Embodiment In the second embodiment, only the position of each joint of the robot is limited within an allowable range. However, in the case of a robot, it may also be necessary to limit the rotation angle of each joint within an allowable range. This is because, in the case of a robot, each joint is realized by a mechanical rotation mechanism. In the case of a mechanical rotation mechanism, the rotation angle of the joint is limited within a certain range depending on the mechanism. It is impossible to rotate the joint beyond that range. Therefore, in the case of a robot, it may be necessary to consider not only the allowable range of the coordinates of each joint, but also the allowable range of the rotation angle. The diffusion model according to the third embodiment solves such a problem.
[0115] 24, a diffusion model 840 according to the third embodiment differs from the diffusion model 750 according to the second embodiment shown in FIG. 20 in that a clipping unit 850 stored in an allowable range storage unit 766 is included instead of the clipping unit 760 shown in FIG. 20. The clipping unit 850 reads an allowable range 768 from the allowable range storage unit 766 and clips the rotation angle of each joint and the three-dimensional coordinates of each joint so that gesture data 762 output from the noise removal unit 152 falls within the boundary of the allowable range 768. The output of the clipping unit 850 is connected to a loop 852, which corresponds to the loop 764 shown in FIG. 20. In the final step of the de-diffusion process in the diffusion model 840, the clipping unit 850 outputs gesture data 842. In the other steps, the output of the clipping unit 850 is re-input to the FiLM 140 via the loop 852.
[0116] Fig. 25 shows the control structure of a program for implementing diffusion model 840 on a computer. This program differs from the program shown in Fig. 21 in that it includes step 900, which repeatedly executes process 902 for t = T, ..., 1, instead of step 780 shown in Fig. 21.
[0117] The process 902 has almost the same configuration as the process 782 shown in Fig. 21. However, in addition to the configuration of the process 782 shown in Fig. 21, the process 902 uses the gesture data x generated by the process up to step 656. t-1 This process differs from process 782 in that it includes step 910, which branches the control flow depending on whether there is an element outside the limit frame among them, and step 912, which, if the determination result in step 910 is positive, corrects the corresponding element to the value of the closest limit frame and then advances control to step 790.
[0118] In this embodiment, x t-1It should be noted that each element in indicates the rotation angle of each joint of the robot. That is, in step 912, when the rotation angle of each joint of the robot is outside the mechanically (physically) rotatable angle range (angle frame), the value is corrected to a value within the angle frame. This angle frame is defined by an angle range determined by upper and lower limits centered on the reference direction of the angle joint. Each joint usually has three angles as degrees of freedom. A frame that restricts the movement of the joint is independently defined for each degree of freedom for each joint. In step 910, it is determined whether each of these three angles of each joint is within the angle frame. If there is even one angle that is not within the angle frame, that angle is corrected in step 912 to be within the angle frame (usually the angle equal to the upper or lower limit closest to the angle of the element that determines the restriction frame).
[0119] If the determination in step 910 is negative, control proceeds directly to step 790. t-1 If the rotation angles for each of the above are within the angle window, control proceeds to step 790.
[0120] The processing from step 790 onwards is the same as in the second embodiment.
[0121] As described above, according to this embodiment, in each step of the de-diffusion process, the rotation angles of all joints included in the gesture data are restricted to fall within the corresponding angle range. In the de-diffusion process, the next rotation angle is probabilistically sampled based on the value obtained for the previous rotation angle in each step. Therefore, even after any angle of any joint is restricted and corrected, it is highly likely that it will change to a different value within the angle range in the next step. This prevents the rotation angles of each joint of the robot from being fixed, resulting in unnatural movement. Furthermore, in each step, the three-dimensional position of each joint determined by the rotation angle of each joint corrected in this manner is similarly determined to be within a predetermined limiting frame, and is corrected to fall within the limiting frame if necessary. This reduces the possibility that the robot's gestures will follow a linear trajectory or become fixed, resulting in more natural robot gestures.
[0122] The above embodiments have been described assuming either an avatar or a robot, for example. However, it is clear that these embodiments are not limited to only one of them. That is, even for an avatar, when a gesture is calculated assuming a three-dimensional skeleton similar to that of a robot, the method described in the second or third embodiment, for example, can be applied. Conversely, even for a robot, if there is no limit to the rotation angle of a joint for forming a gesture (for example, if a certain joint can rotate 360 degrees in a certain plane), an angle frame for restricting the rotation of the joint in that plane is not necessarily required. Conversely, even if a certain joint of a robot can rotate 360 degrees in a certain plane, it may be necessary to intentionally impose a limit on the rotation angle. In such cases, a configuration such as the third embodiment is required.
[0123] Furthermore, in any of the above embodiments, the gesture limit frame of the avatar, robot, etc. is fixed in advance and stored in a storage device. However, the present invention is not limited to such an embodiment. Surrounding objects such as avatars and robots do not always remain stationary. Therefore, the above-mentioned limit frame may be updated in real time according to the movement of surrounding objects or the avatar or robot itself.
[0124] 4. Computer implementation Fig. 26 is an external view of an example of a computer system for realizing the above embodiment, and Fig. 27 is a block diagram showing an example of the hardware configuration of the computer system shown in Fig. 26.
[0125] 26, this computer system 950 includes a computer 970 having a DVD (Digital Versatile Disc) drive 1002, and a keyboard 974, a mouse 976, and a monitor 972 for interacting with a user, all of which are connected to the computer 970. Of course, these are just one example of a configuration for when user interaction is required, and any general hardware and software that can be used for user interaction (for example, a touch panel, voice input, or a general pointing device) can be used.
[0126] 27, the computer 970 includes, in addition to a DVD drive 1002, a CPU (Central Processing Unit) 990 which is a processor, a GPU (Graphics Processing Unit) 992, a bus 1010 connected to the CPU 990, the GPU 992, and the DVD drive 1002, a ROM (Read-Only Memory) 996 connected to the bus 1010 and storing a boot-up program and the like for the computer 970, a RAM (Random Access Memory) 998 connected to the bus 1010 and storing instructions constituting a program, a system program, working data, and the like, and an SSD (Solid State Drive) 1000 which is a non-volatile memory connected to the bus 1010. The SSD 1000 is used to store programs executed by the CPU 990 and the GPU 992, as well as data used by the programs executed by the CPU 990 and the GPU 992. The computer 970 further includes a network I / F (Interface) 1008 that provides connection to a network 986 that enables communication with other terminals, and a USB port 1006 to which a USB (Universal Serial Bus) memory 984 can be attached or detached and that provides communication between the USB memory 984 and each part within the computer 970.
[0127] The computer 970 further includes an input / output I / F 1004 that is connected to a microphone 982, a speaker 980, a motion capture device (not shown), and a bus 1010, and that reads out audio signals, image signals, and text data generated by the CPU 990 and stored in the RAM 998 or SSD 1000 according to instructions from the CPU 990, converts them to analog, amplifies them, and drives the speaker 980, digitizes analog audio signals from the microphone 982, and stores them at any address in the RAM 998 or SSD 1000 specified by the CPU 990, and receives motion capture signals from the motion capture device and stores them at an address specified by the CPU 990.
[0128] In the above embodiment, the programs for realizing the training system 70 shown in Fig. 3, the gesture generation system 110 and each part thereof shown in Fig. 4, the diffusion model 750 and each part thereof shown in Fig. 20, and the diffusion model 840 and each part thereof shown in Fig. 24, the neural network parameters, the neural network program, etc. are all stored in, for example, the SSD 1000, RAM 998, DVD 978, or USB memory 984 shown in Fig. 27, or a storage medium of an external device (not shown) connected via the network I / F 1008 and the network 986. Typically, these data, parameters, etc. are written to the SSD 1000 from the outside, and loaded into the RAM 998 when executed by the computer 970.
[0129] Computer programs for operating this computer system to realize the functions of the training system 70 shown in FIG. 3 , the gesture generation system 110 and each part thereof shown in FIG. 4 , the diffusion model 750 and each part thereof shown in FIG. 20 , and the diffusion model 840 and each part thereof shown in FIG. 24 are stored on a DVD 978 inserted in the DVD drive 1002 and transferred from the DVD drive 1002 to the SSD 1000. Alternatively, these programs may be stored in a USB memory 984, and the USB memory 984 may be inserted in the USB port 1006 and the programs may be transferred to the SSD 1000. Alternatively, the programs may be transmitted to the computer 970 via the network 986 and stored in the SSD 1000.
[0130] The program is loaded into RAM 998 at run time.
[0131] A program that cooperates with the computer 970 to implement the functions of the systems and their components according to the above-described embodiments includes a plurality of instructions written and arranged to cause the computer 970 to operate to implement those functions. Some of the basic functions required to execute those instructions are provided by the operating system (OS) or third-party programs running on the computer 970, or by modules of various toolkits installed on the computer 970. Therefore, the program does not necessarily include all of the functions required to implement the systems and methods of the embodiments. The program need only include instructions that execute the operations of the above-described devices and their components by statically linking appropriate functions or "programming toolkit" functions in a controlled manner to achieve the desired results, or by dynamically linking to those functions during program execution. The method for operating the computer 970 in this manner is well known and will not be repeated here.
[0132] The GPU 992 is capable of parallel processing, allowing it to simultaneously execute large amounts of calculations associated with machine learning and testing in a parallel or pipelined manner.
[0133] The embodiments disclosed herein are merely examples, and the present invention is not limited to the above-described embodiments. The scope of the present invention is defined by the claims in the appended claims, taking into consideration the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wordings described therein. [Explanation of symbols]
[0134] 60 Diffusion Process 62 Dediffusion Process 70 Training System 80 Learning data storage unit 82 Learning control unit 84 Hyperparameter memory 86 Learning data reading unit 88, 122, 126 Noise sampling section 90, 750, 840 diffusion models 92, 114, 700, 730, 762, 842 gesture data 94, 112 audio data 96, 124, 128 Noise 98 Parameter Update Section 120 Generation control unit 140 Films 146 Pointwise FiLM 148 decoder 152 Noise removal section 164 Estimated Noise 170 First Decoder Block 172 Second Decoder Block 174 3rd decoder block 176 4th decoder block 182, 232 ReLU layer 200, 294 frame sequence 210 Embedded part 260 Attention Processing Unit 270 Inter-patch attention mechanism 274 In-patch attention mechanism 332 Depthwise Convolutional Layers 336, 402 Multi-head self-attention mechanism 340 Transposed Convolution Layer 360, 362, 384 1D convolutional layer 380 1D Transposed Convolution Layer 382 Division layer 710 Tolerance Frame 760, 850 clipping section 766 Tolerance Range Memory 950 Computer Systems
Claims
1. a memory for storing instructions; a processor connected to the memory and configured to execute the set of instructions to perform a predetermined process, the processing includes a gesture data generation process for generating gesture data adapted to the voice data from noise conditioned by the voice data by a de-diffusion process using a pre-trained machine learning algorithm; the noise comprises a plurality of frames conditioned by the audio data; The plurality of frames are grouped into a plurality of patches; The machine learning algorithm an inter-patch attention mechanism using inter-patch attention, which is attention between the plurality of patches; an intra-patch attention mechanism that uses intra-patch attention, which is attention between the plurality of frames included in each of the plurality of patches; and generating the gesture data by the de-diffusion process using the
2. The gesture generation system according to claim 1 , wherein each of the plurality of patches includes two or more of the frames that are consecutive to one another.
3. The gesture generation system according to claim 1 or 2, wherein each of the plurality of patches includes an equal number of the frames.
4. The gesture generation system according to claim 1 or claim 2, wherein the machine learning algorithm sequentially executes the inter-patch attention mechanism and the intra-patch attention mechanism on the plurality of frames.
5. The gesture generation system of claim 4 , wherein the machine learning algorithm performs the inter-patch attention mechanism prior to the intra-patch attention mechanism for the plurality of frames.
6. each of the plurality of frames includes a numerical value for controlling a three-dimensional position of each joint of a predetermined skeleton; the despreading process repeats a process of performing a noise removal process on the plurality of frames to generate a plurality of output frames, and further inputting the plurality of output frames to the noise removal process to perform the noise removal process until a predetermined termination condition is met; 3. The gesture generation system according to claim 1, wherein the de-diffusion process further includes a restriction process for restricting the numerical values included in the plurality of output frames obtained in each of the noise removal processes of the de-diffusion process to within a predetermined range.
7. the memory includes an area for storing a restriction value for defining a range of motion of the skeleton; 7. The gesture generation system according to claim 6, wherein the restriction process includes a process of restricting the numerical values included in the plurality of output frames obtained by the de-diffusion process within a range defined by the restriction value in each of the repetitions of the de-diffusion process.
8. the memory includes an area for storing a restriction value for defining a range of motion of the skeleton; 7. The gesture generation system according to claim 6, wherein the restriction process includes a process of restricting the numerical values included in the plurality of output frames obtained by the de-diffusion process to be outside a range defined by the restriction value in each of the repetitions of the de-diffusion process.
9. The gesture generation system according to claim 1 , wherein each of the plurality of frames includes a vector whose elements are rotation angles of each joint of a predetermined skeleton.
10. The gesture generation system according to claim 1 , wherein each of the plurality of frames includes a vector whose elements are three-dimensional coordinates of each joint of a predetermined skeleton.
11. The gesture generation system according to claim 9 or 10, wherein the predetermined skeleton comprises a human skeleton.
12. The gesture generation system according to claim 11 , wherein the predetermined skeleton includes at least an upper body skeleton of a human.
13. a memory for storing at least instructions; a processor connected to the memory and configured to execute the set of instructions to perform a predetermined process, The processing is a gesture generation processing for generating gesture data adapted to voice data from noise conditioned by voice data by a de-diffusion process using a machine learning algorithm that has been trained in advance by a diffusion process, the gesture data includes a time series of numerical values for controlling the three-dimensional position of each joint of the skeleton; The memory stores a limiting value for defining a range of each of the numerical values generated by the gesture generation process during each iteration of the de-diffusion process.
14. A gesture generation method for generating gesture data adapted to voice data from noise conditioned by voice data by a de-diffusion process using a machine learning algorithm that has been trained in advance by a diffusion process, the method comprising: the noise comprises a plurality of frames conditioned by the audio data; The plurality of frames are grouped into a plurality of patches; The machine learning algorithm modifying the plurality of frames using inter-patch attention, the inter-patch attention being attention between the plurality of patches; and a step in which, in each of the plurality of patches, the computer modifies the plurality of frames using intra-patch attention, which is attention between the plurality of frames included in that patch.
15. an initial noise sampling means for sampling an initial noise from a predetermined distribution; a gesture generation unit including a machine learning model that has been trained in advance by a diffusion process to estimate noise based on input data to which noise has been added, the input data includes a plurality of frames; The plurality of frames are grouped into a plurality of patches; the gesture generation unit generates the gesture data by a de-diffusion process consisting of multiple steps, which is a de-diffusion process for the diffusion process; a selection unit that selects either an input to the gesture generation unit or an input from the machine learning model; a conditioning unit that conditions an output of the selection unit using a sigma level that defines a noise schedule that is determined in association with the plurality of steps and input audio data; a sampling means for estimating a new distribution of the gesture data in response to the machine learning model outputting gesture data candidates in accordance with the input conditioned by the conditioning unit, and sampling a new sample from the estimated distribution; a branching unit that selects, depending on whether the plurality of steps have been completed, either a process of outputting the new sample as the gesture data or a process of providing the new sample to the selection unit as the input from the machine learning model; the selection unit selects an input to the gesture generation unit and an output from the branch unit according to whether or not the plurality of steps have started; The machine learning model is an inter-patch attention mechanism that corrects the plurality of frames by inter-patch attention, which is attention between the plurality of patches; A gesture generation system including: a mechanism for modifying the plurality of frames using intra-patch attention, which is attention within the plurality of patches.