Learning processing device, inference processing device, and method therefor

The learning processing device addresses artifacts in rhythmic gestures by classifying audio data and attenuating beat gestures, improving the synchronization and naturalness of generated gestures.

WO2026009633A1PCT designated stage Publication Date: 2026-01-08SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/020451
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-05
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing gesture generation models often produce artifacts such as finger penetration and unnatural temporal fluctuations when executing rhythmic gestures based on audio data, requiring skilled technicians for correction.

Method used

A learning processing device that includes a training dataset storage unit, determination unit, audio feature extraction unit, and teacher data configuration unit to classify and generate learning model parameters for rhythmic gestures, reducing artifacts by attenuating beat gesture movements based on risk assessment and audio characteristics.

Benefits of technology

Improves the accuracy of generating natural gestures by minimizing artifacts and unnatural temporal fluctuations, enhancing the synchronization of gestures with audio rhythms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025020451_08012026_PF_FP_ABST
    Figure JP2025020451_08012026_PF_FP_ABST
Patent Text Reader

Abstract

In this learning processing device, it is determined whether motion data acquired as learning data includes a specific state including at least one of contact between a plurality of parts of a human body model or a pose transition of the human body model. In addition, audio data synchronized with the motion data is classified into at least first audio data and second audio data having a feature different from the first audio data. When the motion data includes a specific state, first teacher data including the first and second audio data is generated, and when the motion data does not include the specific state, second teacher data including the first audio data and not including the second audio data is generated. A learning model parameter for generating a beat gesture corresponding to the rhythm of conversation represented by input audio data is generated on the basis of the first teacher data and the second teacher data.
Need to check novelty before this filing date? Find Prior Art

Description

Learning processing device, inference processing device, and methods thereof

[0001] The present technology relates to a learning processing device, an inference processing device, and methods thereof for generating gestures based on voice data.

[0002] Conventionally, there are known techniques for changing the movement of a model of a person or the like based on speech. For example, Patent Document 1 discloses a method for training a gesture generation model, which employs adversarial learning to switch between a gesture matrix output from a neural network and a gesture matrix of training data divided into a plurality of segments at predetermined times and evaluate the gesture matrix as a gesture matrix relative to the speech matrix, thereby enabling natural use of gestures appropriate to the content of speech (see, for example, paragraphs

[0032] to

[0063] and Figures 3 to 8 of the specification of Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2023-139731

[0004] A commonly known type of gesture is a rhythmic gesture, which is a movement that follows the rhythm of conversation. When a motion model executes a rhythmic gesture based on audio data, artifacts such as the model's fingers penetrating each other may occur. Furthermore, when the model transitions from one pose to another, the rhythmic gesture movement may cause unnatural temporal fluctuations. Correcting such artifacts and unnatural temporal fluctuations may require the work of a skilled video technician.

[0005] In view of the above circumstances, an object of the present technology is to provide a learning processing device, an inference processing device, and methods thereof that can improve the accuracy of generating natural gestures from voice data.

[0006] One aspect of the present technology is a learning processing device including a training dataset storage unit, a determination unit, an audio feature extraction unit, a teacher data configuration unit, and a learning unit. The training dataset storage unit is configured to acquire training data including motion data of a human body model and audio data synchronized with the motion data. The determination unit is configured to determine whether the motion data includes a specific state including at least one of contact between multiple parts of the human body model and a pose transition of the human body model. The audio feature extraction unit is configured to classify the audio data into at least first audio data and second audio data having characteristics different from the first audio data. The teacher data configuration unit is configured to generate first teacher data including the first audio data and the second audio data if the motion data includes the specific state, and to generate second teacher data including the first audio data but not the second audio data if the motion data does not include the specific state. The learning unit is configured to generate learning model parameters for generating rhythmic gestures corresponding to the rhythm of conversation represented by input audio data, based on the first teacher data and the second teacher data.

[0007] One form of the present technology is a learning processing method including the following processes: acquiring learning data including motion data of at least one human body model and audio data synchronized with the motion data; determining whether the motion data includes a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model; classifying the audio data into at least first audio data and second audio data having characteristics different from the first audio data; generating first training data including the first audio data and the second audio data if the motion data includes the specific state; generating second training data including the first audio data but not the second audio data if the motion data does not include the specific state; and generating learning model parameters for generating rhythmic gestures corresponding to the rhythm of conversation represented by input audio data based on the first training data and the second training data.

[0008] One aspect of the present technology is an inference processing device including a voice data acquisition unit and an inference unit. The voice data acquisition unit is configured to acquire input voice data. The inference unit is configured to acquire an inference network for generating at least one rhythmic gesture of a human body model corresponding to the rhythm of a conversation represented by the input voice data, and to acquire learning model parameters of the inference network corresponding to first teacher data representing a plurality of voice data having different characteristics and second teacher data representing voice data having a single characteristic. The inference unit is further configured, based on the input voice data, the inference network, and the learning model parameters, to generate a non-rhythmic gesture synchronized with the input voice data when the input voice data corresponds to the first teacher data, and to generate the rhythmic gesture synchronized with the input voice data when the input voice data corresponds to the second teacher data. The non-rhythmic gesture corresponds to a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model.

[0009] One form of the present technology is an inference processing method including the following processes: acquiring input speech data, acquiring an inference network for generating at least one rhythmic gesture of a human body model corresponding to the rhythm of a conversation represented by the input speech data, acquiring learning model parameters of the inference network corresponding to first training data representing a plurality of speech data having different characteristics and second training data representing speech data having a single characteristic, and generating a non-rhythmic gesture synchronized with the input speech data when the input speech data corresponds to the first training data, and generating the rhythmic gesture synchronized with the input speech data when the input speech data corresponds to the second training data, based on the input speech data, the inference network, and the learning model parameters. The non-rhythmic gesture corresponds to a specific state including at least one of contact between a plurality of parts of the human body model or a pose transition of the human body model.

[0010] 1 is a block diagram showing an example of a configuration of an information processing device according to a first embodiment. FIG. 1 is a block diagram showing an example of a configuration of a risk determination unit. FIG. 2 is a diagram showing a graph related to collision determination. FIG. 2 is a block diagram showing an example of a configuration of a pause transition determination unit. FIG. 3 is a diagram showing a graph related to reliability on a time axis. FIG. 3 is a block diagram showing an example of a configuration of a movement information processing unit. FIG. 4 is a block diagram showing an example of a configuration of a movement feature extraction unit. FIG. 5 is a block diagram showing an example of a configuration of an emotion recognition unit. FIG. 6 is a block diagram showing an example of a configuration of a teacher data configuration unit. FIG. 7 is a flowchart showing processing of a combination data generation unit. FIG. 8 is a graph showing the number of heterogeneous voice feature data based on risk information. FIG. 9 is a schematic diagram showing an example of speech division. FIG. 10 is a block diagram showing an example of a configuration of a movement feature selection unit. FIG. 11 is a block diagram showing an example of a configuration of an information processing device according to a second embodiment. FIG. 12 is a block diagram showing an example of a configuration of a movement information processing unit. FIG. 13 is a schematic diagram in which movement information representing a gesture is likened to a waveform. FIG. 14 is a block diagram showing an example of a configuration of a time smoothing processing unit. FIG. 15 is a block diagram showing an example of a configuration of an information processing device according to a third embodiment. FIG. 16 is a block diagram showing an example of a configuration of a movement feature selection unit. FIG. 17 is a block diagram showing another example of a configuration of a pause transition determination unit. FIG. 18 is a diagram showing a graph related to reliability in speed or acceleration. FIG. 19 is a block diagram showing another example of a configuration of a movement information processing unit. FIG. 19 is a block diagram showing an example of a configuration of a movement feature extraction unit.

[0011] Hereinafter, embodiments of the present technology will be described with reference to the drawings.

[0012] <Definition of Terms> First, the definitions of the terms "pose" and "gesture" that are repeatedly mentioned in the embodiments will be explained. In this disclosure, a pose refers to a sustained posture of a model. Furthermore, a gesture refers to a specific movement, among the dynamic movements of a model, to which some meaning is attached. When a gesture of a model and a pose are associated, they may be referred to as a gesture pose or a gesture pose. For example, a sustained, stationary posture via a gesture of raising both hands may be considered a gesture pose or a gesture pose. It should be noted that the definitions of the terms "pose" and "gesture" are not limited to the definitions described here, but may include dictionary definitions and the contents of each embodiment.

[0013] First Embodiment [Configuration of Information Processing Apparatus] FIG. 1 is a block diagram showing an example configuration of an information processing apparatus 1 according to a first embodiment of the present technology.

[0014] As shown in FIG. 1, the information processing device 1 includes a learning dataset storage unit 2, an audio feature extraction unit 3, a risk assessment unit 4, a movement information processing unit 5, a movement feature extraction unit 6, an emotion recognition unit 7, a movement feature information storage unit 8, a teacher data construction unit 9, a learning unit 10, an audio data division unit 11, a movement feature selection unit 12, and an inference unit 13.

[0015] The information processing device 1 illustrated in Fig. 1 is configured as an embodiment of a learning processing device according to the present technology, and also as an embodiment of an inference processing device according to the present technology. That is, the information processing device 1 illustrated in Fig. 1 is configured to include both components for functioning as a learning processing device according to the present technology and components for functioning as an inference processing device according to the present technology.

[0016] In the example shown in Figure 1, the blocks above the dotted line indicating the boundary between "learning model generation" and "inference processing" in the figure are components that function as a learning processing device according to the present technology. Also, the blocks below the dotted line are components that function as a learning processing device according to the present technology. Note that the movement feature information storage unit 8 located above the dotted line is a component that functions as a learning processing device, and also a component that functions as a learning processing device.

[0017] The learning processing device, which is a learning-side device that generates a learning model, and the inference processing device, which is an inference-side device that performs inference processing, may be configured separately. When the learning-side device and the inference-side device are separate devices, the learning-side device and the inference-side device are configured to each have a movement feature information storage unit 8, and the movement feature information storage unit 8 of the inference-side device is configured to acquire learning parameters from the movement feature information storage unit 8 of the learning-side device.

[0018] The training dataset storage unit 2 includes a training data acquisition unit and a training data storage unit (not shown). The training data acquisition unit acquires voice data and movement data from an external device including a sensor. The training data storage unit stores the voice data and movement data acquired by the training data acquisition unit as training data. Hereinafter, movement data may be referred to as movement information. In this embodiment, the movement information is acquired in synchronization with the voice data and processed into a state where it is subdivided into data of an appropriate data size for training processing. The training dataset storage unit 2 stores a combination of voice data and movement information as a dataset.

[0019] The movement information includes movements of various parts of a person's body, such as the face, head, arms, hands, and legs. Note that the movement information is not limited to movements expressed in three-dimensional space, but may be expressed as two-dimensional movements projected onto a 2D image. Furthermore, the method of expressing the movement is not limited to a specific method, and a method of expressing time-series data of coordinate values ​​or a method of expressing the movement as time-series data of the angles of bones corresponding to joints may be adopted.

[0020] Furthermore, in this disclosure, human parts may be referred to as human body parts. Note that in this disclosure, "person" or "human body" referring to learning or inference processes is not limited to a human or a human body, but may also include human-like robots and virtual objects. Virtual objects may include animated characters that simulate humans and photorealistic digital humans.

[0021] The motion information is not limited to a specific body part or bone structure. Furthermore, the method for acquiring the motion information is not limited to a specific method, and for example, dedicated equipment or detection technology may be used. For example, motion capture may be performed while wearing a dedicated suit. Alternatively, motion information for learning may be extracted from existing 2D video content with audio or multi-viewpoint video content using existing image recognition technology.

[0022] The movement information also includes specific movements that express the speaker's intention or emphasize the conversation, such as movements of the person's hands (fingers) or the entire body. In this disclosure, specific movements that express the speaker's intention or emphasize the conversation may be referred to as beat gestures.

[0023] A rhythmic gesture refers to a gesture that corresponds to the rhythm of a conversation. For example, it includes gestures that express emphasis in speech, movements that keep rhythm with small, quick movements of the hands, and movements that move up, down, left, and right. In other words, a rhythmic gesture can be considered to be a gesture that is unrelated to the content of speech, or a gesture that does not change shape depending on the content of speech.

[0024] The audio data may be acquired (captured) by a sound collector such as a microphone. Furthermore, the audio data may be subjected to various audio signal processing such as noise addition and pitch adjustment for data augmentation. If the duration of the audio data changes, the length of the motion information is changed accordingly. Furthermore, if the learning unit 10 (described later) performs mini-batch learning, the audio data may be read in mini-batch size units (e.g., 32).

[0025] The speech feature extraction unit 3 extracts information indicating speech features based on the speech data 200 input from the training dataset storage unit 2. For example, the speech feature extraction unit 3 outputs, as speech feature information 301, any one of the pitch and energy of the speech signal, or feature quantities such as a Mel-frequency spectrogram and MFCC (Mel-frequency cepstral coefficients) that take human sound perception into consideration, or information combining these.

[0026] The risk determination unit 4 determines whether or not there is a high risk of artifact occurrence for the movement information 201 of the human body model (2D or 3D) input from the learning dataset storage unit 2. Specifically, the risk determination unit 4 determines whether or not the movement information 201 includes a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model. For example, with regard to a gesture in which a hand or finger is in contact with a part of the body, or a movement transitioning from one pose to another, the risk determination unit 4 determines that there is a high risk of artifact occurrence as including the above-mentioned specific state. Note that in the present disclosure, the "human body model" may also be simply referred to as a "model."

[0027] In this embodiment, artifacts include not only artifacts (contact and penetration) of each part of one model, such as the right and left hands, but also artifacts that occur between two models. For example, artifacts that occur due to a gesture performed by model A and a gesture performed by model B, such as clasping hands facing each other, are also included.

[0028] 2 is a block diagram showing an example of the configuration of the risk assessment unit 4. In the present disclosure, the risk assessment unit may be simply referred to as a assessment unit.

[0029] As shown in FIG. 2, the risk determination unit 4 includes a collision determination unit 41 , a pose transition determination unit 42 , and a risk information generation unit 43 .

[0030] The collision determination unit 41 performs collision determination based on the movement information 201. In this embodiment, the collision determination unit 41 sets colliders for parts of a person's gestures that are at high risk of artifact occurrence, and performs collision determination. Specifically, in the case of gestures that are mainly represented by hand movements, artifacts caused by fingers penetrating through the hands are likely to occur, so colliders are set for the hands (fingers).

[0031] The method for setting the collider is not limited, and a box-shaped or spherical collider may be used to set a collider that encompasses the hand, or, for more accurate collision detection, fine capsule-shaped colliders may be set to represent the fingers, first joints, etc. If even higher accuracy and processing speed are acceptable, collision detection may be performed using a hand mesh-based technique such as "FCL: A general purpose library for collision and proximity queries (Jia Pan et al. 2012)."

[0032] 3 is a graph showing collision determination, in which the vertical axis indicates the reliability of collisions and the horizontal axis indicates the number of collision occurrences.

[0033] As shown in FIG. 3, the reliability of a collision is calculated from the characteristics of the graph for the number of collisions included in the frames that make up the motion information 201, and is output as collision determination information 401 (for example, a value range of 0.0 to 1.0).

[0034] Note that the parts to which colliders are set are not limited, and may be applied only to specific parts such as hands, arms, etc. Colliders may also be set to multiple different parts, weighted values ​​may be calculated for each set part, and collision determination may be performed based on the calculated weights.

[0035] On the other hand, if the gesture pose transitions slowly over a relatively long period of time, adding a beat gesture movement makes it easier to perceive unnatural movement fluctuation artifacts. For this reason, the following determination is made by the pose transition determination unit 42.

[0036] FIG. 4 is a block diagram showing an example of the configuration of the pause transition determination unit 42. As shown in FIG.

[0037] The pause transition determination unit 42 determines whether or not a pause transition has occurred. In this embodiment, as shown in FIG. 4 , the pause transition determination unit 42 includes a time axis mid-range extraction unit 44, a time axis high-range extraction unit 45, and a determination unit 46.

[0038] The time axis mid-frequency extraction unit 44 and the time axis high-frequency extraction unit 45 use FIR filters with desired characteristics to extract mid-frequency components and high-frequency components in the time axis direction from the pause information of multiple frames indicated by the motion information 201. The time axis mid-frequency extraction unit 44 outputs the time axis mid-frequency components 402, and the time axis high-frequency extraction unit 45 outputs the time axis high-frequency components 403 to the determination unit 46.

[0039] The determination unit 46 multiplies the mid-frequency reliability MB_Rel calculated from the time axis mid-frequency component 402 according to the characteristics of the graph shown in Figure 5A by the high-frequency reliability HB_Rel calculated from the time axis high-frequency component 403 according to the characteristics of the graph shown in Figure 5B, and outputs the result (MB_Rel x HB_Rel) as pause transition determination information 404 (value range: 0.0 to 1.0).

[0040] Here, the pose information refers to information that indicates what pose the 3D model should assume when performing a predetermined gesture. For example, the poses include poses that express various emotions, such as a pose expressing joy, such as raising both hands, a pose expressing anger, such as clenching fists, a pose expressing fear, such as covering the face, a pose expressing surprise, such as covering the mouth, a pose expressing disgust, such as pointing at an object, or a pose expressing sadness, such as lowering both hands and hanging one's head.

[0041] The 3D model in this embodiment may be, for example, a 3DCG asset including a non-photorealistic model such as an anime character or a photorealistic model such as a digital human. Furthermore, the 3D model is not limited to a 3D model, and a 2D model may also be used.

[0042] The pose transition determination unit 42 may apply each process only to specific body parts (hands or arms), or may calculate a weighted value for each body part and perform the determination process.

[0043] 2 , the risk information generation unit 43 outputs the larger of the input collision determination information 401 and pose transition determination information 404 as risk information 405. In other words, the risk information may be considered as information indicating the possibility of contact between models due to a beat gesture or the occurrence of artifacts including unnatural movement fluctuations.

[0044] The movement information processing unit 5 performs processing to reduce the amount of movement components of the beat gesture when it is determined that the risk is high based on the movement information 201 and the risk information 405. In this embodiment, the movement information processing unit 5 attenuates the movement components of the beat gesture included in the movement information 201 in accordance with the value of the risk information 405.

[0045] Specifically, an encoder model that compresses the motion information into low-dimensional information and a decoder model that restores it to the original dimensions are separately learned and obtained in advance using VAE (Variational Autoencoder) technology, such as "Speech Drives Templates: Co-Speech Gesture Synthesis with Learned Templates" (ICCV2021), using all the motion information 201 included in the training dataset storage unit 2. In this embodiment, only the VAE decoder is used, and is provided as a VAE decoder unit 50, which will be described later.

[0046] Fig. 6A is a block diagram showing an example of the configuration of the movement information processing unit 5. Fig. 6B is a diagram showing an example of the characteristics of a graph for setting the attenuation coefficient α of the attenuation coefficient setting unit.

[0047] As shown in FIG. 6, the motion information processing unit 5 includes a VAE decoder unit 50 , an attenuation coefficient setting unit 51 , and an amplifier 52 .

[0048] The VAE decoder unit 50 receives low-dimensionally expressed movement feature information 601 from the movement feature extraction unit 6. This movement feature extraction unit 6 functions as a VAE encoder corresponding to the VAE decoder, as will be described later with reference to FIG. 7 . The VAE decoder unit 50 executes VAE decoding processing to restore the movement feature information 601 to its original dimensions. The output data of this decoder represents the general characteristics of the movement information 201, and the component obtained by subtracting this output data from the original movement information 201 may be considered as the movement component of the beat gesture.

[0049] In this embodiment, the data output from the VAE decoder unit 50 is set as a non-beat gesture movement component 501. The component obtained by subtracting the non-beat gesture movement component 501 from the movement information 201 is set as a beat gesture movement component 502.

[0050] The attenuation coefficient setting unit 51 sets the attenuation coefficient α from the risk information 405 in accordance with the characteristics of the graph shown in Fig. 6B. In this embodiment, as shown in Fig. 6B, the attenuation coefficient α is set to be larger as the risk of artifacts increases.

[0051] The amplifier 52 outputs a beat gesture attenuation movement component 503 calculated by multiplying the beat gesture movement component 502 by an attenuation coefficient α. Furthermore, the movement information processing unit 5 outputs processed movement information 504 by subtracting the movement information 201 from the output beat gesture attenuation movement component 503. In this manner, the higher the risk of artifacts, the more the beat gesture movement is attenuated, thereby suppressing artifacts and unnatural temporal fluctuations caused by the beat gesture. In the present disclosure, processed movement information may also be referred to as processed movement data.

[0052] The movement feature extraction unit 6 outputs movement feature information 601 indicating features such as rough changes in gestures in response to the movement information 201. The movement feature information 601 is stored in the movement feature information storage unit 8 for use in the inference process described below.

[0053] FIG. 7 is a block diagram showing an example of the configuration of the movement feature extraction unit 6. As shown in FIG.

[0054] As shown in FIG. 7, the motion feature extraction unit 6 has a VAE encoder unit 60 corresponding to the VAE decoder unit 50 of the motion information processing unit 5 .

[0055] The VAE encoder unit 60 expresses the rough characteristics of the input motion information 201 as low-dimensional information, and outputs this low-dimensional information as motion characteristic information 601. In this embodiment, the motion characteristic information 601 is output to the motion information processing unit 5, the motion characteristic information storage unit 8, and the teacher data configuration unit 9.

[0056] The emotion recognition unit 7 generates emotion recognition information 701 by applying various emotion recognition techniques to the input audio data 200 and movement information 201. The emotion recognition information 701 is linked to the movement characteristic information 601 and stored in the movement characteristic information storage unit 8.

[0057] FIG. 8 is a block diagram showing an example of the configuration of the emotion recognition unit 7.

[0058] As shown in FIG. 8, the emotion recognition unit 7 includes a voice emotion recognition unit 70 , a text conversion unit 71 , a text emotion recognition unit 72 , a motion emotion recognition unit 73 , and an emotion recognition integration unit 74 .

[0059] The voice emotion recognition unit 70 recognizes the emotions (joy, anger, fear, surprise, disgust, sadness, etc.) contained in the input voice data 200. For example, voice emotion recognition technology such as "Recognition of Emotions Included in Voice" (Journal of the Acoustical Society of Japan, 71, 484-489 (2015)) is used to output voice emotion estimation data 702.

[0060] The text conversion unit 71 converts the input voice data 200 into text information 703 by using a voice recognition technique such as speech-to-text.

[0061] The text emotion recognition unit 72 applies text-based emotion recognition technology (e.g., "Emotion Analysis of Japanese Sentences Using an Emotional Word Dictionary," Visualization Information Vokl41 No. 161 (2021)), etc. to the input text information 703, and outputs text emotion estimation data 704.

[0062] The motion emotion recognition unit 73 applies motion-based emotion recognition technology (e.g., "Emotion Recognition from Skeletal Movements," Entropy 29 June 2019) to the input movement information 201 to output motion emotion estimation data 705.

[0063] The emotion recognition integrating unit 74 outputs emotion recognition information 701 to the movement feature information storage unit 8 based on the emotion recognition results of the voice emotion estimation data 702, text emotion estimation data 704, and motion emotion estimation data 705. For example, when the emotion recognition results are composed of reliability (e.g., 0.0 to 1.0) of six types of emotions, namely joy, anger, fear, surprise, disgust, and sadness, the emotion recognition integrating unit 74 sets the emotion recognition information 701 as a weighted average of the voice emotion estimation data 702, text emotion estimation data 704, and motion emotion estimation data 705.

[0064] The movement feature information storage unit 8 stores movement feature information 601 and emotion recognition information 701. Here, the processing on the learning side of the movement feature information storage unit 8 (which refers to the learning dataset storage unit 2, the audio feature extraction unit 3, the risk determination unit 4, the movement information processing unit 5, the movement feature extraction unit 6, the emotion recognition unit 7, the movement feature information storage unit 8, the teacher data configuration unit 9, and the learning unit 10) will be described.

[0065] In this embodiment, the movement feature information storage unit 8 associates emotion recognition information 701 with movement feature information 601 and stores the information in various storage media with an index for identification assigned thereto. The movement feature information storage unit 8 is used as a condition input for the inference side (referring to the voice data division unit 11, the movement feature selection unit 12, and the inference unit 13). The condition input in this embodiment may be considered as input data representing the motion context or situation of the human body model during the inference process that drives the human body model in response to the voice input. Hereinafter, the condition input may be referred to as context information. The context information may be represented as a condition vector.

[0066] In the following, the learning-side processing performed by the learning dataset storage unit 2, the audio feature extraction unit 3, the risk assessment unit 4, the movement information processing unit 5, the movement feature extraction unit 6, the emotion recognition unit 7, the movement feature information storage unit 8, the teacher data configuration unit 9, and the learning unit 10 may be referred to as "learning time." Similarly, the inference-side processing performed by the audio data division unit 11, the movement feature selection unit 12, and the inference unit 13 may be referred to as "inference time." For example, when it is stated that certain data is used during inference, it means that the data is used when inference-side processing is performed. Note that the processing performed during learning corresponds to processing executed as an inference processing method related to the present technology, and the processing performed during inference corresponds to processing executed as an inference processing method related to the present technology.

[0067] The teacher data configuration unit 9 configures teacher data to be used for learning based on the audio feature information 301, the processed movement information 504, and the movement feature information 601.

[0068] In this embodiment, the teacher data construction unit 9 is configured to add not only voice data but also movement feature information representing pose changes according to gestures as Condition input, similar to "Speech Drives Templates: Co-Speech Gesture Synthesis with Learned Templates" (ICCV2021), and generate movement information representing gestures from that information.

[0069] When the teacher data construction unit 9 receives, as training data, movement feature information 601 that may have a high risk of artifacts during inference, it ignores the corresponding audio data to suppress the occurrence of artifacts during inference. More specifically, the teacher data construction unit 9 constructs the teacher data based on audio training data 903, processed movement training data 904, and movement feature training data 905, which will be described later. In the present disclosure, audio data that is substantially ignored in constructing the teacher data may be referred to as second audio data. Furthermore, teacher data including first audio data and second audio data may be referred to as first teacher data.

[0070] FIG. 9 is a block diagram showing an example of the configuration of the teacher data configuration unit 9.

[0071] As shown in FIG. 9, the teacher data configuration unit 9 includes a voice feature clustering unit 90 , a heterogeneous voice feature data selection unit 91 , and a combination data generation unit 92 .

[0072] The teacher data construction unit 9 acquires audio feature information 301, risk information 405, processed motion information 504, and motion feature information 601. The audio feature information 301, risk information 405, processed motion information 504, and motion feature information 601 are each segment data (hereinafter referred to as learning segment data) corresponding to a certain period of time (e.g., 4 seconds), and are all synchronized in chronological order. Each synchronized data is assigned a corresponding identification index (hereinafter referred to as segment identification index).

[0073] The speech feature clustering unit 90 performs cluster analysis on all segment data of the speech feature information 301, where the total number of data is N, and groups data having the same properties. For example, the cluster analysis may be performed using the K-means method. The speech feature clustering unit 90 also outputs cluster group information 901 to which each piece of speech feature information 301 belongs to, to the heterogeneous speech feature data selection unit 91.

[0074] The heterogeneous speech feature data selection unit 91 selects data having different properties (features) for each piece of speech feature information 301. For example, the heterogeneous speech feature data selection unit 91 selects multiple cluster groups (hereinafter referred to as heterogeneous cluster groups) that are distant from the cluster center of gravity of the cluster group to which each piece of speech feature information 301 belongs. Furthermore, the heterogeneous speech feature data selection unit 91 randomly selects multiple pieces of data from each heterogeneous cluster group and outputs heterogeneous speech feature data 902. In the present disclosure, speech data having different characteristics may be distinguishably expressed as first speech data and second speech data.

[0075] The combination data generation unit 92 constructs audio training data 903, processed motion training data 904, and motion feature training data 905 as training data based on the heterogeneous audio feature data 902, audio feature information 301, risk information 405, processed motion information 504, and motion feature information 601.

[0076] FIG. 10 is a flowchart showing the process of the combination data generating unit 92.

[0077] As shown in Figure 10, for all learning segment data of audio feature information 301, processed movement information 504, and movement feature information 601 to which a segment identification index is attached, it is determined whether the value of the corresponding risk information 405 is greater than or equal to a threshold value Th (step 101).

[0078] If the value of the risk information 405 is equal to or greater than the threshold Th (YES in step 101), the teacher data construction unit 9 determines that the multiple pieces of heterogeneous audio feature data 902 are training data with a high risk of artifacts during inference, and adds them as teacher data for the audio feature information 301 (step 102). At this time, the teacher data construction unit 9 associates the multiple pieces of input movement feature information 601 and the corresponding multiple pieces of audio feature information 301 with one correct answer data (corresponding to the output data during inference) of the processed movement information 504 in a many-to-one relationship. In other words, the teacher data construction unit 9 assigns one piece of processed movement information 504 to multiple pieces of audio data with various properties.

[0079] As described above, in this embodiment, a plurality of pieces of heterogeneous audio feature data 902 are learned in a many-to-one relationship with the processed motion information 504. Therefore, in the inference process for driving a human body model in response to audio input, if the motion feature information satisfies a high-risk condition, substantially a single piece of processed motion information 504 is always output in response to input of audio data with different features.

[0080] In this embodiment, training data obtained by adding a plurality of heterogeneous audio feature data 902 as training data for the audio feature information 301 in addition to the audio feature information 301, processed motion information 504, and motion feature information 601 may be referred to as first training data. In the first training data, a plurality of heterogeneous audio feature data 902 is associated with one piece of processed motion information 504.

[0081] If the risk information value is less than the threshold Th (NO in step 101), the foreign audio feature data 902 is not added, and training data is set using only the normal audio feature information 301 (step 103). That is, the audio feature information 301 is learned one-to-one with respect to the processed movement information 504 corresponding to the audio feature information 301. In this embodiment, training data in which the audio feature information 301 and the processed movement information 504 are set in a one-to-one relationship may be referred to as second training data. Here, the second training data may be considered to further include movement feature information 601 in addition to the audio feature information 301 and the processed movement information 504, which have a one-to-one relationship. Note that, because the risk information value is less than the threshold Th, attenuation is not applied to the rhythm gestures related to the second training data, unlike the rhythm gestures related to the first training data.

[0082] The teacher data construction unit 9 repeats the processes from steps 101 to 103 up to the total number of data N (steps 104 and 105). The teacher data construction unit 9 outputs the voice training data 903, the processed movement training data 904, and the movement feature training data 905 set by the above steps to the learning unit 10 as teacher data.

[0083] When the value of the risk information 405 is equal to or greater than the threshold Th and a plurality of pieces of heterogeneous audio feature data 902 are added as training data for the audio feature information 301, the training data composing unit 9 may change the number of pieces of heterogeneous audio feature data 902 to be added according to the value of the risk information 405, as shown in the graph in Fig. 11. As a result, the higher the risk of artifacts in the movement feature information 601, the more likely it is that the input audio data 200 will be ignored.

[0084] The learning unit 10 calculates the weights of the learning network from the voice learning data 903 and the movement feature learning data 905 so as to generate a movement that is close to the processed movement learning data 904 .

[0085] In this embodiment, the learning unit 10 performs supervised learning as a regression problem of converting the voice training data 903 and the movement feature training data 905 into the processed movement training data 904. For example, as the network architecture, a convolutional neural network (CNN) such as U-Net or a recurrent neural network (RNN) such as LSTM may be used, or other configurations may be used.

[0086] When the learning unit 10 performs mini-batch learning, the speech training data 903, processed motion training data 904, and motion feature training data 905 are output to the learning unit 10 in mini-batch size units (e.g., 32). In this embodiment, in order to perform the learning operation of the above-mentioned many-to-one relationship, not only is the number of input data simply increased, but multiple pieces of information with properties different from the input speech feature information 301 are efficiently sampled, thereby significantly reducing the learning time.

[0087] The loss function for the time-series data of the processed motion learning data 904 is calculated and set as a weighted average using both or either of intra-frame information at a certain point in time and inter-frame information at different points in time. The loss function may be set by taking into account a generative adversarial network (GAN).

[0088] Through these learning processes, learning model parameters 1001 of a neural network capable of generating human movements from input voice are generated and output to the inference unit 13 .

[0089] The audio data dividing unit 11 divides audio test data 202 captured by a sound collector such as a microphone in a manner that takes into consideration the timing of gesture switching, and outputs the divided audio test data 1101 .

[0090] The voice data segmentation unit 11 extracts gesture switching timings from the voice test data 202. Specifically, the voice data segmentation unit 11 segments the voice data using existing voice recognition speech segmentation technology (e.g., "Statistical Language Model for Speech Segmentation in the Speech Recognition Process," Nakajima et al., Transactions of the Information Processing Society of Japan, Vol. 42 (11), pp. 2681-2688, 2001-11-15), and measures and accumulates the speech time for each utterance. Furthermore, the unit 11 performs morphological analysis of natural language processing to extract the timings at which keywords corresponding to conjunctions are spoken and the timing at which silence ends when the duration of a silent section exceeds a threshold (see reference numeral 1102 in FIG. 12 ).

[0091] In addition, the audio data division unit 11 divides the audio test data 202 when one of the following conditions is met: when the cumulative speech time exceeds a lower threshold and just before a conjunction is uttered (see reference number 1103 in Figure 12), when a silent period exceeds a predetermined threshold (see reference number 1102 in Figure 12) and the silent period ends (see reference number 1104 in Figure 12), or when the cumulative speech time exceeds an upper threshold, and outputs the divided audio test data 202 to the movement feature selection unit 12 and the inference unit 13 as divided audio test data 1101.

[0092] FIG. 13 is a block diagram showing an example of the configuration of the movement feature selection unit 12. As shown in FIG.

[0093] As shown in FIG. 13 , the movement feature selection unit 12 includes an emotion recognition unit 120 , a selected movement feature information storage unit 121 , and a selection information setting unit 122 .

[0094] The emotion recognition unit 120 performs emotion recognition processing on the input divided audio test data 1101, and outputs inferred emotion recognition information 1201 to the selection information setting unit 122. For example, the emotion recognition processing may be performed using a configuration similar to that of the emotion recognition unit 7, but excluding the motion emotion recognition unit 73. Note that the emotion recognition technology is not limited, and may be the same as that used on the learning side, or a different method.

[0095] The selected movement characteristic information storage unit 121 stores the selected movement characteristic information 1204 output to the inference unit 13 described later in memory, and outputs the selected movement characteristic information 1202 to the selection information setting unit 122 as past selected movement characteristic information for the current input.

[0096] The selection information setting unit 122 integrates the inferred emotion recognition information 1201 and the selected movement characteristics information 1202 and outputs the result as movement selection information 1203 to the movement characteristics information storage unit 8 .

[0097] Here, we will explain the operation of the movement feature information storage unit 8 during inference. During inference, movement selection information 1203 is input to the movement feature information storage unit 8, and the inferred emotion recognition information 1201 and selected movement feature information 1202 included in that data are used, and the characteristics of the movement feature information 601 and emotion recognition information 701 stored during learning are taken into consideration to select from multiple pieces of data.

[0098] Specifically, the similarity Edist is calculated between each piece of stored emotion recognition information 701 and the inferred emotion recognition information 1201. For example, if the emotion recognition result is expressed as a six-dimensional vector with elements representing the reliability (e.g., 0.0 to 1.0) of six types of emotions, namely, joy, anger, fear, surprise, disgust, and sadness, the distance between the vectors (e.g., L1 norm, Euclidean distance, cosine similarity, etc.) is calculated.

[0099] Furthermore, the similarity Mdist between the movement feature information 601 linked to each emotion recognition information 701 and the selected movement feature information 1202 is calculated. When Mdist is expressed as a vector, it is defined as the distance between the vectors, similar to Edist. A cost value (Cost = Edist + λ × Mdist, λ is a positive coefficient) is calculated from the two similarities Edist and Mdist, and the movement feature information 601 corresponding to the smallest cost value is selected and output to the movement feature selection unit 12 and the inference unit 13 as selected movement feature information 1204.

[0100] The larger the λ in the cost calculation formula, the more movement feature information that takes into account the continuity of the gesture is selected. Conversely, when λ is small, the gesture is more likely to reflect the emotion recognition results.

[0101] As a result, selected movement feature information 1204 suitable for expressing the emotion of the divided audio test data 1101 input at the time of inference is selected so that no temporal discontinuity occurs.

[0102] The inference unit 13 sets the weight of the inference network based on the learning model parameters 1001 , and outputs inferred motion information 1301 when the divided audio test data 1101 and selected motion feature information 1204 are input.

[0103] In this embodiment, the inference unit 13 has a network architecture corresponding to the learning unit 10. The parameters of the inference network are set by the learning model parameters 1001 output from the learning unit 10, and the inference network is input with the divided audio test data 1101 and the selected movement feature information 1204 as the condition input, whereby the inference network outputs the inferred movement information 1301.

[0104] The inferred emotion recognition information 1201 obtained by performing emotion recognition processing on the input divided audio test data 1101 can also be considered inferred emotion information indicating the emotion associated with the input audio data. Furthermore, the emotion recognition information 701 stored during learning can also be considered learned emotion information used when learning the inference network. Furthermore, the similarity Edist can also be considered the similarity between the inferred emotion information and the learned emotion information. In this embodiment, the inference unit 13 can select inferred movement information 1301 as movement feature information corresponding to the gesture to be performed by the human body model, based on the similarity between the inferred emotion information and the learned emotion information.

[0105] As described above, the information processing device 1 according to this embodiment determines the possibility of contact between models or unnatural fluctuation of movement according to a beat gesture based on the motion data of the models. Based on the determination result, first training data including multiple pieces of heterogeneous audio feature data 902 for a single piece of motion data and second training data including a single piece of audio feature information 301 for a single piece of motion data are generated. This makes it possible to improve the accuracy of generating natural gestures.

[0106] Conventionally, when changing the movement of a human or other model based on audio, artifacts such as finger penetration or unnatural temporal fluctuations that react to the volume and intonation of the audio when transitioning from one pose to another can occur. On the other hand, when removing these artifacts through post-processing using CG software, etc., finger penetration and other artifacts have been addressed using the IK (Inverse Kinematics) function, and unnatural temporal fluctuations have been addressed using time smoothing processing.

[0107] This technology uses machine learning to generate human gesture movements using audio data as input. It determines whether the training data includes movements involving hand contact or pose transitions, and performs training that ignores the input audio data. By simply replacing the trained model during inference, it is possible to generate movements free of artifacts such as finger penetration and unnatural shaking during pose transitions, as described above.

[0108] More specifically, by suppressing the amount of movement components of rhythmic gestures that are highly correlated with audio data, and by configuring the training data so that learning operations ignore paired audio data, it becomes possible to suppress the occurrence of artifacts during inference.

[0109] Furthermore, when generating human movements from voice data using deep learning, it is possible to efficiently perform learning by determining whether the training data includes movements involving hand contact or pose transitions, and then ignoring the input voice data, or by using a number of training data samples with different characteristics determined according to the risk of artifacts.Furthermore, during inference, simply by replacing the trained model, it becomes possible to generate gesture movements free of artifacts such as the above-mentioned finger penetration and unnatural shaking during pose transitions.

[0110] Second Embodiment An information processing device according to a second embodiment of the present technology will be described. In the following description, descriptions of parts having the same configuration and operation as those of the information processing device 1 described in the above embodiment will be omitted or simplified.

[0111] FIG. 14 is a block diagram showing an example of the configuration of an information processing device 1B according to the second embodiment of the present technology.

[0112] 14, the second embodiment differs from the above-described embodiments in the processing content of the motion information processing unit 5 and in that motion feature information 601 is not input from the motion feature extraction unit 6 to the motion information processing unit 5. Hereinafter, the motion information processing unit 5 according to the second embodiment will be referred to as motion information processing unit 5B.

[0113] Fig. 15 is a block diagram showing an example of the configuration of the movement information processing unit 5B. Fig. 16 is a schematic diagram in which movement information representing a gesture is likened to a waveform.

[0114] As shown in FIG. 15, the movement information processing unit 5B includes a time smoothing processing unit 53, an attenuation coefficient setting unit 51, and an amplifier 52.

[0115] In the first embodiment, the VAE encoder 60 is used in the motion feature extractor 6 to represent the rough features of the motion information 201 as low-dimensional information. Therefore, the corresponding processing is applied to the motion information processor 5.

[0116] Furthermore, the purpose of the movement feature extraction unit 6 is to extract features that can be distinguished from other data as condition inputs for the network, and it is not essential to associate it with the process of removing and reducing the movement components of beat gestures contained in the movement information 201 executed by the movement information processing unit 5. Using the VAE encoder unit 60 has another effect of reducing the capacity required to store data in the movement feature information storage unit 8, the data size of the network of the learning unit 10 and the inference unit 13, and the calculation cost by compressing the data into low-dimensional information.

[0117] In the second embodiment, the time smoothing processing unit 53 removes and reduces the movement components of the beat gesture included in the movement information 201 .

[0118] Specifically, the time smoothing processor 53 applies processing based on a time-direction FIR low-pass filter that does not generate time phase distortion. As shown in Fig. 16, the small-amplitude time-varying movement surrounded by the dotted line 15 can be considered to be the movement component of a beat gesture. If a simple time-direction FIR low-pass filter is applied, as shown in graph 16 in Fig. 16, when the gesture pose suddenly changes significantly (for example, dotted line 17 in Fig. 16), such movement will be dulled, and even the original movement that is intended to be expressed will be distorted.

[0119] Therefore, in the second embodiment, an edge-preserving smoothing filter for an image is applied to one-dimensional data in the time direction. FIG. 17 is a block diagram showing an example of the configuration of the time smoothing processing unit 53.

[0120] As shown in FIG. 17, the time smoothing processing unit 53 includes a time FIR low-pass filter 54 , a time edge reliability calculation unit 55 , and a blending unit 56 .

[0121] The temporal FIR low-pass filter 54 applies an FIR low-pass filter to the input movement information 201 corresponding to each body part, and outputs temporally smoothed movement information 505. Note that the frequency characteristics of the filter may be changed according to the properties of the movement components of the rhythmic gestures that make up the movement information 201, and multiple low-pass filters may be applied by switching between them.

[0122] The temporal edge reliability calculation unit 55 determines whether an abrupt change in pause has occurred, as indicated by the dotted line portion 17 in Fig. 16. For example, the temporal edge reliability calculation unit 55 makes the determination based on the magnitude of the adjacent time difference or the like, and outputs a temporal edge reliability 506, which is an occurrence reliability indicating whether an abrupt change in pause has occurred.

[0123] The blending unit 56 blends the original motion information 201 and the time-smoothed motion information 505 so that the greater the time edge reliability 506, the greater the weighting of the motion information 201. The non-beat gesture motion component 507, which is the calculation result, is output from the time-smoothing processing unit 53.

[0124] For example, the non-beat gesture movement component 507 can be calculated by the following formula: Note that, hereinafter, the time edge reliability 506 is represented by Trel (value range: 0.0 to 1.0).

[0125] Non-beat gesture movement component 507 = Trel × movement information 201 + (1.0 − Trel) × time-smoothed movement information 505

[0126] In this way, by introducing the time smoothing processing unit 53 into the movement information processing unit 5B, it is possible to remove and reduce the movement components of rhythmic gestures with high precision without degrading the movements of non-rhythmic gestures contained in the movement information 201.

[0127] Third Embodiment An information processing device according to a third embodiment of the present technology will be described below. Fig. 18 is a block diagram showing an example of the configuration of an information processing device 1C according to the third embodiment of the present technology.

[0128] 18 , in the third embodiment, the emotion recognition unit 7 is replaced with a text conversion unit 7C, and the processing contents of the movement characteristic information storage unit 8 and the movement characteristic selection unit 12 are different from those in the first embodiment. Hereinafter, the movement characteristic information storage unit 8 and the movement characteristic selection unit 12 according to the third embodiment will be referred to as the movement characteristic information storage unit 8C and the movement characteristic selection unit 12C.

[0129] The text conversion unit 7C uses only the text conversion unit 71 of the emotion recognition unit 7 in the first embodiment, and converts the audio data 200 into text information 701C by using a speech recognition technique such as Speech-to-text. The text information 701C is output to the movement feature information storage unit 8C.

[0130] In the movement characteristic information storage unit 8C, the text information 701C is linked to the movement characteristic information 601, and is stored in various storage media with an index for identification assigned thereto.

[0131] FIG. 19 is a block diagram showing an example of the configuration of the movement feature selection unit 12C.

[0132] As shown in FIG. 19, the movement characteristic selecting unit 12C includes a text converting unit 123, a selected movement characteristic information storage unit 121, and a selection information setting unit 122C.

[0133] The text conversion unit 123 converts the audio test data 202 into inference text information 1205 by using a voice recognition technique such as speech-to-text on the audio test data 202. The inference text information 1205 is output to the selection information setting unit 122C.

[0134] As in the first embodiment, the selected movement characteristic information storage unit 121 stores the selected movement characteristic information 1204 output to the inference unit 13 in memory, and outputs the selected movement characteristic information 1202 to the selection information setting unit 122C as past selected movement characteristic information for the current input.

[0135] The selection information setting unit 122C integrates the inference text information 1205 and the selected movement characteristic information 1202, and outputs the result as movement selection information 1203C to the movement characteristic information storage unit 8C.

[0136] The operation of the movement characteristic information storage unit 8C during inference in the third embodiment will be described below. During inference, movement selection information 1203C is input to the movement characteristic information storage unit 8C, and inference text information 1205 and selected movement characteristic information 1202 included in the data are used, and the characteristics of text information 701C and movement characteristic information 601 stored during learning are taken into consideration to select from a plurality of pieces of data.

[0137] Specifically, the text similarity Tdist is calculated between the dialogue of each piece of stored text information 701C and the inference text information 1205. For example, "Sentence-BERT: Sentence Embeddings using Siamese BERT Networks (2019)" can be used as a technique for calculating the text similarity.

[0138] Furthermore, similarly to the first embodiment, the similarity Mdist between the movement characteristic information 601 linked to each piece of text information 701C and the selected movement characteristic information 1202 is calculated.

[0139] A cost value (Cost = Tdist + λ × Mdist, λ is a positive coefficient) is calculated from these two similarities, and the motion feature information 601 corresponding to the smallest cost value is selected and output to the motion feature selection unit 12C and the inference unit 13 as selected motion feature information 1204.

[0140] In this way, the gesture poses included in the learning data when a speech similar to the lines of the divided voice test data 1101 input at the time of inference is expressed.

[0141] Other Embodiments The present technology is not limited to the above-described embodiments, and various other embodiments can be realized.

[0142] In the above embodiment, the movement information 201 is extracted into mid-frequency components and high-frequency components in the time axis direction by the time axis mid-frequency extraction unit 44 and the time axis high-frequency extraction unit 45 provided in the pose transition determination unit 42. However, the present invention is not limited to this, and the pose transition determination information 404 may be output from the speed and acceleration of each part that indicate the movement of a person.

[0143] FIG. 20 is a block diagram showing another example of the configuration of the pause transition determination unit 42.

[0144] As shown in FIG. 20, the pause transition determination unit 42 includes a speed measurement unit 47, an acceleration measurement unit 48, and a determination unit 46.

[0145] The speed measurement unit 47 and the acceleration measurement unit 48 measure the speed and acceleration of each part indicating the movement of the person for the pose information of multiple frames indicated by the movement information 201. The determination unit 46 multiplies the speed reliability V_Rel calculated according to the characteristics of the graph shown in Fig. 21A by the acceleration reliability A_Rel calculated according to the characteristics of the graph shown in Fig. 21B using different thresholds, and outputs the result (V_Rel x A_Rel) as pose transition determination information 404. Note that each process may be applied only to a specific part (hand or arm), or a weighted value may be calculated for each part and the determination process may be performed.

[0146] In the above embodiment, the VAE technique is used in the motion information processing unit 5. However, the present invention is not limited to this, and principal component analysis (hereinafter referred to as PCA) may also be used.

[0147] FIG. 22 is a block diagram showing another example of the configuration of the movement information processing unit 5. In FIG.

[0148] As shown in FIG. 22, the motion information processing unit 5 includes a PCA inverse transform unit 57 , an attenuation coefficient setting unit 51 , and an amplifier 52 .

[0149] Even when PCA is used, the rough characteristics of the motion information 201 can be expressed as low-dimensional information, and therefore it can be used for the same purpose.

[0150] Furthermore, when the motion information processing unit 5 is provided with the PCA inverse transform unit 57, the motion feature extraction unit 6 may be provided with a PCA forward transform unit 61 corresponding to the PCA inverse transform unit 57 (see FIG. 23 ). In this case, too, as in the first embodiment, the input motion information 201 is expressed as low-dimensional information representing rough features, and this low-dimensional information is output as motion feature information 601.

[0151] The configurations of the risk assessment unit, teacher data configuration unit, learning unit, etc. described with reference to the drawings are merely one embodiment and can be modified as desired without departing from the spirit of the present technology. In other words, any other configurations, algorithms, etc. for implementing the present technology may be adopted.

[0152] It should be noted that the effects described in this disclosure are merely examples and are not limiting, and other effects may also be present. The description of multiple effects above does not necessarily mean that these effects are exhibited simultaneously. It means that at least one of the effects described above can be obtained depending on the conditions, etc., and of course, effects not described in this disclosure may also be exhibited.

[0153] It is also possible to combine at least two of the characteristic features of each embodiment described above. In other words, the various characteristic features described in each embodiment may be combined in any manner without distinguishing between the embodiments.

[0154] The present technology may also be configured as follows: (1) A learning processing device comprising: a learning dataset storage unit that acquires learning data including motion data of at least one human body model and audio data synchronized with the motion data; a determination unit that determines whether the motion data includes a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model; an audio feature extraction unit that classifies the audio data into at least first audio data and second audio data having features different from the first audio data; a teacher data configuration unit that generates first teacher data including the first audio data and the second audio data if the motion data includes the specific state, and generates second teacher data including the first audio data but not the second audio data if the motion data does not include the specific state; and a learning unit that generates learning model parameters for generating rhythmic gestures corresponding to the rhythm of conversation represented by input audio data, based on the first teacher data and the second teacher data. (2) The learning processing device according to (1), further comprising a processing unit that generates processed movement data in which the rhythmic gesture is attenuated when the movement data includes the specific state. (3) The learning processing device according to (2), wherein in the first teacher data, a single piece of processed movement data corresponds many-to-one to a plurality of pieces of audio data including the first audio data and the second audio data, and in the second teacher data, a single piece of processed movement data corresponds one-to-one to the first audio data and the movement data representing the rhythmic gesture to which attenuation has not been applied. (4) The learning processing device according to (2), further comprising an extraction unit that reduces the dimension of the movement data and extracts feature information indicative of a feature of a gesture, and the processing unit generates the processed movement data based on the reduced-dimensional movement data. (5) The learning processing device according to (4), further comprising an emotion recognition unit that recognizes emotion information related to emotions associated with the audio data and the movement data based on the audio data and the movement data, and a storage unit that stores the emotion information in association with the feature information.(6) The learning processing device according to (1), wherein the determination unit extracts mid-frequency components and high-frequency components of the motion data in a time axis direction and determines whether the motion data includes the specific state based on the mid-frequency components and the high-frequency components. (7) A learning processing method including: acquiring training data including motion data of at least one human body model and audio data synchronized with the motion data, determining whether the motion data includes a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model, classifying the audio data into at least first audio data and second audio data having characteristics different from the first audio data, generating first teacher data including the first audio data and the second audio data if the motion data includes the specific state, generating second teacher data including the first audio data but not the second audio data if the motion data does not include the specific state, and generating, based on the first teacher data and the second teacher data, learning model parameters for generating rhythmic gestures corresponding to the rhythm of conversation represented by input audio data. (8) An inference processing device comprising: a voice data acquisition unit that acquires input voice data; and an inference network that acquires an inference network for generating a rhythmic gesture of at least one human body model corresponding to the rhythm of a conversation represented by the input voice data; and acquires learning model parameters of the inference network that correspond to first teacher data representing a plurality of voice data having different characteristics and second teacher data representing voice data having a single characteristic; and an inference unit that, based on the input voice data, the inference network, and the learning model parameters, generates a non-beat gesture synchronized with the input voice data when the input voice data corresponds to the first teacher data, and generates the beat gesture synchronized with the input voice data when the input voice data corresponds to the second teacher data, wherein the non-beat gesture corresponds to a specific state including at least one of contact between a plurality of parts of the human body model or a pose transition of the human body model.(9) The inference processing device according to (8), wherein the inference unit selects movement feature information corresponding to a gesture to be performed by the human body model based on a similarity between inferred emotion information indicating an emotion associated with the input voice data and learned emotion information used when training the inference network. (10) An inference processing method including: acquiring input speech data; acquiring an inference network for generating at least one rhythmic gesture of a human body model corresponding to the rhythm of a conversation represented by the input speech data; acquiring learning model parameters of the inference network corresponding to first teacher data representing a plurality of speech data having different characteristics and second teacher data representing speech data having a single characteristic; and generating, based on the input speech data, the inference network, and the learning model parameters, a non-rhythmic gesture synchronized with the input speech data when the input speech data corresponds to the first teacher data, and generating the rhythmic gesture synchronized with the input speech data when the input speech data corresponds to the second teacher data, wherein the non-rhythmic gesture corresponds to a specific state including at least one of contact between a plurality of parts of the human body model or a pose transition of the human body model.

[0155] DESCRIPTION OF SYMBOLS 1... Information processing device 4... Risk determination unit 5... Movement information processing unit 6... Movement feature extraction unit 8... Movement feature information storage unit 9... Teacher data construction unit 10... Learning unit 12... Movement feature selection unit 13... Inference unit

Claims

1. A learning processing device having: a learning dataset storage unit that acquires learning data including movement data of at least one human body model and audio data synchronized with the movement data; a determination unit that determines whether the movement data includes a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model; an audio feature extraction unit that classifies the audio data into at least first audio data and second audio data having characteristics different from the first audio data; a teacher data construction unit that generates first teacher data including the first audio data and the second audio data if the movement data includes the specific state, and generates second teacher data including the first audio data but not the second audio data if the movement data does not include the specific state; and a learning unit that generates learning model parameters for generating rhythmic gestures corresponding to the rhythm of conversation represented by input audio data based on the first teacher data and the second teacher data.

2. The learning processing device according to claim 1, further comprising a processing unit that generates processed movement data in which the beat gesture is attenuated when the movement data includes the specific state.

3. The learning processing device described in claim 2, wherein in the first training data, there is a many-to-one correspondence between a plurality of pieces of audio data including the first audio data and the second audio data and a single piece of processed movement data, and in the second training data, there is a one-to-one correspondence between the first audio data and the movement data representing the beat gesture to which no attenuation has been applied.

4. The learning processing device according to claim 2, further comprising an extraction unit that reduces the dimension of the motion data and extracts feature information indicating gesture features, and the processing unit generates the processed motion data based on the reduced-dimensional motion data.

5. The learning processing device according to claim 4, further comprising: an emotion recognition unit that recognizes emotion information relating to emotions associated with the voice data and the movement data based on the voice data and the movement data; and a memory unit that stores the emotion information in association with the feature information.

6. The learning processing device according to claim 1, wherein the determination unit extracts mid-range and high-range components of the motion data in the time axis direction, and determines whether the motion data includes the specific state based on the mid-range and high-range components.

7. A learning processing method comprising: acquiring learning data including motion data of at least one human body model and audio data synchronized with the motion data; determining whether the motion data includes a specific state including at least one of contact between multiple parts of the human body model or a pose transition of the human body model; classifying the audio data into at least first audio data and second audio data having characteristics different from the first audio data; generating first training data including the first audio data and the second audio data if the motion data includes the specific state; generating second training data including the first audio data but not the second audio data if the motion data does not include the specific state; and generating learning model parameters for generating rhythmic gestures corresponding to the rhythm of conversation represented by input audio data based on the first training data and the second training data.

8. An inference processing device comprising: a voice data acquisition unit that acquires input voice data; and an inference network that acquires an inference network for generating a rhythmic gesture of at least one human body model corresponding to the rhythm of a conversation represented by the input voice data; and acquires learning model parameters of the inference network that correspond to first teacher data representing a plurality of voice data having different characteristics and second teacher data representing voice data having a single characteristic; and an inference unit that, based on the input voice data, the inference network, and the learning model parameters, generates a non-beat gesture synchronized with the input voice data if the input voice data corresponds to the first teacher data, and generates the beat gesture synchronized with the input voice data if the input voice data corresponds to the second teacher data, wherein the non-beat gesture corresponds to a specific state including at least one of contact between a plurality of parts of the human body model or a pose transition of the human body model.

9. The inference processing device according to claim 8, wherein the inference unit selects movement feature information corresponding to a gesture to be performed by the human body model based on the similarity between inferred emotion information indicating the emotion associated with the input voice data and learned emotion information used when training the inference network.

10. An inference processing method comprising: acquiring input speech data; acquiring an inference network for generating rhythmic gestures of at least one human body model corresponding to the rhythm of conversation represented by the input speech data; acquiring learning model parameters of the inference network corresponding to first teacher data representing a plurality of speech data having different characteristics and second teacher data representing speech data having a single characteristic; and generating, based on the input speech data, the inference network, and the learning model parameters, a non-rhythmic gesture synchronized with the input speech data if the input speech data corresponds to the first teacher data, and generating the rhythmic gesture synchronized with the input speech data if the input speech data corresponds to the second teacher data, wherein the non-rhythmic gesture corresponds to a specific state including at least one of contact between a plurality of parts of the human body model or a pose transition of the human body model.

Citation Information

Patent Citations

  • Android gesture generating device and computer program

    JP2020006482A

  • Posture data generation device, learning tool, computer program, learning data, posture data generation method and learning model generation method

    JP2020082246A

  • Learning model generation device, inference processing device, learning model generation method, and inference processing method

    WO2024014318A1