Information processing apparatus, information processing method, and recording medium

By segmenting speech data into high-frequency and low-frequency components and correcting gesture information, the one-to-many problem of gesture generation in existing technologies is solved, enabling diversified expression of gestures and improving control accuracy, simplifying network configuration and reducing learning time.

CN120883249APending Publication Date: 2025-10-31SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480019774.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-27
Filing Date
2024-02-13
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies suffer from a one-to-many problem when generating gestures for parts other than the face based on speech. This makes it difficult to achieve diverse expressions and control accuracy of gestures. Furthermore, the large amount of learning data required and the complex network configuration lead to inaccurate generated actions.

Method used

Gesture information is inferred by inputting speech data into a learned model and segmenting it into high-frequency and low-frequency components. These components are then corrected using correction parameters to generate corrected gesture information, which is then combined with the speech data to execute the gesture synchronously.

Benefits of technology

It enables more diverse gesture expressions and improves control accuracy, simplifies network configuration, and reduces learning time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120883249A_ABST
    Figure CN120883249A_ABST
Patent Text Reader

Abstract

This information processing device is equipped with a dividing unit and a correcting unit. The inferring unit inferres gesture information related to a gesture by inputting voice data into a training model. The dividing unit divides the gesture information into a high-frequency component representing a beat gesture corresponding to the dialogue rhythm and a low-frequency component representing a gesture different from the beat gesture. The correction unit generates corrected gesture information by correcting at least one of the high-frequency component and the low-frequency component based on the correction parameter. This enables multiple gesture expressions and improved gesture control accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to information processing devices, information processing methods, and recording media that can be applied to generate animations such as human movements. Background Technology

[0002] Conventionally, techniques for altering mouth movements based on speech are known (see Patent Document 1). However, there are problems: when attempting to generate movements of parts other than the face based on speech, different movements are generated based on the same spoken speech, which is known as a one-to-many problem, and the output of the inference processing may be the average movement of the learning data. Furthermore, as shown in Table 2 of Patent Document 1, classification methods using iconic, metaphorical, beat, deictic, and emmblem as gesture classification labels are known.

[0003] Furthermore, similar to Non-Patent Document 1, multiple actions are generated by designing datasets and configuring networks. However, Non-Patent Document 1 requires a large amount of training data, has a more complex network configuration, and consumes a significant amount of time for learning. Moreover, even if multiple actions can be generated, the desired actions cannot always be obtained, and it is difficult to improve the control accuracy of the generated action data.

[0004] Citation List

[0005] Patent documents

[0006] Patent Document 1: Japanese Patent Application Publication No. 2020-6482

[0007] Non-patent literature

[0008] Non-patent literature 1: "Speech 6Drives Templates: Co-Speech Gesture Synthesis with Learned Templates" (ICCV2021) Summary of the Invention

[0009] Technical issues

[0010] Therefore, it is desirable to provide an information processing device, an information processing method, and a recording medium describing a program, which can achieve diversified expression of gestures and improved control accuracy of gestures when generating human actions based on speech.

[0011] In view of the above, the purpose of this technology is to provide an information processing device, an information processing method, and a recording medium that can realize diversified expression of gestures and improve the accuracy of gesture control.

[0012] Solution to the problem

[0013] To achieve the aforementioned objectives, the information processing apparatus according to embodiments of the present technology includes an inference unit, a segmentation unit, and a correction unit.

[0014] The inference unit infers gesture information related to gestures by inputting speech data into a learned model.

[0015] The segmentation unit divides gesture information into high-frequency components representing rhythmic gestures that correspond to the rhythm of the dialogue and low-frequency components representing gestures that are different from rhythmic gestures.

[0016] The correction unit generates corrected gesture information by correcting at least one of the high-frequency and low-frequency components based on correction parameters.

[0017] In this information processing device, gesture information related to gestures is inferred by inputting speech data into a learned model. This gesture information is then segmented into high-frequency components representing rhythmic gestures corresponding to the dialogue rhythm and low-frequency components representing gestures different from rhythmic gestures. Furthermore, corrected gesture information is generated by correcting at least one of the high-frequency or low-frequency components based on correction parameters. Therefore, diversified expression of gestures and improved accuracy of gesture control become possible.

[0018] The information processing device may also include a playback control unit that synchronizes the corrected gesture information with the voice data and enables the avatar to perform gestures based on the corrected gesture information.

[0019] The correction parameters may include at least one of the following: an emphasis level indicating the degree of emphasis on the beat gesture, replacement information for causing the avatar to perform a predetermined gesture, or collision information related to the collision of the avatar.

[0020] The correction unit can emphasize beat gestures based on the degree of emphasis.

[0021] The correction unit can add beat gestures that are emphasized based on emphasis to gestures represented by low-frequency components.

[0022] The correction unit can replace gestures represented by low-frequency components based on replacement information.

[0023] Gestures represented by low-frequency components can include gestures accompanied by speech as well as gestures that indicate temporal or spatial orientation.

[0024] The information processing device may also include a setting unit that sets speech information related to the speaker who utters the speech based on the speech data.

[0025] Discourse information may include at least one of the following: the speaker's emotions in the discourse, the speaker's mood in the discourse, keywords included in the discourse, or the speaker's characteristics in the discourse.

[0026] The information processing device may also include a control unit that controls correction parameters based on utterance information. In this case, the control unit may control the emphasis based on at least one of emotion, mood, or characteristic.

[0027] The control unit can control the timing of replacing gestures represented by low-frequency components based on the timing of saying the keyword.

[0028] The information processing device may also include a detection unit that determines whether the avatar is colliding. In this case, the collision information may include information about whether the avatar is colliding and information about the location where the collision is occurring.

[0029] The correction unit can reduce the intensity based on collision information to prevent collisions from occurring.

[0030] The information processing method according to embodiments of this technology is an information processing method executed by a computer system, and includes: inferring gesture information related to gestures by inputting voice data into a learned model.

[0031] Gesture information is segmented into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures different from rhythmic gestures. Corrected gesture information is generated by correcting at least one of the high-frequency or low-frequency components based on correction parameters.

[0032] The recording medium of the recording program according to the embodiments of the present technology causes the computer system to perform the following steps.

[0033] The steps involved in inferring gesture-related information by inputting speech data into a learned model.

[0034] The steps involve segmenting gesture information into high-frequency components representing rhythmic gestures that correspond to the rhythm of the conversation and low-frequency components representing gestures that differ from rhythmic gestures.

[0035] The step of generating corrected gesture information by correcting at least one of the high-frequency component or the low-frequency component based on the correction parameters. Attached Figure Description

[0036] Figure 1 A block diagram showing an example configuration of an information processing apparatus according to a first embodiment.

[0037] Figure 2 A block diagram showing an example configuration of a learning unit.

[0038] Figure 3 A block diagram showing an example configuration of the voice data processing unit.

[0039] Figure 4 A block diagram showing an example configuration of the inference unit.

[0040] Figure 5 A block diagram showing an example configuration of a segmentation unit.

[0041] Figure 6 A block diagram showing an example configuration of the emphasis unit.

[0042] Figure 7 A block diagram showing an example configuration of the beat gesture reduction unit.

[0043] Figure 8 A block diagram showing a configuration example of the discourse information setting unit.

[0044] Figure 9 A schematic diagram illustrating an example of discourse segmentation.

[0045] Figure 10 A block diagram showing an example configuration of the action generation unit.

[0046] Figure 11 A block diagram showing an example configuration of the emphasis setting unit.

[0047] Figure 12 A block diagram showing an example configuration of the action determination unit.

[0048] Figure 13 A block diagram showing an example configuration of the posture information setting unit.

[0049] Figure 14 A block diagram showing an example configuration of the action segment setting unit.

[0050] Figure 15 A block diagram showing an example configuration of the reproduction control unit.

[0051] Figure 16 A block diagram showing an example configuration of an information processing apparatus according to a second embodiment.

[0052] Figure 17 A block diagram showing an example configuration of a speech synthesis unit.

[0053] Figure 18 A block diagram showing a configuration example of the discourse information setting unit.

[0054] Figure 19 A block diagram showing an example configuration of an information processing apparatus according to a third embodiment. Detailed Implementation

[0055] In the following description, embodiments according to the present technology will be described with reference to the accompanying drawings.

[0056] <First Implementation Method>

[0057] [Configuration of Information Processing Device]

[0058] Figure 1 This is a block diagram illustrating an example configuration of an information processing apparatus 1 according to a first embodiment of the present technology.

[0059] The information processing device 1 includes a learning dataset storage unit 2 and a learning unit 3 for generating a learning model, as well as a speech data processing unit 100 and an inference unit 200, a segmentation unit 300, an emphasis unit 400, a speech information setting unit 500, an action generation unit 600, a posture information setting unit 700, an action segment setting unit 800, and a reproduction control unit 900 for performing inference processing.

[0060] The learning dataset storage unit 2 stores voice data and motion information. In this embodiment, motion information is acquired synchronously with the voice data, and the motion information is processed to be subdivided into appropriate data sizes for learning processing. The learning dataset storage unit 2 stores voice data and motion information as paired datasets.

[0061] Motion information includes movements of various body parts such as the face, head, arms, hands, and legs. It should be noted that motion information is not limited to movements represented in three-dimensional space, but can also be represented as two-dimensional movements projected onto a 2D video. Furthermore, the method of representing motion is unrestricted, and motion can be represented as time-series data of coordinate values, or as time-series information of skeletal angles corresponding to joints.

[0062] It should be noted that motion information is not limited to specific body parts or skeletal structures. Furthermore, there are no restrictions on the methods used to obtain motion information, and specialized equipment and detection technologies can be used. For example, specialized kits can be used for motion capture.

[0063] It should be noted that speech data can be obtained through sound pickup devices such as microphones. Furthermore, speech data can undergo various types of speech signal processing, such as pitch adjustment and noise addition for data enhancement. It should also be noted that the length of the action information changes with the duration of the speech data. Additionally, when the learning unit, described later, performs mini-batch learning, speech data can be read in small batches (e.g., 32).

[0064] The learning unit 3 reads the necessary data size (e.g., small batch size) of speech data 10 and action information 11 from the learning dataset storage unit 2 as learning data and performs learning processing.

[0065] Figure 2 This is a block diagram showing an example configuration of learning unit 3.

[0066] like Figure 2 As shown, learning unit 3 includes feature extraction unit 4 and neural network learning unit 5.

[0067] Feature extraction unit 3 extracts feature information 12 representing speech features from speech data 10. It should be noted that the method used to extract feature information 12 is not limited and can be expressed using pitch or energy, or can be represented as a Mel spectrogram, Mel frequency cepstral coefficients (MFCC), etc., taking into account human sound perception. Feature information 12 is output to neural network learning unit 5.

[0068] The neural network learning unit 5 learns the regression problem of transforming feature information 12 into action information. For example, a convolutional neural network (CNN) such as U-Net or a recurrent neural network (RNN) such as LSTM can be used as the network architecture.

[0069] In this embodiment, the loss function for the time-series motion information is set using either intra-frame information at a specific time point or inter-frame information at different time points. Alternatively, a generative adversarial network (GAN) can be considered to set the loss function.

[0070] Through this learning process, an action generation model 13 is generated, which is a neural network capable of generating human actions based on the input speech. The action generation model 13 is then output to the inference unit 200.

[0071] The voice data processing unit 100 performs processing such as data size adjustment on the input voice test data 50. For example, the voice test data 50 is actual speech and is obtained through a sound pickup device such as a microphone. It should be noted that the voice test data 50 can be a pre-recorded audio file format or can be streaming data transmitted in any way.

[0072] Figure 3 This is a block diagram illustrating a configuration example of the voice data processing unit 100.

[0073] like Figure 3 As shown, the speech data processing unit 100 includes a learning data adjustment unit 101 and a speech data segmentation unit 102.

[0074] The learning data adjustment unit 101 applies speech signal processing to the input speech test data 50. For example, it performs adjustments to adapt to the characteristics of the learning data (speech data and action information) used in the action generation model 13. Furthermore, the speech signal processing includes, for example, noise reduction processing, bandwidth limiting processing, and volume normalization processing. Additionally, when the learning data is the speech of a specific person, speech conversion processing using deep learning can be applied, and speech signal processing can be performed to approximate the speech quality of the speaker used in the learning data.

[0075] The speech data segmentation unit 102 segments the speech test data 111, which is adjusted based on the speech information 520 output from the speech information setting unit 500, which will be described later.

[0076] The speech test data 112, processed by the learning data adjustment unit 101 and the speech data segmentation unit 102, is output to the inference unit 200.

[0077] The inference unit 200 performs inference processing based on the action generation model 13 output by the learning unit 3. In this embodiment, the inference unit 200 infers action generation data 212 by receiving inputs from speech test data 112 and the action generation model 13. The action generation data 212 includes time-series data of the posture of the human skeleton. Furthermore, the action generation data 212 is output to the segmentation unit 300.

[0078] Here, action generation data 212 includes actions performed to indicate the speaker's intention or to emphasize the dialogue. These actions include gestures that make specific movements. It should be noted that, for the sake of distinction, actions performed to indicate the speaker's intention or to emphasize the dialogue may be referred to as gestures in this disclosure.

[0079] Gestures can be divided into conventionally defined actions (such as the OK sign and numbers, symbols) and spontaneous gestures. Conventionally defined actions vary in shape and meaning depending on the society. Spontaneous gestures are those whose shape and meaning are not entirely conventionally defined, as well as gestures that occur naturally with speech.

[0080] In this embodiment, spontaneous gestures are divided into rhythmic gestures and symbolic gestures.

[0081] Rhythmic gestures are hand gestures that correspond to the rhythm of a conversation. For example, rhythmic gestures include gestures that express emphasis, actions that create rhythm by moving the hands quickly, and rhythmic movements of the body up, down, left, and right. In other words, rhythmic gestures can also be described as gestures that are independent of the content of the speech, or gestures whose shape does not change according to the content of the speech.

[0082] Symbolic gestures are gestures that indicate features in terms of time or space. Symbolic gestures can be divided into iconographic gestures that express the shape and size of an object, metaphorical gestures that express abstract concepts, and indicative gestures that indicate actions that point to a specific direction, location, or object (see Table 2 in Japanese Patent Application Publication No. 2020-6482).

[0083] It should be noted that the motion generation data 212 in this embodiment corresponds to gesture information related to gestures.

[0084] Figure 4 This is a block diagram illustrating an example configuration of the inference unit 200.

[0085] like Figure 4 As shown, the inference unit 200 includes a feature extraction unit 201 and a neural network inference unit 202.

[0086] The feature extraction unit 201 extracts feature information 211 representing speech features from the speech test data 112 output by the speech data processing unit 100. It should be noted that the method used to extract the feature information 211 is not limited and can be expressed using pitch or energy, or can be represented as a Mel spectrogram, Mel frequency cepstral coefficients (MFCC), etc., considering human sound perception. The feature information is output to the neural network inference unit 202.

[0087] The neural network inference unit 202 infers action generation data 212 based on feature information 211. In this embodiment, the neural network inference unit 202 is composed of a network architecture corresponding to the neural network learning unit 5, and the parameters of the inference network are set according to the action generation model 13. By inputting feature information into the inference network, action generation data 212 is generated.

[0088] The segmentation unit 300 segments the motion generation data 212 into low-frequency and high-frequency components in the time direction. In this embodiment, the segmentation unit 300 segments the motion generation data 212 such that beat gestures are high-frequency components, and symbolic and symbolic gestures are low-frequency components. That is, gestures other than beat gestures are classified as low-frequency components.

[0089] Furthermore, in this embodiment, symbols and symbolic gestures (hereinafter referred to as low-frequency information 306) as low-frequency components and beat gestures 307 as high-frequency components are output to the emphasis unit 400.

[0090] Figure 5 This is a block diagram illustrating an example configuration of the segmentation unit 300.

[0091] like Figure 5 As shown, the segmentation unit 300 includes a low-frequency component extraction unit 301 and an adder 302.

[0092] The low-frequency component extraction unit 301 applies a low-pass filter to the time-series data of the human skeleton posture shown in the input motion generation data 212 in the time direction, and outputs the low-frequency component of the motion data as low-frequency information 306.

[0093] Adder 302 subtracts the time-series data indicating low-frequency information 306 from the motion generation data 212 to determine the high-frequency component (beat gesture 307) of the motion data.

[0094] It should be noted that mean filters, Gaussian filters, etc., can be used as low-pass filters.

[0095] The low-frequency information 306 and the beat gesture 307, the beat gesture emphasis 650 output from the motion generation unit 600 (described later), the posture information 712 output from the posture information setting unit 700 (described later), and the collision information 913 output from the reproduction control unit 900 (described later) are input to the emphasis unit 400.

[0096] Furthermore, the emphasis unit 400 performs emphasis processing on the beat gesture 307 based on the beat gesture emphasis level 650. Additionally, the emphasis unit 400 adds the emphasized beat gesture to the gesture with a low-frequency component set by the low-frequency information 306 or the posture information 712. In this embodiment, the action with the added low-frequency component of the emphasized beat gesture (hereinafter referred to as post-action generation information 440) is output to the reproduction control unit 900.

[0097] Here, gesture information 712 is information indicating what type of gesture the 3D model should perform when it performs a predetermined gesture. For example, gestures include gestures (actions) that express a variety of emotions, such as gestures that express joy, such as raising both hands; gestures that express anger, such as clenching fists; gestures that express fear, such as covering the face; gestures that express surprise, such as covering the mouth; gestures that express disgust, such as pointing a finger at a target; or gestures that express sadness, such as lowering the head and letting the arms hang down.

[0098] It should be noted that posture includes spontaneous movements like those in a still image and temporal sequences of movements like those in a video. Furthermore, posture also includes the process of achieving a predetermined action. For example, posture includes both a static standing position with hands raised and the action of raising hands from a low position.

[0099] It should be noted that in this embodiment, the 3D model refers to model information 65, which is a virtual object on which animation reproduction data 912 output from information processing device 1 is applied. For example, the 3D model can be a 3DCG asset, such as non-photorealistic models of anime characters, or realistic models of digital humans. Furthermore, the model is not limited to 3D models, and 2D models can also be used.

[0100] Model information 65 is information related to the 3D model performing the gesture. For example, model information 65 includes mesh data for items such as the shape of various parts of the face, hands, and legs; joint positions, such as knee height; and various types of texture data, such as albedo, specular reflection, and roughness. Additionally, model information 65 may include parts representing features of the model, such as the shape of hair, clothing, and accessories. Furthermore, bones can be set for each of the hair, clothing, accessories, etc., allowing for physical swaying under collisions or gravity. Moreover, rotation direction and angle can be set for each joint, and upper limits can be set for the swaying (movement in various directions) of clothing (such as skirts) and hair.

[0101] In this embodiment, the posture information 712 is configured such that identification data indicating low-frequency information, replacement posture data, or data indicating invalidity are applied. It should be noted that in this embodiment, the posture converted to a pre-prepared (captured) posture is referred to as replacement. In other words, replacement posture data is data indicating which of a variety of predetermined postures is used for replacement.

[0102] Figure 6 This is a block diagram illustrating an example configuration of the emphasis unit 400.

[0103] like Figure 6 As shown, the emphasis unit 400 includes a beat gesture reduction unit 410, a beat gesture emphasis unit 401, a low-frequency component selection unit 402, an adder 403, and an invalid data determination unit 404.

[0104] The beat gesture reduction unit 410 performs a reduction of beat gesture emphasis 650 on joints near the location where a mesh collision occurs, based on collision information 913 output from the reproduction control unit 900, which will be described later.

[0105] Collision information 913 is information related to collisions with the model or collisions between other objects. In this embodiment, collision information 913 includes whether there is a mesh collision that causes artifacts (such as fingers or clothing passing through each other), and the location where such a collision occurs.

[0106] The beat gesture emphasis level of 650 indicates the degree of emphasis on the beat gestures. Specific settings will be provided later. Figure 10 and Figure 11 As described in the text.

[0107] Here, we will refer to Figure 7 This describes a configuration example for the beat gesture reduction unit 410. Figure 7 This is a block diagram illustrating an example configuration of the beat gesture reduction unit 410.

[0108] like Figure 7 As shown, the beat gesture reduction unit 410 includes a coefficient setting unit 411, a first multiplier 412, a second multiplier 413, an adder 414, and a delay memory 415.

[0109] The coefficient setting unit 411 sets the coefficients of the beat gesture emphasis output to the first multiplier 412 and the second multiplier 413 based on the collision information 913. In this embodiment, when a predetermined joint is close to the position where a mesh collision occurs, the coefficient setting unit 411 sets the coefficient of the beat gesture emphasis 420 input to the first multiplier 412 to 0, and sets the coefficient of the beat gesture emphasis 421 input to the second multiplier 413 to a value less than 1 (e.g., 0.9).

[0110] Furthermore, when collisions between the predetermined joint and the mesh are irrelevant, the coefficient setting unit 411 sets the coefficient of the beat gesture emphasis 420 input to the first multiplier 412 to 1.0, and sets the coefficient of the beat gesture emphasis 421 input to the second multiplier 413 to 0.

[0111] Furthermore, as mesh collisions continue to occur over time, the second multiplier 413 multiplies the past values ​​in the delay memory 415. In other words, the smaller the coefficient of the beat gesture emphasis, the smaller it gradually becomes, and the beat gesture emphasis is weakened until the point at which mesh collisions no longer occur.

[0112] By setting a coefficient for the beat gesture emphasis for all joints, the coefficient is set to remain unchanged for joints unrelated to mesh collisions, and set to a smaller coefficient for joints near where a mesh collision has already occurred. The coefficient of the beat gesture emphasis 422, taking into account the collision information 913, is input to the beat gesture emphasis unit 401.

[0113] The beat gesture emphasis unit 401 emphasizes the beat gesture 307 output from the segmentation unit 300 based on the beat gesture emphasis level 422. In this embodiment, the value of the beat gesture emphasis level 422 is multiplied by the time series data shown by the beat gesture 307. The emphasized beat gesture 430 is output to the adder 403.

[0114] The low-frequency component selection unit 402 selects either low-frequency information 306 or posture information 712 output from the posture information setting unit 700, which will be described later. In this embodiment, the low-frequency component selection unit 402 selects low-frequency information 306 when the identification data indicating low-frequency information has already been set as posture information 712. Furthermore, the low-frequency component selection unit 402 selects posture information 712 when the replacement posture data has already been set as posture information 712. In this embodiment, the data 431 of either low-frequency information 306 or posture information 712 is output to the adder 403.

[0115] For example, the identification data includes data not included in the preset predefined posture; for instance, all rotation angles showing joint movement are set to values ​​outside the defined range.

[0116] Adder 403 adds the emphasized beat gesture 430 to the data 431 (low-frequency information 306 or gesture information 712) output by low-frequency component selection unit 402. The additional information 432 resulting from this addition is provided to invalidation determination unit 404.

[0117] If the input posture information 712 is determined to be invalid, the invalidity determination unit 404 outputs invalid data (e.g., data indicating that all rotation angles of the joint movement have been set to values ​​outside the defined range) as post-motion generation information 440. Furthermore, if the input posture information 712 is not determined to be invalid, the invalidity determination unit 404 outputs additional information 432 as post-motion generation information 440.

[0118] It should be noted that the data when posture information 712 is determined to be invalid is different from the identification data of low-frequency information.

[0119] It should be noted that the emphasis unit in this embodiment corresponds to the correction unit, which generates corrected gesture information by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.

[0120] Furthermore, the beat gesture emphasis 650, posture information 712, and collision information 913 in this embodiment correspond to the correction parameters.

[0121] Based on the speech test data 50 and the setting information 55 related to the speaking task, the speech information setting unit 500 sets speech information 520, which is information related to the emotion that can be estimated based on the speech and information related to the speaker (e.g., characteristics).

[0122] In this embodiment, the discourse information 520 includes: information related to the speaker, such as emotions, moods, keywords and characteristics; and information related to the content of the speaker's speech.

[0123] Figure 8 This is a block diagram illustrating a configuration example of the speech information setting unit 500.

[0124] like Figure 8 As shown, the discourse information setting unit 500 includes a discourse segmentation unit 501, an emotion recognition unit 502, a discourse analysis unit 503, a keyword extraction unit 504, a feature analysis unit 505, and a discourse information integration unit 506.

[0125] The speech segmentation unit 501 extracts gesture switching timing for the speech test data 50. Specifically, speech segmentation is performed using existing speech recognition-based speech segmentation techniques (e.g., "Statistical Language Model for Speech Division in Speech Recognition", Nakajima et al., Transactions of the Information Processing Society of Japan, 42(11), 2681-2688, 2001-11-15), and speech time is measured and accumulated for each speech. In addition, morphological analysis of natural language processing is performed, and the timing of speaking keywords corresponding to conjunctions and silent periods exceeding a threshold are extracted (see [link to relevant documentation]). Figure 9 The timing of the end (see attached figure 531).

[0126] Furthermore, the discourse segmentation unit 501 outputs the time code as segmentation data 511 to the discourse information integration unit 506, which will be described later, when one of the following conditions is met: the timing immediately preceding the utterance of the conjunction when the accumulated discourse time exceeds a lower threshold (see...). Figure 9 (See attached figure 530); exceeding a predetermined threshold during a quiet period (see attached figure 530); Figure 9 Under the state indicated by reference numeral 531 in the attached diagram, the timing of the end of the silent period (see Figure 531). Figure 9 (See attached figure 532); or the time when the accumulated speech time exceeds the upper threshold.

[0127] The emotion recognition unit 502 identifies the speaker's emotions from the speech test data 50. For example, the speaker's emotions include feelings such as joy, anger, fear, surprise, disgust, or sadness, and are constituted by the reliability of these emotions (0.0 to 1.0). In addition, mixtures of these emotions (e.g., joy 0.6, surprise 0.3, etc.) are also included in the speaker's emotions.

[0128] It should be noted that there are no limitations on the methods for identifying emotions. For example, discourse emotion recognition techniques such as “Recognition of emotions contained in speech” (Acoustical Society of Japan, 71, 484-489 (2015)) can be applied. In addition, emotion recognition can be performed based on text information processed through speech-to-text (e.g., “Multimodal emotion estimation of text and speech 2019”, Morikawa et al., Information Processing Society of Japan Research Report (SLP) (2019)).

[0129] In this embodiment, data including speaker emotions identified by the emotion recognition unit 502 (hereinafter referred to as emotion data 512) is output to the discourse information integration unit 506.

[0130] The discourse analysis unit 503 uses discourse-to-text to convert the speech test data 50 into text and applies natural language sentiment analysis techniques (e.g., see Devlin, J. et al., "BERT: Pre-training of..."). DeepBidirectional Transformers for Language Understanding (2018) analyzes the emotion in discourse. For example, emotion includes states such as positive, negative, or neutral, and is composed of the reliability of these emotions (0.0 to 1.0). In addition, emotion also includes mixtures of the emotions mentioned above.

[0131] In this embodiment, data including discourse emotion analyzed by the discourse analysis unit (hereinafter referred to as emotion data 513) is output to the discourse information integration unit 506.

[0132] Keyword extraction unit 504 uses discourse-to-text to convert speech test data 50 into text, performs natural language processing morphological analysis, and extracts pre-registered keywords from interjections such as emotional expressions, calls, responses, greetings or shouts, as well as numbers.

[0133] In this embodiment, data including discourse keywords extracted by the keyword extraction unit 504 (hereinafter referred to as keyword data 514) is output to the discourse information integration unit 506.

[0134] The feature analysis unit 505 analyzes the speaker's characteristics (personality) based on the speech test data 50. For example, the feature analysis unit 505 uses "Automatic estimation of various personality traits contained in speech," Journal of the Acoustical Society of Japan (2015), as a technique for identifying gender and age from speech, and "Speech Personality Recognition Based on Annotation Classification Using Log-Likelihood Distance and Extraction of Essential Audio Features" (2021), as a technique for identifying human personality, to analyze the speech.

[0135] Here, the speaker's characteristics include their age, gender, nationality, and the Big Five personality traits (openness, agreeableness, conscientiousness, extraversion, and neuroticism), and are constituted by the reliability of these characteristics (0.0 to 1.0). Furthermore, the speaker's characteristics also include a mixture of these characteristics.

[0136] Furthermore, given that the speaker's characteristics are known, the user can set characteristic-related settings 55 and input these settings 55 into the characteristic analysis unit. Additionally, the characteristic analysis unit can analyze the characteristics using the speech test data 50 and the settings 55.

[0137] In this embodiment, data including speaker characteristics analyzed by the feature analysis unit 505 (hereinafter referred to as feature data 515) is output to the discourse information integration unit 506.

[0138] The discourse information integration unit 506 outputs discourse information 520, which is obtained by reusing and integrating discourse segmentation data 511, sentiment data 512, emotion data 513, keyword data 514 and feature data 515.

[0139] It should be noted that the speech information setting unit 500 in this embodiment corresponds to the setting unit, which sets speech information related to the speaker based on speech data.

[0140] The action generation unit 600 sets and outputs beat gesture emphasis 650 and action control information 680, which will be described later, based on the utterance information 520. Specifically, refer to... Figure 10 Describe it.

[0141] Figure 10 This is a block diagram showing a configuration example of the action generation unit 600.

[0142] like Figure 10 As shown, the action generation unit 600 includes a speech information segmentation unit 601, an emphasis setting unit 610, and an action determination unit 660.

[0143] The discourse information segmentation unit 601 segments the discourse information 520 into discourse segmentation data 511, sentiment data 512, discourse analysis unit 503, emotion data 513, keyword data 514, and feature data 515, which are then reused.

[0144] In this embodiment, emotional data 512, mood data 513, and characteristic data 515 are provided to the emphasis setting unit 610.

[0145] Furthermore, in this embodiment, speech segmentation data 511, emotion data 512, keyword data 514, and feature data 515 are provided to the action determination unit 660.

[0146] The emphasis setting unit 610 sets the beat gesture emphasis 650 based on emotional data 512, mood data 513 and characteristic data 515.

[0147] Figure 11 This is a block diagram illustrating a configuration example of the emphasis setting unit 610.

[0148] like Figure 11 As shown, the emphasis setting unit 610 includes an emotion setting unit 612, an emotion setting unit 613, a feature setting unit 615, and an emphasis integration unit 630.

[0149] The emotion setting unit 612 sets the emotion emphasis 622 based on the emotion data 512. For example, the emotion setting unit 612 sets a higher emphasis when the reliability of joy or anger is high, and sets a lower emphasis when the reliability of fear or sadness is high.

[0150] The emotion setting unit 613 sets the emotion emphasis 623 based on the emotion data 513. For example, the emotion setting unit 613 sets a higher emphasis when the positive reliability is high, and a lower emphasis when the negative reliability is high.

[0151] The feature setting unit 615 sets the feature emphasis 625 based on the feature data 515. For example, the feature setting unit 615 sets the emphasis higher when the speaker is younger and lower when the speaker is older. Furthermore, for example, when the feature data includes the Big Five personality traits, the emphasis is set high when the reliability of indicating extraversion is high. That is, the gesture becomes vivid and strong (“Generation of Agent Gestures that Express Personality Traits”, see Nakano et al., Human Interface Society Journal, Vol. 23, No. 2, pp. 153-164, 2021)). On the other hand, when the reliability of indicating neuroticism is high, the emphasis is set low and is set to vary slightly over time. Therefore, the gesture can be suppressed, and unstable and awkward gestures can be expressed.

[0152] Furthermore, when the feature data 515 includes multiple features (e.g., gender, age, personality, etc.), the feature setting unit 615 integrates the multiple features and sets the feature emphasis 625. Specifically, the emphasis of each feature can be calculated by weighting the average of the emphasis, multiplying multiple emphasis values, or using a calculation formula obtained from a subjective evaluation experiment.

[0153] The emphasis integration unit 630 integrates three types of emphasis output from the emotion setting unit 612, the mood setting unit 613, and the characteristic setting unit 615. For example, the emphasis integration unit 630 can integrate the emphasis by weighting the average value, multiplying multiple emphasis values, or using a calculation formula obtained from a subjective evaluation experiment. The integrated emphasis (beat gesture emphasis 650) is output to the emphasis unit 400.

[0154] It should be noted that in this embodiment, a beat gesture emphasis 650 is set for each joint of the human skeleton. Furthermore, different weights can be assigned to each joint, and the emphasis can be set based on motion control information 680 output from the motion determination unit 660, which will be described later. For example, in the case of making a gesture of raising the right arm and lowering the left arm while speaking, assuming that the right arm and right hand will move significantly, the weight of each joint (right shoulder, right elbow, right wrist, fingers, etc.) is set to large, and the weight of the left arm is set to small.

[0155] The action determination unit 660 sets action control information 680 based on speech segmentation data 511, emotion data 512, keyword data 514 and feature data 515.

[0156] Motion control information 680 is information indicating posture data used for animation reproduction. In this embodiment, motion control information 680 is information obtained by integrating timing information 671 output from the switching timing setting unit 661 (described later) and motion type information 672 output from the motion selection unit 662.

[0157] Figure 12 This is a block diagram showing an example configuration of the action determination unit.

[0158] like Figure 12 As shown, the action determination unit 660 includes a switching timing setting unit 661, an action selection unit 662, and an action integration unit 663.

[0159] The switching timing setting unit 661 measures the cumulative time used to determine the action switching based on the speech segmentation data 511. Furthermore, the switching timing setting unit 661 outputs the timing when the aforementioned cumulative time exceeds a predetermined threshold (indicating the timing of the next speech segmentation data 511) as timing information 671. In other words, timing information 671 is the time when the cumulative time meets the predetermined threshold; or, in other words, timing information 671 is information about the timing of the switch between speech segmentation data and the next speech segmentation data.

[0160] It should be noted that if the reliability of neuroticism in the input characteristic data 515 is equal to or greater than a predetermined threshold, small errors can be randomly added to the value of the timing information. Therefore, neurotic actions can be expressed by switching gestures at awkward moments.

[0161] Based on sentiment data 512, keyword data 514, and feature data 515, the action selection unit 662 sets action type information 672, which is information for selecting preset fixed actions (hereinafter referred to as action segments).

[0162] In this embodiment, when a word or phrase corresponding to a keyword pre-registered in keyword data 514 is included, the action selection unit 662 selects an action segment corresponding to that keyword. For example, the action segment may include actions using the whole body, such as bowing, or hand gestures, such as peace signs. Typically, action segments include actions that convey a common meaning across regions, countries, etc.

[0163] It should be noted that the words corresponding to the action segments can be set arbitrarily. For example, if the words contain the keyword "thank you," a bow can be selected. Furthermore, if the words contain keywords indicating a number such as "two," a symbol (e.g., a peace sign) can be selected to indicate raising the finger corresponding to that number.

[0164] Furthermore, in the absence of words corresponding to keywords pre-registered in keyword data 514, the action selection unit 662 selects actions included in the action generation data 212, rather than action segments. In this embodiment, based on emotion data 512 and feature data 515, the action selection unit 662 selects gestures corresponding to low-frequency information 306 output from the segmentation unit 300.

[0165] Specifically, the following table is defined, in which the poses actually performed by the performer are pre-captured, and a reliability value for each of the following factors is assigned: emotion, gender, age, and personality. For example, the table is defined as follows.

[0166] Posture Name: A

[0167] Emotions (Joy: 0.9, Anger: 0.1, Fear: 0.05, Surprise: 0.4, Disgust: 0.09, Sadness: 0.08)

[0168] Mood (Positive: 0.8, Negative: 0.1, Neutral: 0.4)

[0169] Characteristics (Male: 0.8, Female: 0.2, Young Adults: 0.6, Middle-aged Adults: 0.33, Older Adults: 0.07, Openness: 0.54, Agreeableness: 0.31, Conscientiousness: 0.13, Extraversion: 0.73, Neuroticism: 0.09)

[0170] Choose the pose that minimizes the sum of errors between these values ​​of reliability of emotion, mood, and traits and the values ​​of reliability output from the emotion recognition unit, discourse analysis unit, and trait analysis unit.

[0171] It should be noted that there are no restrictions on the method of assigning reliability when defining a table, and it can be set based on subjective evaluation.

[0172] Here, the gesture actually performed by the performer refers to the gesture performed by a performer with various individual characteristics (such as gender, age, and personality) included in the characteristic data, for each of multiple predetermined gestures representing the emotions (such as joy, anger, fear, surprise, disgust, or sadness) included in the emotional data. For example, a 20-year-old male performer with an outgoing personality performs a gesture of joy by raising his hands, and this gesture is captured.

[0173] In other words, the action type information 672 includes selection information, which is used to reproduce a fixed action (action fragment) corresponding to a pre-registered keyword or an action included in the action generation data 212. In this embodiment, the action type information 672 is output to the action integration unit 663.

[0174] The motion integration unit 663 integrates the timing information 671 output from the switching timing setting unit 661 and the motion type information 672 output from the motion selection unit 662, and outputs them as motion control information 680.

[0175] It should be noted that when the scenario is fixed according to the usage, the timing information 671 and action type information 672 can be forcibly set according to the scenario 60 specified by the user.

[0176] The posture information setting unit 700 outputs posture information 712 based on the input motion control information 680. Specifically, refer to... Figure 13 Describe it.

[0177] Figure 13 This is a block diagram illustrating a configuration example of the posture information setting unit 700.

[0178] like Figure 13 As shown, the posture information setting unit 700 includes an action type information segmentation unit 701, a posture information generation unit 702, and a posture information database 703.

[0179] The motion type information segmentation unit 701 segments timing information 671 and motion type information 672 from the input motion control information 680. In this embodiment, the timing information 671 and motion type information 672 are output to the posture information generation unit 702.

[0180] When the action type information 672 includes information indicating replacement posture data, the posture information generation unit 702 searches the data stored in the posture information database 703 and obtains the corresponding data 713, wherein the replacement posture data is used to reproduce the animation using the action. Furthermore, the posture information generation unit 702 outputs the data 713 obtained from the aforementioned posture information database 703 as posture information 712 at the timing indicated by the timing information 671.

[0181] Furthermore, if the action type information 672 includes information indicating the reproduction of an action segment, the posture information generation unit 702 sets invalid data to be applied to the posture information 712. Additionally, if the action type information 672 is recognition data indicating low-frequency information, the posture information generation unit 702 sets the recognition data to be applied to the posture information 712.

[0182] The motion segment setting unit 800 outputs motion segment information 812 based on the input motion control information 680.

[0183] Motion segment information 812 is information indicating the movement of various parts of a person (e.g., face, head, arms, and legs) in a preset fixed action. For example, motion segment information 812 includes information on control operations, such as how much each part is moved and how many times each joint is rotated.

[0184] Furthermore, the motion segment information 812 can be set with a time frame preceding the predetermined action. For example, in the case of performing a predetermined action within two seconds from an initial state (e.g., standing upright), a time allocation for performing predetermined operations on various parts and joints can be set, such as the speed and acceleration of raising an arm. Of course, an upper limit can be set to reproduce realistic motion.

[0185] Figure 14 This is a block diagram showing a configuration example of the action segment setting unit 800.

[0186] like Figure 14 As shown, the motion segment setting unit 800 includes a motion type information segmentation unit 801, a motion segment generation unit 802, and a motion segment database 803. It should be noted that the description of the motion type information segmentation unit 801 will be omitted because its execution is related to... Figure 13 The action type information segmentation unit 701 in the posture information setting unit 700 in the middle is similar to the processing.

[0187] When the input motion type information 672 includes information indicating the reproduction of a motion segment, the motion segment generation unit 802 searches for data stored in the motion segment database 803 and obtains the corresponding data 813. Furthermore, the motion segment generation unit 802 outputs the data 813 obtained from the aforementioned posture information database 803 as motion segment information 812 at the timing indicated by the timing information 671.

[0188] Furthermore, if the motion type information 672 indicates that the motion included in the motion generation data 212 is reproduced, the motion segment generation unit 802 sets invalid data (e.g., data in which all rotation angles indicating joint motion are set to values ​​outside the defined range) to be applied to the motion segment information 812.

[0189] The reproduction control unit 900 applies the motion determined based on the post-motion generation information 440 and the motion segment information 812 to the model information 65, wherein additionally input animation is applied to the model information 65. Furthermore, the reproduction control unit 900 outputs animation reproduction data 912, in which the 3D character model set in the model information 65 changes (operates) in a time sequence.

[0190] In this embodiment, the reproduction control unit 900 outputs collision information 913 to the emphasis unit 400 to provide feedback on the collision information. The emphasis unit 400 adjusts and controls the applied beat gesture emphasis 650 to prevent grid collisions.

[0191] Figure 15 This is a block diagram illustrating a configuration example of the reproduction control unit 900.

[0192] like Figure 15 As shown, the reproduction control unit 900 includes an effective action setting unit 901, a reproduction data generation unit 902, and a collision detection unit 903.

[0193] The valid motion setting unit 901 outputs the subsequent motion generation information 440 or motion fragment information 812 to the reproduction data generation unit 902. Specifically, since the subsequent motion generation information 440 or motion fragment information 812 always indicates invalid data, the invalid data is ignored, and data indicating the validity of the other is output. In the following text, the data indicating the validity of the subsequent motion generation information 440 or motion fragment information 812 will be referred to as valid data 911.

[0194] The reproduction data generation unit 902 applies the validity data 911 to the input model information 65 and generates animation reproduction data 912.

[0195] The collision detection unit 903 determines whether a mesh collision has occurred based on the mesh data included in the input animation reproduction data 912. In this embodiment, the collision detection unit 903 outputs collision information 913 as information indicating whether a collision has occurred and the location of the collision to the emphasis unit 400.

[0196] It should be noted that the collision detection unit 903 in this embodiment corresponds to the detection unit for determining whether the avatar is colliding.

[0197] The emphasis unit 400 reduces the value of the beat gesture emphasis 650 based on the collision information 913. Therefore, grid collision is prevented.

[0198] It should be noted that the collision determination method is not limited, and techniques such as “FCL: A general purpose library for collision and proximity queries (Jia Pan et al., 2012)” can be used.

[0199] It should be noted that in this embodiment, the motion generation unit 600 and the posture information setting unit 700 correspond to the control unit that controls the correction parameters based on speech information.

[0200] In the preceding text, the information processing apparatus 1 according to this embodiment infers action generation data 212 by inputting speech data into the action generation model 13, such that the action generation data 212 is segmented into high-frequency components representing beat gestures 307 corresponding to the rhythm of the dialogue and low-frequency information 306 representing gestures different from beat gestures 307. Furthermore, post-action generation information 440 is generated by correcting at least one of the high-frequency and low-frequency components based on beat gesture emphasis 650, posture information 712, and collision information 913. Therefore, diverse expressions of gestures can be achieved, and the control accuracy of gestures can be improved.

[0201] Conventionally, techniques for altering mouth movements based on speech are known. However, since the same speech may produce different movements, there is a problem: the mouth movements output from the inference processing tend to be the average of the learned data.

[0202] In the speech-driven gesture generation method based on speech-generated gestures according to this technology, when generating human actions (speech, facial expressions, gestures, etc.) from speech data using deep learning, the time-series data of the human skeleton's posture output from the inference processing is segmented into low-frequency and high-frequency components in the time direction for post-processing, and high-frequency component emphasis processing and low-frequency component replacement processing are applied. This allows for the addition of learning data and the achievement of diversified expressions solely through inference-side processing without altering the learning processing. Furthermore, the amplitude of gesture movements can be controlled, enabling the expression of actions in multiple postures.

[0203] Furthermore, this technology can provide the following effect: by analyzing the input speech data and generating actions that match the speaker and the content of the speech, the actions in the dialogue feel natural.

[0204] Furthermore, by controlling the motion generation process for each local part of the human based on whether mesh collisions occur during animation reproduction, unnatural animation reproduction can be prevented.

[0205] Furthermore, as an application of this embodiment, remote customer support operations can be considered. For example, by allowing customer service operators to respond solely through voice without displaying their physical appearance, operator privacy can be protected, and customer service efficiency can be achieved when performing tasks such as keyboard operations. Additionally, on the customer side, a near-face-to-face service experience can be provided, as customers can receive services through an avatar displayed as an animation generated through speech-driven actions.

[0206] <Second Implementation Method>

[0207] Figure 16This is a block diagram illustrating an example configuration of an information processing apparatus according to a second embodiment of the present technology. In the following text, descriptions of parts with configurations and operations similar to those described in the above-mentioned embodiments will be omitted or simplified. Furthermore, for simplicity, reference numerals similar to those used in the information processing apparatus described in the above-mentioned embodiments may be used in this embodiment, but the configuration is not necessarily limited to the above-mentioned embodiments.

[0208] like Figure 16 As shown, the information processing device 70 includes a speech synthesis unit 150, but not a speech data processing unit 100.

[0209] The speech synthesis unit 150 synthesizes speech based on text test data 75, which is input in place of the speech test data 50 (which is actual speech). The synthesized speech 154 is output to the inference unit 200. It should be noted that the text test data 75 includes text information that instructs the generation of silent periods, which are periods between utterances, according to the specifications of the text-to-speech (TTS) processing unit 152, which will be described later.

[0210] Figure 17 This is a block diagram illustrating an example configuration of the speech synthesis unit 150.

[0211] like Figure 17 As shown, the speech synthesis unit 150 includes a speech text segmentation unit 151 and a TTS processing unit 152.

[0212] The speech text segmentation unit 151 uses natural language processing to perform morphological analysis on the text test data 75 and counts the number of characters. In this embodiment, under the condition that the total number of counted characters falls within a predetermined range, the speech text segmentation unit 151 detects and segments sentence boundaries, giving priority to points immediately preceding the occurrence of conjunctions. The segmented text 153 is output to the TTS processing unit 152.

[0213] TTS processing unit 152 performs text-to-speech processing. In this embodiment, TTS processing unit 152 uses techniques such as speech cloning (e.g., “Deep Voice 3: Scaling text-to-speech with Convolutional Sequence Learning”, Wei Ping et al., (2017)) to approximate the speech quality of the speaker used in the learning data.

[0214] Therefore, there is no need to change the dataset on the learning side or relearn, and the action generation model learned from actual speech can be used as is. Furthermore, since the functionality of the speech data processing unit 100 is included in the speech synthesis unit 150, it is unnecessary to provide a separate speech data processing unit 100.

[0215] Figure 18 This is a block diagram illustrating a configuration example of the speech information setting unit 500.

[0216] like Figure 18 As shown, the discourse information setting unit 500 includes an emotion recognition unit 502, a discourse analysis unit 503, a keyword extraction unit 504, a feature analysis unit 505, and a discourse information integration unit 506. Figure 18 Since there is no need to segment the speech data, there is also no need for speech segmentation unit 501 or speech-to-text processing in each block.

[0217] For example, in this embodiment, the emotion recognition unit 502 uses text-based emotion recognition technology, such as the method described in, for example, “Text-based emotion detection: Advances, challenges, and opportunities,” Francisca Adoma Acheampong et al., (Computer Science and Engineering Report 2020), to identify the speaker’s (text) emotion.

[0218] Furthermore, for example, the feature analysis unit 505 analyzes features using text-based recognition techniques, such as those described in “Age and Gender prediction in Open Domain Text”, Emad E. Abdallah et al., (Procedia Computer Science, 2020) and “Attempts and considerations for utilizing writer's personality trait information estimated from text”, (Nasukawa et al., Proceedings of the 26th Annual Meeting of the Japanese Language Processing Society (2020)).

[0219] Therefore, even if no one speaks the speech data used for learning, or even if the person's speech is unavailable during inference, actions can be generated based on multiple dialogue speech data using synthesized speech of similar quality to the speech used for learning. In other words, it can be applied to AI agent services that use animated expressions of virtual characters.

[0220] <Third Implementation Method>

[0221] Figure 19 This is a block diagram illustrating a configuration example of an information processing apparatus 1000 according to a third embodiment of the present technology.

[0222] exist Figure 19 In, with Figure 1 Blocks and data similar to those in the information processing apparatus 1 shown can be described as "first". For example, the learning dataset storage unit 2, learning unit 3, inference unit 200, segmentation unit 300, and emphasis unit 400 are described as first learning dataset storage unit 2, first learning unit 3, first inference unit 200, first segmentation unit 300, and first emphasis unit 400. Similarly, Figure 19 The added blocks and data can be described as "second". Furthermore, the added blocks and data (e.g., beat gestures and low-frequency information) in this embodiment are assigned with reference numerals from the 1000 series. For example, Figure 19 The reference numeral 1306 shown in the figure refers to low-frequency information 1306.

[0223] like Figure 19 As shown, in addition to the first learning dataset storage unit 2, the first learning unit 3, the speech data processing unit 100, the first inference unit 200, the first segmentation unit 300, the first emphasis unit 400, the speech information setting unit 500, the action generation unit 600, the posture information setting unit 700, the action segment setting unit 800, and the reproduction control unit 900, the information processing device 1000 also includes a second learning dataset storage unit 1002, a second learning unit 1003, a second inference unit 1200, a second segmentation unit 1300, a second emphasis unit 1400, and an action integration unit 1020.

[0224] In other words, in the third embodiment, different learning processes are performed during the learning process. Specifically, the datasets of paired second speech data 1010 and second action information 1011 are stored in the second learning dataset storage unit 2. It should be noted that the second speech data 1010 is speech data delivered by the same person as the first speech data 10. Furthermore, the second action information 1011 is generated based on the performance of the same person as the first action information 11.

[0225] In the third embodiment, the characteristics of the second voice data 1010 and the second motion information 1011 are created to be different from the characteristics of the first voice data and the first motion information.

[0226] Specifically, the emotion recognition unit 502 sets the voice data and motion information when uttering lines of high reliability indicating joy or anger as first voice data 10 and first motion information 11. Furthermore, the speech analysis unit 503 sets the voice data and motion information when uttering lines of high positive reliability as first voice data 10 and first motion information 11.

[0227] On the other hand, the emotion recognition unit 502 sets the voice data and motion information when uttering lines of fear or sadness with high reliability as second voice data 1010 and second motion information 1011. Furthermore, the voice analysis unit 503 sets the voice data and motion information when uttering lines of negative reliability as second voice data 1010 and second motion information 1011.

[0228] In this way, two types of datasets are used to perform learning processing, and two types of action generation models are generated. That is, the first learning unit 3 generates a first action generation model 13 and outputs it to the first inference unit 200. In addition, the second learning unit 1003 generates a second action generation model 1013 and outputs it to the second inference unit 1200.

[0229] Similar to the first post-action generation information 440 inferred by the first inference unit 200, the second inference unit 1200 infers the second post-action generation data 1440 based on the second action generation model 1013.

[0230] The action integration unit 1020 generates third post-action generation data 1025 based on the first post-action generation information 440 and the second post-action generation information 1440, the emotion data 512 and the mood data 513.

[0231] Specifically, the first post-action generation information 440 and the second post-action generation information 1440 are integrated based on each reliability of the emotional data (such as joy, anger, fear, surprise, disgust, or sadness) and each reliability of the emotional data (such as positive, negative, or neutral).

[0232] Furthermore, the action integration unit 1020 can output first post-action generation information 440 or second post-action generation signal 1440 as third post-action generation data 1025. For example, if the reliability of joy, anger, and positivity is determined to be high, the first post-action generation information 440 can be output as third post-action generation information 1025. Similarly, if the reliability of fear, sadness, and negativity is determined to be high, the second post-action generation information 1440 can be output as third post-action generation information 1025.

[0233] It should be noted that there is no limit to the number of learning units and the number of inference units corresponding to the learning units, and multiple learning dataset storage units and multiple learning units can be used.

[0234] As shown in the first embodiment, when only one type of action generation model is used, there is a problem that the cycle of the beat actions tends to be monotonous. However, in the third embodiment, since the beat actions are generated based on sentiment data and emotion data from two types of action generation models generated using a network structure, more varied actions can be expressed.

[0235] <Other Implementation Methods>

[0236] This technology is not limited to the implementation methods mentioned above, and can be implemented in many other ways.

[0237] In the embodiments mentioned above, the post-action generation information 440 is applied to the model information 65, which is a virtual object. This technique is not limited to this and can also be applied to label data used in deep learning image generation techniques (e.g., “Video-to-VideoSynthesis” (2018)).

[0238] The configurations described with reference to the accompanying drawings, such as the inference unit, segmentation unit, action generation unit, and emphasis unit, are merely implementation methods and can be arbitrarily modified without departing from the essence of this technology. That is, any other configuration, algorithm, etc., used to execute this technology can be employed.

[0239] It should be noted that the effects described in this disclosure are exemplary and not limiting, and other effects may also be provided. The above description of multiple effects does not necessarily imply that all of these effects are provided simultaneously. This means that at least any of the effects mentioned above are obtained depending on conditions, and of course, effects not described in this disclosure may also be provided.

[0240] At least two features from the aforementioned implementation methods can be combined. That is, the various features described in each implementation method can be arbitrarily combined across each implementation method.

[0241] It should be noted that this technology can also be configured as follows.

[0242] (1) An information processing device, comprising:

[0243] An inference unit that infers gesture information related to gestures by inputting speech data into a learned model;

[0244] A segmentation unit, wherein the segmentation unit segments the gesture information into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures different from the rhythmic gestures; and

[0245] A correction unit generates corrected gesture information by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.

[0246] (2) The information processing apparatus according to (1) further includes:

[0247] A reproduction control unit synchronizes the corrected gesture information with the voice data and enables the avatar to perform the gesture based on the corrected gesture information.

[0248] (3) The information processing apparatus according to (1), wherein,

[0249] The correction parameters include at least one of the following: an emphasis level indicating the degree of emphasis on the beat gesture, replacement information for causing the avatar to perform a predetermined gesture, or collision information related to the collision of the avatar.

[0250] (4) The information processing apparatus according to (3), wherein,

[0251] The correction unit emphasizes the beat gesture based on the emphasis level.

[0252] (5) The information processing apparatus according to (4), wherein,

[0253] The correction unit adds the beat gesture, which is emphasized based on the emphasis, to the gesture represented by the low-frequency component.

[0254] (6) The information processing apparatus according to (4), wherein,

[0255] The correction unit replaces the gesture represented by the low-frequency component based on the replacement information.

[0256] (7) The information processing apparatus according to (1), wherein,

[0257] Gestures represented by the low-frequency components include gestures accompanied by speech and gestures that indicate temporal or spatial orientation.

[0258] (8) The information processing apparatus according to (1) further includes:

[0259] The setting unit sets speech information related to the speaker who utters the speech based on the speech data.

[0260] (9) The information processing apparatus according to (8), wherein,

[0261] The discourse information includes at least one of the following: the speaker's emotion in the discourse, the speaker's mood in the discourse, keywords included in the discourse, or the speaker's characteristics in the discourse.

[0262] (10) The information processing apparatus according to (9) further includes:

[0263] The control unit controls the correction parameters based on the speech information, wherein,

[0264] The control unit controls the emphasis based on at least one of the emotion, the mood, or the characteristic.

[0265] (11) The information processing apparatus according to (9), wherein,

[0266] The control unit controls the timing of replacing the gesture represented by the low-frequency component based on the timing of saying the keyword.

[0267] (12) The information processing apparatus according to (3) further includes:

[0268] The detection unit determines whether the avatar is colliding, wherein...

[0269] The collision information includes information about whether the avatar is colliding and information about the location where the collision is occurring.

[0270] (13) The information processing apparatus according to (12), wherein,

[0271] The correction unit reduces the intensity based on the collision information so that the collision does not occur.

[0272] (14) An information processing method, comprising:

[0273] By computer system,

[0274] Gesture-related information is inferred by feeding speech data into a learned model;

[0275] The gesture information is segmented into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures different from the rhythmic gestures; and

[0276] Corrected gesture information is generated by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.

[0277] (15) A recording medium for recording a program that causes a computer system to perform the following steps:

[0278] Gesture-related information is inferred by feeding speech data into a learned model;

[0279] The gesture information is segmented into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures different from the rhythmic gestures; and

[0280] Corrected gesture information is generated by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.

[0281] List of reference numerals

[0282] 1. Information processing device

[0283] 13 Action Generation Model

[0284] 200 inference units

[0285] 212 Action Generation Data

[0286] 300 segmentation units

[0287] 306 Low-frequency information

[0288] 307 beat gestures

[0289] 400 Emphasis Unit

[0290] 440 Post-action generation information

[0291] 500 discourse information setting units

[0292] 520 Discourse Message

[0293] 600 motion generation units

[0294] 650 beats gesture emphasis

[0295] 712 Posture Information

[0296] 900 Reproduction Control Unit

[0297] 913 Collision Information

Claims

1. An information processing apparatus, comprising: An inference unit that infers gesture information related to gestures by inputting speech data into a learned model; The segmentation unit segments the gesture information into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures different from the rhythmic gestures. as well as A correction unit generates corrected gesture information by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.

2. The information processing apparatus according to claim 1, further comprising: A reproduction control unit synchronizes the corrected gesture information with the voice data and enables the avatar to perform the gesture based on the corrected gesture information.

3. The information processing apparatus according to claim 1, wherein, The correction parameters include at least one of the following: an emphasis level indicating the degree of emphasis on the beat gesture, replacement information for causing the avatar to perform a predetermined gesture, or collision information related to the collision of the avatar.

4. The information processing apparatus according to claim 3, wherein, The correction unit emphasizes the beat gesture based on the emphasis level.

5. The information processing apparatus according to claim 4, wherein, The correction unit adds the beat gesture, which is emphasized based on the emphasis, to the gesture represented by the low-frequency component.

6. The information processing apparatus according to claim 4, wherein, The correction unit replaces the gesture represented by the low-frequency component based on the replacement information.

7. The information processing apparatus according to claim 1, wherein, Gestures represented by the low-frequency components include gestures accompanied by speech and gestures that indicate temporal or spatial orientation.

8. The information processing apparatus according to claim 1, further comprising: The setting unit sets speech information related to the speaker who utters the speech based on the speech data.

9. The information processing apparatus according to claim 8, wherein, The discourse information includes at least one of the following: the speaker's emotion in the discourse, the speaker's mood in the discourse, keywords included in the discourse, or the speaker's characteristics in the discourse.

10. The information processing apparatus according to claim 9, further comprising: The control unit controls the correction parameters based on the speech information, wherein... The control unit controls the emphasis based on at least one of the emotion, the mood, or the characteristic.

11. The information processing apparatus according to claim 9, wherein, The control unit controls the timing of replacing the gesture represented by the low-frequency component based on the timing of saying the keyword.

12. The information processing apparatus according to claim 3, further comprising: The detection unit determines whether the avatar is colliding, wherein... The collision information includes information about whether the avatar is colliding and information about the location where the collision is occurring.

13. The information processing apparatus according to claim 12, wherein, The correction unit reduces the intensity based on the collision information so that the collision does not occur.

14. An information processing method, comprising: By computer system, Gesture-related information is inferred by feeding speech data into a learned model; The gesture information is segmented into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures that are different from the rhythmic gestures. as well as Corrected gesture information is generated by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.

15. A recording medium for recording a program that causes a computer system to perform the following steps: Gesture-related information is inferred by feeding speech data into a learned model; The gesture information is segmented into high-frequency components representing rhythmic gestures corresponding to the rhythm of the dialogue and low-frequency components representing gestures that are different from the rhythmic gestures. as well as Corrected gesture information is generated by correcting at least one of the high-frequency component or the low-frequency component based on correction parameters.