Generation apparatus, generation method, and generation program
The generation device and method provide virtual characters with human-like movements and speech through speech recognition and motion coordination, improving customer interactions.
Patent Information
- Application Number
- JP2024090861
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-06-04
AI Technical Summary
Existing virtual characters lack the ability to move and communicate like real humans, hindering their widespread adoption in customer interactions.
A generation device and method that includes a recognition unit for speech processing, a generation unit for synthetic speech and motion coordination, and an output control unit to generate and control the digital human's movements and responses, utilizing machine learning models to mimic human-like interactions.
Enables virtual characters to communicate and move similarly to real humans, enhancing customer service interactions.
Smart Images

Figure 2025183014000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a generating device, a generating method, and a generating program. [Background technology]
[0002] With labor shortages and productivity improvements becoming issues, in recent years, there has been increasing consideration of utilizing advanced technologies such as XR (Cross Reality), the metaverse, and robots. Among these, the introduction of chatbots and virtual characters that respond to customers via voice has become widespread at customer contact points such as contact centers and stores. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Nobukatsu Hojo, Yusuke Ijima, and Hideyuki Mizuno. "DNN-based speech synthesis using speaker codes." IEICE TRANSACTIONS on Information and Systems 101.2 (2018): 462-472. [Retrieved July 27, 2023], Internet <URL:https: / / www.jstage.jst.go.jp / article / transinf / E101.D / 2 / E101.D_2017EDP7165 / _pdf / -char / en> Summary of the Invention [Problem to be solved by the invention]
[0004] However, there are still many customers who want to connect with real people, and in order to further popularize virtual characters, there is a demand for virtual characters that can move and communicate more like real people.
[0005] The present invention has been made in consideration of the above, and aims to provide a generation device, a generation method, and a generation program that provide a virtual character that can communicate, including movements that are similar to those of a real human. [Means for solving the problem]
[0006] In order to solve the above-mentioned problems and achieve the object, the generation device of the present invention is characterized by having a recognition unit that acquires a speech voice of a first interlocutor and performs speech recognition of the acquired speech voice to recognize a first speech text of the speech voice; a generation unit that uses a generation model to generate first synthetic speech data of a first answer text to the first speech text and first motion coordinate information indicating a first motion corresponding to the content of the first answer text; and an output control unit that outputs synthetic speech based on the first synthetic speech data from a user interface used by the first interlocutor and controls the digital human to perform the first motion on a display on which the digital human is displayed based on the first motion coordinate information. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide a virtual character that can communicate as a whole, including movements that are similar to those of a real human. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of providing a digital human by the generation system according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of providing a digital human by the generation system according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of a configuration of a generation system according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a configuration of the generating device illustrated in FIG. 3. [Figure 5] FIG. 5 is a diagram illustrating the generation of an image of a digital human. [Figure 6] FIG. 6 is a diagram illustrating an example of the configuration of the voice synthesis model shown in FIG. [Figure 7] FIG. 7 is a diagram for explaining the learning of the motion generation model shown in FIG. [Figure 8] FIG. 8 is a diagram for explaining input and output of the motion generation model shown in FIG. [Figure 9] FIG. 9 is a diagram illustrating the flow of the generation process of the generation system according to the first embodiment. [Figure 10] FIG. 10 is a sequence diagram illustrating a processing procedure of the generation process according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a configuration of a generation system according to the second embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the configuration of the generating device illustrated in FIG. [Figure 13] FIG. 13 is a diagram illustrating an example of the configuration of the voice synthesis model shown in FIG. [Figure 14] FIG. 14 is a diagram for explaining the learning of the motion generation model shown in FIG. [Figure 15] FIG. 15 is a diagram for explaining input and output of the motion generation model shown in FIG. [Figure 16] FIG. 16 is a sequence diagram illustrating a processing procedure of the generation process according to the second embodiment. [Figure 17] FIG. 17 is a diagram illustrating an example of a configuration of a generating device according to the third embodiment. [Figure 18] FIG. 18 is a diagram illustrating the processing of the generating device shown in FIG. [Figure 19] FIG. 19 is a sequence diagram illustrating a processing procedure of the generation process according to the third embodiment. [Figure 20] FIG. 20 is a diagram illustrating an example of a computer that implements the generating device by executing a program. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.
[0010] [Embodiment 1] 1 and 2 are diagrams illustrating an example of providing a digital human by the generation system according to the first embodiment.
[0011] As shown in FIGS. 1 and 2, in the embodiment, a digital human 50 is displayed on a life-size display 22 placed opposite an interlocutor 40 (first interlocutor). In the embodiment, a response from the digital human 50 according to the content of an utterance made by the interlocutor 40 into a microphone 21 is output as a synthetic voice from a speaker 23. Furthermore, in the first embodiment, the digital human 50 is made to express a movement according to the response. For example, in the first embodiment, the digital human 50 is made to express a movement according to the response with its entire body.
[0012] In this way, the first embodiment can provide a digital human 50 that communicates with users while moving in a manner similar to that of a real human, thereby improving customer service.
[0013] Here, digital human 50 is a digital character that reproduces the appearance and movements of a person on a display using technology such as computer graphics. In addition, in the first embodiment, a conversation between interlocutor 40 and digital human 50 is realized by outputting synthesized voice as the digital human 50's response from speaker 23 placed near interlocutor 40.
[0014] [Generation System] Next, a description will be given of a generation system that realizes the provision of digital human 50. Fig. 3 is a diagram showing an example of the configuration of the generation system according to the first embodiment.
[0015] As shown in FIG. 3, the generation system 100 according to the first embodiment has a configuration in which a user interface 20 installed on the interlocutor side and a generation device 10 that generates and controls the operation of a digital human 50 are communicatively connected.
[0016] The user interface 20 has a microphone 21 that collects the speech of the interlocutor 40, a life-size display 22 that displays the digital human 50, and a speaker 23 that outputs synthesized speech as the digital human 50's response. Note that, hereinafter, an example will be described in which the intention of the interlocutor 40 is acquired by collecting the speech using the microphone 21, but this is not limiting. For example, text corresponding to the intention of the interlocutor 40 may be input from an input device such as a keyboard. Furthermore, if the interlocutor 40 expresses his or her intention using sign language, the intention of the interlocutor 40 may be acquired by analyzing images of the hands, fingers, face, etc. of the interlocutor 40.
[0017] The user interface 20 has a communication interface (not shown) for transmitting and receiving various information to and from another device (the generation device 10) connected via a network or the like. The user interface 20 uses the communication interface to transmit the speech of the interlocutor collected by the microphone 21 to the generation device 10. The user interface 20 receives, via the communication interface, the synthesized speech transmitted from the generation device 10 and motion data for controlling the motions to be expressed by the digital human 50.
[0018] The generation device 10 acquires the speech of the interlocutor 40 from the user interface 20, and generates synthetic speech data indicating a response to the acquired speech and motion coordinate information indicating a motion corresponding to the response. The generation device 10 causes the user interface 20 of the interlocutor 40 to output the synthetic speech, and controls the digital human 50 to perform a motion corresponding to the response on the display.
[0019] [Generation System] Fig. 4 is a diagram showing an example of the configuration of the generating device 10 shown in Fig. 3. As shown in Fig. 4, the generating device 10 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0020] The communication unit 11 is a communication interface that transmits and receives various information to and from other devices connected via a network, etc. The communication unit 11 is realized by a NIC (Network Interface Card) or the like, and performs communication between other devices (e.g., the user interface 20) and the control unit 13 (described later) via electric communication lines such as a LAN (Local Area Network) or the Internet.
[0021] The storage unit 12 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, and stores a processing program that operates the generation device 10, data used during execution of the processing program, etc. The storage unit 12 has digital human generation data 121.
[0022] The digital human generation data 121 is image data, a generation program, and the like used when generating the body (or the entire body) including the facial features of the digital human 50 using computer graphics.
[0023] 5 is a diagram illustrating the generation of a digital human image. In the first embodiment, the digital human 50 is created by scanning images of the faces and bodies of multiple realistic people (e.g., nine people) ((1) and (2) in FIG. 5), aggregating, and processing them to create an original digital human 50 that even expresses realistic textures ((3) in FIG. 5).
[0024] Creating an original digital human's appearance from scratch is difficult, and requires careful consideration and design of the desired appearance. Furthermore, simply creating a digital human from a single person can lead to issues with portrait rights and other issues. Therefore, in the first embodiment, by aggregating and processing multiple people who actually work for a business that provides customer support, it is possible to generate a digital human 50 that resembles an employee of the business or embodies the business.
[0025] The control unit 13 controls the entire generating device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 13 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory. The control unit 13 also functions as various processing units by running various programs. The control unit 13 has a voice recognition unit 131, a generation unit 132, and an output control unit 136.
[0026] The speech recognition unit 131 acquires speech data of the speech of the interlocutor 40 via communication with the user interface 20, performs speech recognition on the acquired speech, and recognizes a first speech text of the speech. The speech recognition unit 131 performs speech recognition processing using a trained speech recognition model (machine learning model) that outputs speech text from the speech data.
[0027] The generation unit 132 uses a generation model 133 (machine learning model) to generate first synthetic voice data of a first answer text to a first spoken text and first motion coordinate information indicating a first motion corresponding to the content of the first answer text.
[0028] The first synthetic voice data is output from the user interface 20 as the voice of the digital human 50 under the control of the output control unit 136. The first synthetic voice data is provided with time information indicating the output time of the synthetic voice. The first motion coordinate information is information indicating the coordinate position of each joint of the digital human 50. The first motion coordinate information associates the coordinate position of each joint of the digital human 50 with time information indicating the placement time for placing the joint at this coordinate position. The time information of the first synthetic voice data and the first motion coordinate information is set so that the voice and movement of the digital human 50 are linked.
[0029] The generative model 133 is a model that has learned the relationship between the utterance text of the interlocutor's speech and the answer text, and the relationship between the motion of the digital human corresponding to the content of the answer text. When a first utterance text is input, the generative model 133 outputs first synthetic speech data of the first answer text to the first utterance text and first motion coordinate information indicating a first motion corresponding to the content of the first answer text. The generative model 133 includes, for example, a speech synthesis model 134 and a motion generation model 135.
[0030] The speech synthesis model 134 learns a plurality of pairs of data, each consisting of an utterance text of a speech of an interlocutor and a response text corresponding to the utterance text, as training data. The speech synthesis model 134 performs speech synthesis of a first response text corresponding to an input first utterance text based on the speech of a plurality of people, and outputs first synthetic speech data.
[0031] Fig. 6 is a diagram illustrating an example of the configuration of the speech synthesis model 134 shown in Fig. 4. As shown in Fig. 6, the speech synthesis model 134 includes a dialogue model 1341 and a speech synthesis unit 1342 that vocalizes the input first answer text with synthetic speech.
[0032] The dialogue model 1341 is a machine learning model that learns the relationship between an input utterance text and a response text to the utterance text. FAQ (Frequently Asked Questions) data, which are expected to be questions from interlocutors, and the response data to the utterance text are set as training data for the dialogue model 1341. The dialogue model 1341 learns the relationship between an input utterance text and a response text to the utterance text using supervised learning based on the training data ((1) in FIG. 6). The speech synthesis unit 1342 is a speech synthesizer that outputs a similar speech based on a live speech actually uttered by a person ((3) in FIG. 6).
[0033] When the dialogue model 1341 receives a first utterance text of the speech of the interlocutor 40, it outputs a first reply text to the input first utterance text ((2) in FIG. 6). The speech synthesis unit 1342 converts the first reply text output by the dialogue model 1341 into speech and outputs first synthetic speech data ((4) in FIG. 6).
[0034] When the first utterance text of interlocutor 40 and the first synthetic voice data by voice synthesis model 134 are input, motion generation model 135 outputs first motion coordinate information indicating a first motion corresponding to the first synthetic voice data. Fig. 7 is a diagram explaining the learning of motion generation model 135 shown in Fig. 4. Fig. 8 is a diagram explaining the input and output of motion generation model 135 shown in Fig. 4.
[0035] The motion generation model 135 is a model that has learned the relationship between the interlocutor's spoken text, synthetic speech data of a reply text to the interlocutor's spoken text, and motion coordinate information that indicates the motion corresponding to the synthetic speech data. The motion generation model 135 learns the speech of a person for motion learning, the speech text corresponding to this person's speech, the physical movements of the person for motion learning, and synthetic speech data obtained by vocalizing the reply text corresponding to the speech text spoken by the person for motion learning as learning data (FIG. 7).
[0036] When the first utterance text of interlocutor 40 and the first synthetic voice data by voice synthesis model 134 are input, motion generation model 135 outputs first motion coordinate information indicating a first motion corresponding to the first utterance text of interlocutor 40 and the first synthetic voice data (FIG. 8). Motion generation model 135 generates motion coordinates according to the input first synthetic voice data and the first utterance text of interlocutor 40 ((1) in FIG. 8).
[0037] The output control unit 136 outputs the first synthetic voice generated by the generation unit 132 from the user interface 20 of the interlocutor 40, and also controls the digital human 50 to perform the first motion on the display 22 based on the first motion coordinate information generated by the generation unit 132.
[0038] Based on the first synthetic voice data and first motion coordinate information generated by the generation unit 132, the output control unit 136 adjusts the image of the digital human 50 to be output to the display 22 and the synthetic voice to be output from the speaker 23. In this way, the generation device 10 reproduces the first synthetic voice data and first motion corresponding to the utterance of the interlocutor 40 on the digital human 50 ((2) in FIG. 8).
[0039] [Generation process] Next, a description will be given of the generation processing of the generation system 100 according to the embodiment 1. Fig. 9 is a diagram illustrating the flow of the generation processing of the generation system 100 according to the embodiment 1. Fig. 10 is a sequence diagram showing the processing procedure of the generation processing according to the embodiment 1.
[0040] The user interface 20 receives the speech of the interlocutor 40 from the microphone 21 ((1) in FIG. 9), and transmits the speech of the interlocutor 40 to the generation device 10 ((2) in FIG. 9, step S1 in FIG. 10).
[0041] In the generation device 10, the speech recognition unit 131 acquires the speech of the interlocutor 40, performs speech recognition on the acquired speech, and recognizes a first speech text of the speech (step S2 in FIG. 10). The speech recognition unit 131 outputs the first speech text to the generation unit 132 ((3) in FIG. 9, step S3 in FIG. 10).
[0042] The generation unit 132 uses the generation model 133 to generate a first response text to the first spoken text, and generates the generated first synthetic voice data and first motion coordinate information indicating a first motion corresponding to the content of the first response text (steps S4, S5, and S6 in FIG. 10).
[0043] In the generation model 133, the speech synthesis model 134 generates a first response text in response to the first utterance text, generates the generated first synthetic speech data, and outputs it to the motion generation model 135 ((4) in FIG. 9). When the first utterance text and the first synthetic speech data are input, the motion generation model 135 outputs the first motion coordinate information together with the first synthetic speech data to the output control unit 136 ((5) in FIG. 9, step S7 in FIG. 10).
[0044] In the generating device 10, the output control unit 136 generates output data for the user interface 20 (step S8 in FIG. 10). For example, the output control unit 136 generates an image of the digital human 50 performing a first motion corresponding to the first motion coordinate information, and generates control data for outputting the synthetic voice of the first synthetic voice data.
[0045] The output control unit 136 uses control data including the first synthetic voice data and motion data related to the first motion to output the synthetic voice of the first synthetic voice data from the user interface 20 of the interlocutor 40, and causes the digital human 50 to reproduce the first motion on the display 22 based on the first motion coordinate information generated by the generation unit 132 ((6-1) and (6-2) in Figure 9, steps S9 and S10 in Figure 10).
[0046] [Effects of the First Embodiment] Therefore, the generation device 10 can provide a digital human 50 that responds with synthetic voice in response to the content of the speech of the interlocutor 40, including movements similar to those of a real human. Therefore, the generation device 10 can provide a digital human 50 that can communicate smoothly with the interlocutor 40 by behaving and speaking like a real human, which is expected to improve customer service.
[0047] [Embodiment 2] Next, a description will be given of embodiment 2. Fig. 11 is a diagram showing an example of the configuration of a generation system according to embodiment 2.
[0048] In the generation system 200 according to the second embodiment, the generation device 210 is connected to the emotion determination device 230 so as to be able to communicate with each other.
[0049] The emotion determination device 230 analyzes the input speech of the interlocutor 40 and determines the type of emotion of the interlocutor 40, who is the person who uttered the speech. The emotion determination device 230 transmits first emotion information indicating the type of emotion of the interlocutor 40 to the generation device 210.
[0050] The emotion determination device 230 determines the type of emotion of the person who uttered the speech by acquiring and analyzing physical features of the speech, such as intonation and volume of the voice, and the spoken text of the speech. The types of emotions include, for example, "joy," "anxiety," "anger," "sadness," and "surprise."
[0051] The emotion determination device 230 also analyzes the input answer text and determines the type of emotion indicated by the answer text. The emotion determination device 230 has machine learning models such as a natural language processing model and a voice recognition model.
[0052] Based on the first utterance text of the interlocutor 40 and the emotion information (first emotion information) of the interlocutor 40 determined by the emotion determination device 230, the generation device 210 outputs synthetic speech data of a first reply text to the first utterance text according to the first emotion information, and first motion coordinate information indicating a first motion corresponding to the content of the first reply text. The generation device 210 then reproduces a digital human 50 that responds with a tone of voice and gestures based on the emotion of the interlocutor 40.
[0053] [Generation device] Fig. 12 is a diagram showing an example of the configuration of the generating device 210 shown in Fig. 11. The generating device 210 has a control unit 213 instead of the control unit 13 shown in Fig. 4.
[0054] The control unit 213 includes an emotion information acquisition unit 2131 , a voice recognition unit 131 , a generation unit 2132 , and an output control unit 2136 .
[0055] The emotion information acquisition unit 2131 inputs the speech of the interlocutor 40 to the emotion determination device 230 and makes a determination request to determine the emotion information of the speech. By receiving the emotion information determined by the emotion determination device 230, the emotion information acquisition unit 2131 acquires first emotion information indicating the type of emotion of the interlocutor 40 determined based on the speech of the interlocutor 40.
[0056] The generation unit 2132 uses the generation model 2133 to generate first synthetic speech data of a first answer text to a first utterance text and first emotional information, and first motion coordinate information indicating a first motion corresponding to the content of the first answer text and the first emotional information.
[0057] The generative model 2133 (machine learning model) is a model that learns the relationship between emotion information indicating the type of emotion of a conversation partner and a response text to the conversation partner's utterance text, and the relationship between the emotion information indicating the type of emotion of the conversation partner and a motion of a digital human corresponding to the content of the response text. When a first utterance text and first emotion information are input, the generative model 2133 outputs synthetic speech data of a first response text to the first utterance text according to the first emotion information, and first motion coordinate information indicating a first motion corresponding to the first emotion information and the content of the first response text. The generative model 2133 has a voice synthesis model 2134 and a motion generation model 2135.
[0058] When a first utterance text and first emotion information are input, the speech synthesis model 2134 generates synthetic speech data of a first reply text to the first utterance text according to the first emotion information. The speech synthesis model 2134 learns, as training data, a plurality of pair data of utterance texts of the interlocutor's speech and reply texts corresponding to the utterance texts, each of which is accompanied by emotion information (e.g., emotion tags) indicating the type of emotion.
[0059] Fig. 13 is a diagram illustrating an example of the configuration of the speech synthesis model 2134 shown in Fig. 12. As shown in Fig. 13, the speech synthesis model 2134 has a dialogue model 21341 and a speech synthesis unit 21342 that vocalizes the input first answer text with synthetic speech.
[0060] The dialogue model 21341 is a machine learning model that learns the relationship between utterance text and response text to the utterance text according to various emotions. The dialogue model 21341 is data to which emotion tags indicating the type of emotion are attached, and FAQ data expected as questions from the interlocutor and their response data are set as training data. The dialogue model 21341 learns the relationship between utterance text and response text to the utterance text according to each type of emotion using supervised learning based on the training data ((1) in FIG. 13). The speech synthesis unit 21342 is a speech synthesizer that outputs similar speech based on the actual raw speech uttered by a person ((3) in FIG. 13).
[0061] When the first utterance text and first emotional information of the speech of interlocutor 40 are input, dialogue model 21341 outputs the first emotional information and a first reply text and first emotional information to the first utterance text ((2) in FIG. 13). The speech synthesis unit 21342 vocalizes the first reply text output by dialogue model 21341 in accordance with the first emotional information, and outputs first synthetic speech data and the first emotional information ((4) in FIG. 13).
[0062] Furthermore, the generation unit 2132 may request the emotion determination device 230 to determine the emotion of the first answer text generated by the dialogue model 21341, and determine whether the first answer text has an emotion that is in line with the first utterance text. In this case, if the emotion of the first answer text determined by the emotion determination device 230 is in line with the first utterance text, the generation unit 2132 converts the first answer text into synthetic speech.
[0063] For example, if the first emotion information indicates strong "anxiety," the dialogue model 21341 generates a polite answer text. If the first emotion information indicates strong "anxiety," the dialogue model 21341 may generate a first answer text and make an emotion determination request until the emotion of the first answer text is determined to be "relief."
[0064] For example, if the first emotional information indicates a strong sense of "anxiety," the speech synthesis unit 21342 vocalizes the first answer text in a slow tone and outputs first synthetic speech data. Furthermore, if the first emotional information indicates a strong sense of "joy," the dialogue model 21341 generates answer text in a friendly tone. If the first emotional information indicates a strong sense of "joy," the speech synthesis unit 21342 vocalizes the first answer text in a cheerful manner and outputs first synthetic speech data.
[0065] The speech synthesis unit 21342 converts the first answer text into synthesized speech in a tone of voice such as emphasizing the speech tone, speaking sadly, speaking happily, speaking calmly, etc., depending on the first emotional information and / or the emotion determination result of the first spoken text.
[0066] When the first utterance text of interlocutor 40, first synthetic voice data by voice synthesis model 134, and first emotion information are input, motion generation model 2135 outputs first motion coordinate information indicating a first motion corresponding to the first synthetic voice data and the first emotion information. Fig. 14 is a diagram explaining the learning of motion generation model 2135 shown in Fig. 12. Fig. 15 is a diagram explaining the input and output of motion generation model 2135 shown in Fig. 12.
[0067] The motion generation model 2135 is a model that has learned the relationship between the interlocutor's spoken text, an emotion tag indicating the type of emotion of the interlocutor, synthetic speech data of an answer text based on the emotion, and motion coordinate information indicating the motion corresponding to the synthetic speech data. The motion generation model 135 learns the speech of a person for motion learning, an emotion tag indicating the type of emotion of the person for motion learning, the speech text corresponding to this speech, the physical movements of the person for motion learning, and synthetic speech data obtained by vocalizing the answer text corresponding to the speech text spoken by the person for motion learning as training data (FIG. 14).
[0068] When first emotion information of interlocutor 40, a first utterance text of interlocutor 40, and first synthetic voice data based on the emotion by voice synthesis model 134 are input, motion generation model 2135 outputs first motion coordinate information indicating a first motion corresponding to the emotion information of interlocutor 40, the first utterance text of interlocutor 40, and the first synthetic voice data (FIG. 15). Motion generation model 135 generates motion coordinates according to the first emotion information, the first synthetic voice data, and the first utterance text of interlocutor 40 ((1) in FIG. 15).
[0069] For example, if the first emotion information indicates a strong "anxiety," the motion generation model 2135 outputs first motion coordinate information indicating a first motion showing polite gestures. Also, if the first emotion information indicates a strong "joy," the motion generation model 2135 outputs first motion coordinate information indicating a first motion showing vigorous gestures.
[0070] Furthermore, when emotion determination is performed on the first answer text, the motion generation model 2135 may acquire the emotion determination result for the first answer text and set a first motion corresponding to the first emotion information and / or the emotion determination result for the first answer text. For example, when the emotion of the first answer text is determined to be negative, the motion generation model 2135 sets a motion of standing upright with hands clasped without making any gestures. When the emotion of the first answer text is determined to be positive, the motion generation model 2135 sets a motion of making a wide gesture.
[0071] The output control unit 2136 controls the user interface 20 of the interlocutor 40 to output the first synthetic voice generated by the generation unit 132, and controls the digital human 50 to perform a first motion on the display 22 based on the first motion coordinate information generated by the generation unit 132. The output control unit 2136 adjusts the image of the digital human 50 to be output to the display 22 and the synthetic voice to be output from the speaker 23 based on the first synthetic voice data based on the emotion of the interlocutor 40 and the first motion coordinate information based on the emotion of the interlocutor 40 generated by the generation unit 2132 (FIG. 15).
[0072] As a result, the generating device 210 causes the digital human 50 to perform a first motion generated based on the first emotional information indicating the type of emotion of the interlocutor 40, and causes the speaker 23 to output a first synthetic voice generated based on the first emotional information.
[0073] Therefore, for example, the generation device 210 can cause the digital human 50 to respond to a participant 40 who is highly "anxious" with a polite tone of voice and polite gestures. Also, the generation device 210 can cause the digital human 50 to respond to a participant 40 who is highly "joyed" with a friendly tone of voice and vigorous gestures.
[0074] [Generation process] Next, a description will be given of the generation process of the generation system 200 according to the second embodiment. Fig. 16 is a sequence diagram showing the processing procedure of the generation process according to the second embodiment.
[0075] Step S21 in Fig. 16 is the same process as step S1 in Fig. 10. Steps S25 and S26 in Fig. 16 are the same process as steps S2 and S3 in Fig. 10.
[0076] The generation device 210 transmits the speech and a request to determine the type of emotion of the person who uttered the speech to the emotion determination device 230 (steps S22 and S23).
[0077] The emotion determination device 230 analyzes the input speech and determines the type of emotion of the interlocutor 40 who uttered the speech (step S24). The emotion determination device 230 transmits first emotion information indicating the type of emotion of the interlocutor 40 (step S27).
[0078] The generation unit 2132 generates a first reply text to the first utterance text according to the first emotion information, based on the first utterance text and first emotion information of the interlocutor 40 (step S28). The generation unit 2132 transmits a request for emotion determination for the generated reply text to the emotion determination device 230 (step S29). The emotion determination device 230 analyzes the input first reply text and determines the type of emotion of this first reply text (step S30).
[0079] The generation unit 2132 receives emotion information returned by the emotion determination device 230 (step S31). If the emotion type of the first answer text matches the first emotion information, the generation unit 2132 generates synthetic voice data of the first answer text (step S32). Next, the generation unit 2132 generates first motion coordinate information indicating a first motion corresponding to the first emotion information and the content of the first answer text (step S33).
[0080] Steps S34 to S37 in FIG. 16 are the same processes as steps S7 to S10 in FIG.
[0081] [Effects of the second embodiment] For example, the generation device 210 adjusts the tone of voice of the digital human 50 according to the emotions of the interlocutor 40, such as speaking more strongly, speaking sadly, speaking happily, or speaking calmly. The generation device 201 can make the digital human 50 respond to an interlocutor 40 who is highly "anxious" with a polite tone of voice and polite gestures. The generation device 210 can also make the digital human 50 respond to an interlocutor 40 who is highly "joyful" with a friendly tone of voice and intense gestures.
[0082] In this way, the generation device 210 can provide a digital human 50 that adopts a tone of voice and motions that correspond to the emotions of the interlocutor 40. Therefore, the generation device 210 can provide a digital human 50 that can communicate more smoothly with the interlocutor 40.
[0083] [Embodiment 3] Next, a description will be given of embodiment 3. Fig. 17 is a diagram showing an example of the configuration of a generation device according to embodiment 3. Compared to generation system 200 shown in Fig. 10, the generation system according to embodiment 3 has a generation device 310 shown in Fig. 17.
[0084] The generating device 310 includes a control unit 313 having an output control unit 3136 instead of the control unit 213 shown in FIG.
[0085] The output control unit 3136 causes the user interface 20 of the interlocutor 40 to output the first synthetic voice. At the same time, the output control unit 3163 controls the digital human 50 to perform a first motion with an expression corresponding to the first emotional information of the interlocutor determined by the emotion determination device 230 based on the speech of the interlocutor. Figure 18 is a diagram illustrating the processing of the generation device 310 shown in Figure 17.
[0086] The output control unit 3136 adjusts the image of the facial expression of the digital human 50 to be output to the display 22 and the synthetic voice to be output from the speaker 23 based on the first synthetic voice data based on the emotion of the interlocutor 40 and the first motion coordinate information based on the emotion of the interlocutor 40 generated by the generation unit 2132 (Figure 18).
[0087] For example, if the emotion of the interlocutor 40 is negative, the output control unit 3136 controls so that the digital human 50 has a sad expression. If the emotion of the interlocutor 40 is positive, the output control unit 3136 controls so that the digital human 50 has a happy expression. Depending on the emotion of the interlocutor 40, the output control unit 3136 controls so that the digital human 50 has an angry expression, a sad expression, a happy expression, or the like.
[0088] When emotion determination is performed on the first answer text, the output control unit 3136 may acquire the emotion determination result of the first answer text and perform control so that a digital human 50 with a facial expression is output. When the emotion determination result of the first answer text is negative, the output control unit 3136 performs control so that a digital human 50 with a sad expression is output. When the emotion determination result of the first answer text is positive, the output control unit 3136 performs control so that a digital human 50 with a happy expression is output.
[0089] This allows the generation device 310 to make the digital human 50 produce synthetic voice and make motions that correspond to the emotions of the interlocutor 40, and also make facial expressions that correspond to the emotions of the interlocutor 40.
[0090] [Generation process] Next, a description will be given of the generation process of the generation system according to Embodiment 3. Fig. 19 is a sequence diagram showing the processing procedure of the generation process according to Embodiment 3.
[0091] Steps S41 to S54 shown in Fig. 19 are the same processes as steps S21 to S24 shown in Fig. 16. Output control unit 3136 generates synthetic speech, motion data, and a facial expression image based on the emotion of interlocutor 40 as output data from user interface 20 (step S55). Then, output control unit 3136 causes user interface 20 of interlocutor 40 to output synthetic speech of the first synthetic speech data with a facial expression corresponding to the expression of interlocutor 40, and causes digital human 50 to reproduce the first motion (steps S56 and S57).
[0092] [Embodiment 3] In this way, generation device 310 can provide digital human 50 that uses facial expressions and speech patterns and motions that correspond to the emotions of interlocutor 40. Therefore, generation device 310 can provide digital human 50 that can communicate more smoothly with interlocutor 40.
[0093] In addition, the generation devices 210, 310 may be linked to CRM (Customer Relationship Management) or the like, and may control the output of the digital human 50 so that the digital human 50 responds with emotional tone of voice, motion, and / or facial expressions in accordance with the visit information and other attribute information of the user who will be the interlocutor 40.
[0094] Furthermore, the generation devices 10, 210, and 310 may perform image recognition processing on images of the face-to-face interlocutor 40 to provide personalized information. For example, the generation devices 10, 210, and 310 may change the content of speech of the digital human 50 based on gender, age, facial expression, and the like, or may output advertisements tailored to the interlocutor 40 from the display 22 or the like to guide the interlocutor 40. Furthermore, for example, the generation devices 10, 210, and 310 may adjust the line of sight and body orientation of the digital human 50 to match the line of sight and body orientation of the interlocutor 40.
[0095] [System configuration of the embodiment] The components of the generation devices 10, 210, and 310 are conceptual functional components and do not necessarily need to be physically configured as shown in the drawings. In other words, the specific forms of distribution and integration of the functions of the generation devices 10, 210, and 310 are not limited to those shown in the drawings, and all or part of them can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.
[0096] Furthermore, all or any part of the processes performed by the generation devices 10, 210, and 310 may be realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Furthermore, each process performed by the generation devices 10, 210, and 310 may be realized as hardware using wired logic.
[0097] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.
[0098] [program] 20 is a diagram showing an example of a computer in which generation devices 10, 210, and 310 are realized by executing a program. Computer 1000 includes, for example, memory 1010 and CPU 1020. Computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0099] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0100] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the generation device 10, 210, 310 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the generation device 10, 210, 310 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0101] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.
[0102] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0103] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]
[0104] 10,210,310 generator 11 Communications Department 12 Storage section 13,213,313 Control Unit 100,200 Generation System 121 Data for generating digital humans 131 Voice Recognition Unit 132 Generation part 133,2133 Generative Model 134,2134 speech synthesis model 135,2135 Motion Generation Model 136,2136,3136 Output control section 230 Emotion determination device 1341,21341 Dialogue Model 1342,21342 Speech synthesis unit
Claims
1. a recognition unit that acquires a speech of a first interlocutor and performs speech recognition on the acquired speech to recognize a first spoken text of the speech; a generation unit that generates, using a generation model, first synthetic speech data of a first response text to the first spoken text and first motion coordinate information that indicates a first motion corresponding to the content of the first response text; an output control unit that outputs synthetic speech based on the first synthetic speech data from a user interface used by the first interlocutor, and controls the digital human to perform the first motion on a display on which the digital human is displayed, based on the first motion coordinate information; A generating device comprising:
2. the generative model is a model that learns a relationship between a response text and an utterance text of a speech of an interlocutor, and a relationship between a motion of a digital human corresponding to the content of the response text, The generating device according to claim 1, characterized in that, when the first spoken text is input, the generating model outputs first synthetic speech data of a first answer text to the first spoken text and first motion coordinate information indicating a first motion corresponding to the content of the first answer text.
3. The generative model is a speech synthesis model that learns a plurality of pairs of data, each pair consisting of an utterance text of a speech of an interlocutor and a response text corresponding to the utterance text, as training data, performs speech synthesis of a first response text corresponding to the input first utterance text based on the speeches of a plurality of people, and outputs the first synthetic speech data; a motion generation model that learns, as training data, a speech voice of a person, a speech text corresponding to the speech voice of the person, a body movement of the person, and synthetic speech data obtained by converting an answer text corresponding to the speech text spoken by the person into a voice, and that, when the first speech text and first synthetic speech data generated by the speech synthesis model are input, outputs first motion coordinate information indicating a first motion corresponding to the first speech text and the first synthetic speech data; The generating device according to claim 2, further comprising:
4. The communication device further includes an acquisition unit that acquires first emotion information indicating a type of emotion of the first interlocutor determined based on the speech of the first interlocutor, 2. The generation device according to claim 1, wherein the generation unit uses the generation model to generate the first synthetic speech data of the first answer text to the first utterance text and the first emotional information, and first motion coordinate information indicating a first motion corresponding to the content of the first answer text and the first emotional information.
5. the generative model is a model that has learned a relationship between emotion information indicating a type of emotion of a conversation partner and a response text to an utterance text of the conversation partner, and a relationship between the emotion information indicating a type of emotion of the conversation partner and a motion of a digital human corresponding to the content of the response text; 5. The generation device according to claim 4, wherein, when the first spoken text and the first emotional information are input, the generation model outputs synthetic speech data of a first answer text to the first spoken text according to the first emotional information, and first motion coordinate information indicating a first motion corresponding to the content of the first answer text.
6. The generation device according to claim 4, characterized in that the output control unit outputs the synthetic speech from a user interface of the first interlocutor and controls the digital human to perform the first motion with an expression corresponding to the first emotional information.
7. The generating device according to claim 1 , wherein the digital human is a reproduction of the appearance and movements of a person generated by aggregating and processing the appearances of a plurality of people.
8. A generation method executed by a generation device, comprising: acquiring a speech of a first interlocutor, and performing speech recognition on the acquired speech to recognize a first spoken text of the speech; generating, using a generative model, first synthetic speech data of a first response text to the first spoken text and first motion coordinate information indicating a first motion corresponding to the content of the first response text; outputting synthetic speech based on the first synthetic speech data from a user interface used by the first interlocutor, and controlling the digital human to perform the first motion on a display on which the digital human is displayed, based on the first motion coordinate information; A generating method comprising:
9. acquiring a speech of a first interlocutor, and performing speech recognition on the acquired speech to recognize a first spoken text of the speech; generating, using a generative model, first synthetic speech data of a first response text to the first spoken text and first motion coordinate information indicating a first motion corresponding to the content of the first response text; outputting synthetic speech based on the first synthetic speech data from a user interface used by the first interlocutor, and controlling the digital human to perform the first motion on a display on which the digital human is displayed based on the first motion coordinate information; A generating program for causing a computer to execute the above.
Citation Information
Patent Citations
Speech synthesizer, speech synthesis method and speech synthesis program
JP2010128103A
Interactive response method and computer system using the same
JP2020047240A
Avatar creation device, mobile terminal, accessory matching system, avatar creation method, and program
JP2021051618A
Speech synthesis device, speech synthesis program, and speech synthesis method
JP2021099454A
Method, device, apparatus and medium for man-machine interactions
JP2021168139A