Information processing device and information processing method
The information processing device automates the integration of dialogue data from multiple communication modes by identifying and integrating text data and timing information, addressing the inefficiency of manual data comparison and integration in existing techniques.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NT T INC
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-07
AI Technical Summary
Existing techniques for generating dialogue data between individuals in different communication modes, such as voice and text, or sign language and text, require manual comparison and integration of data, which is burdensome and inefficient.
An information processing device and method that automatically acquires dialogue data by identifying and integrating text data and timing information from multiple speakers, regardless of their communication mode, using a control unit to process first and second information, including speaker identification data.
Reduces the burden of manually creating dialogue data by automating the integration process, enabling efficient acquisition of dialogue data across different communication modes.
Smart Images

Figure JP2024038530_07052026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus and Information Processing Method
[0001] The present invention relates to an information processing apparatus and an information processing method.
[0002] In a virtual space, there is a technique (Non-Patent Document 1) in which an avatar makes a statement that makes the listener feel as if the user is present even though the user is absent. Such a statement by the avatar is generated by a mathematical model obtained based on the data of the conversation that the user himself / herself had with others in the virtual space, a mathematical model that has learned the user's personality. Therefore, collection of conversation data is required for such a technique. Since the way of expressing words is often common regardless of the user in the virtual space, collection of conversation data is easy.
[0003] NTT Corporation, “Individuality Reproduction Dialogue Technology for efficientreproduction of an individual's speech using LLMsDigital alter egos can be generated at low cost using NTT LLM"tsuzumi"”, January 17, 2024
[0004] By the way, such a technique in which an avatar makes a statement as if it were the user's words even though the user is absent does not necessarily need to be used only within the virtual space. For example, it is also expected to be applied to conversations between people existing in different spaces such as video distribution services and online learning.
[0005] However, in the case of conversations between people existing in different spaces, unlike conversations in the virtual space, not only are the spaces where the distributor and the viewer exist different, but the ways of expressing words are often not the same. For example, in a video distribution service, the distributor expresses words by voice, and the viewer expresses words by text. Specifically, for example, the distributor sees the words transmitted by the listener in text, and the distributor responds to it by voice.
[0006] In such cases, for example, the history of the viewer's words remains as text data, but this history does not contain any information indicating the broadcaster's words. Also, since the audio data recording the broadcaster's speech does not contain any information indicating the viewer's words, even if the audio data is converted to text, it does not contain any information indicating the viewer's words. As a result, those who want dialogue data have to collect both types of data, compare them, and manually create the dialogue data themselves, which can be a burdensome task.
[0007] Furthermore, this phenomenon was common not only when one party expressed words verbally and the other through text, but also in cases where one party expressed words through sign language and the other through text, or when one party expressed words verbally and the other through sign language, or in any other case where the form of expression of words differed between the two parties.
[0008] In view of the above circumstances, the present invention aims to provide a technology that reduces the burden required to acquire dialogue data.
[0009] One aspect of the present invention is an information processing device comprising a control unit that performs a dialogue data acquisition process to acquire dialogue data, which includes dialogue time series information showing the words indicated by the first timing information and the words indicated by the second timing information in chronological order, and first speaker identification data indicating which of the words shown in the dialogue time series information are the words of the first speaker, based on first information including first text data which is text data indicating words expressed by a first speaker and first timing information indicating the timing at which the first speaker expressed the words, and second information including second text data which is text data indicating words expressed by a second speaker and second timing information indicating the timing at which the second speaker expressed the words, and first information including dialogue time series information showing the words indicated by the first timing information and first speaker identification data indicating which of the words shown in the dialogue time series information are the words of the first speaker.
[0010] One aspect of the present invention is an information processing method executed by an information processing device, comprising a control unit that performs a dialogue data acquisition process, which acquires dialogue data, based on first information including first text data which is text data indicating words expressed by a first speaker and first timing information indicating the timing at which the first speaker expressed the words; second information including second text data which is text data indicating words expressed by a second speaker and second timing information indicating the timing at which the second speaker expressed the words; dialogue time series information which shows the words indicated by the first timing information and the words indicated by the second timing information in chronological order; and first speaker identification data which shows which of the words shown in the dialogue time series information is the words of the first speaker, the information processing method comprising a dialogue data acquisition step in which the control unit performs the dialogue data acquisition process.
[0011] This invention makes it possible to reduce the burden required to acquire dialogue data.
[0012] An explanatory diagram illustrating the information processing device of the embodiment. A diagram showing an example of the hardware configuration of the information processing device of the embodiment. A diagram showing an example of the configuration of the control unit of the embodiment. A flowchart showing an example of the processing flow executed by the information processing device of the embodiment. A diagram showing an example of the configuration of the control unit in a modified example. A diagram showing an example of the configuration of the information processing system in a modified example.
[0013] (Embodiment) Figure 1 is an explanatory diagram illustrating an information processing device 1 of an embodiment. The information processing device 1 includes a control unit 11 which comprises a processor 91 such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or NPU (Neural Network Processing Unit) connected by a bus, and a memory 92. The control unit 11 executes dialogue data acquisition processing.
[0014] The dialogue data acquisition process is a process that acquires dialogue data based on first information and second information. The first information includes first text data and first timing information. The first text data is text data that shows the words expressed by the first speaker. The first timing information shows the timing at which the first speaker expressed the words.
[0015] The second set of information includes second text data and second timing information. The second set of text data is text data that shows the words spoken by the second speaker. The second timing information shows the timing at which the second speaker spoke the words.
[0016] Dialogue data includes dialogue time-series information and first speaker identification data. Dialogue time-series information shows the words indicated by the first timing information and the words indicated by the second timing information in chronological order. First speaker identification data indicates which of the words shown in the dialogue time-series information are the words of the first speaker.
[0017] Furthermore, the words expressed by the first speaker may be expressed in any way. Therefore, the first speaker may express the words by voice, sign language, or text.
[0018] Furthermore, the words expressed by the second speaker may be expressed in any way. Therefore, the second speaker may express the words by voice, sign language, or text. The first speaker and the second speaker express the words in different ways.
[0019] For example, if the words expressed by the first speaker are words expressed by a text, then the first text data is the text data of that text.
[0020] For example, if the words expressed by the first speaker are words expressed through speech, the first text data is text data obtained based on the speech data of that speech. Such first text data may be obtained, for example, by a well-known technique for converting the speech data into text.
[0021] For example, if the words expressed by the first speaker are words expressed through speech, the first text data is text data obtained based on the speech data of that speech. Such first text data may be obtained, for example, by a well-known technique for converting the speech data into text.
[0022] For example, if the words expressed by the first speaker are words expressed through sign language, the first text data is text data obtained based on image data of an image representing that sign language. Such first text data may be obtained, for example, by a technique that generates text data indicating the content of sign language based on image data of an image representing sign language.
[0023] The technology for generating text data representing the content of sign language based on image data of an image representing sign language may be, for example, a technology that executes a pre-trained mathematical model obtained by machine learning, which infers text data based on image data of an image representing sign language. The training of such a mathematical model may be, for example, a process in which image data of an image representing sign language is input into the mathematical model to be trained, the results of the mathematical model's inference based on the image data are compared with the correct data, and the mathematical model to be trained is updated to reduce the difference from the correct data.
[0024] The mathematical model being trained here is a mathematical model that infers text data based on image data of images representing sign language (hereinafter referred to as the "sign language content inference model"). The ground truth data is text data that represents the content of the sign language. The trained sign language content inference model is the result of such training being carried out until predetermined conditions for the termination of training (hereinafter referred to as the "training termination conditions") are met.
[0025] The learning termination condition can be any predetermined condition relating to the termination of learning. For example, the learning termination condition may be that the mathematical model being learned has been updated a predetermined number of times through learning. For example, the learning termination condition may be that the change in the mathematical model being learned after the update is smaller than a predetermined change.
[0026] For example, if the words expressed by the second speaker are words expressed by the text, then the second text data is the text data of that text.
[0027] For example, if the words expressed by the second speaker are words expressed through speech, the second text data is text data obtained based on the speech data of that speech. Such second text data may be obtained, for example, by well-known techniques for converting speech data into text.
[0028] For example, if the words expressed by the second speaker are words expressed through sign language, the second text data is text data obtained based on image data of an image representing that sign language. Such second text data may be obtained by techniques that generate text data representing the content of sign language based on image data of an image representing sign language, such as running the trained sign language content inference model described above.
[0029] The first speaker may be, for example, a video streamer. The video streamer may be, for example, a VTuber. The video streamer may also be, for example, an instructor in e-learning such as online English conversation. In such a case, the second speaker may be, for example, a viewer who watches the video streamed by that video streamer. This viewer expresses their words by, for example, typing text into a device such as their own smartphone or computer. In this case, the viewer's words are expressed, for example, as text.
[0030] The first and second speakers may be, for example, people who are having a video conference using a web conferencing system. In this case, for example, the first speaker may be someone who has difficulty hearing and communicates using sign language. In this case, the second speaker's monitor displays the sign language performed by the first speaker, so the second speaker can understand the words expressed by the first speaker. The second speaker may then communicate with the first speaker by, for example, typing text into their smartphone.
[0031] Figure 1 shows an explanatory diagram illustrating an example of the specific content performed during the dialogue data acquisition process when the first speaker is the video streamer and the second speaker is the viewer.
[0032] The “video streamer” in Figure 1 is an example of the first speaker. The “video viewer” in Figure 1 is an example of the second speaker. Data D101 in Figure 1 is an example of the first information. Data D101 shows text such as “I’m hungry, I want to eat something” and a time such as “11:00”. The text shown in Data D101 is an example of the text shown in the first text data. The time shown in Data D101 is an example of the first timing information.
[0033] Data D102 in Figure 1 is an example of the second type of information. Data D102 shows text such as "What do you want to eat?" and a time such as "11:01". The text shown in Data D102 is an example of the text shown in the second type of text data. The time shown in Data D102 is an example of the second type of timing information.
[0034] Data D103 in Figure 1 is an example of dialogue data. In Data D103 in Figure 1, texts such as "I'm hungry, I want something to eat" (indicated by Data D101) and "What do you want to eat" (indicated by Data D102) are shown in chronological order. In Figure 1, the texts indicated by Data D101 in Data D103 are highlighted. Therefore, Data D103 indicates which words in the dialogue time-series information are spoken by the first speaker by using highlighting. This is a result of Data D103 containing information indicating whether or not a word should be highlighted.
[0035] Thus, the text shown in data D103 is an example of dialogue time-series information, and the information indicating whether or not it is subject to highlighting is an example of first speaker identification data. The highlighting of data D103 is an example of the result of data D103 containing first speaker identification data.
[0036] <Effects of the dialogue data acquisition process> By executing the dialogue data acquisition process, dialogue data is generated based on the first information and the second information. Therefore, those who want dialogue data do not need to compare the first information and the second information and manually create the dialogue data. As a result, the information processing device 1 can reduce the burden required to acquire dialogue data.
[0037] <Example of Hardware Configuration> Figure 2 shows an example of the hardware configuration of the information processing device 1 according to the embodiment. The information processing device 1 includes a control unit 11 which is a control unit equipped with a processor 91 such as a CPU, GPU or NPU connected by a bus and a memory 92, and executes a program. The information processing device 1 functions as a device comprising the control unit 11, interface unit 12 and storage unit 13 by executing the program.
[0038] More specifically, the processor 91 reads the program stored in the storage unit 13 and stores the read program in the memory 92. By executing the program stored in the memory 92, the information processing device 1 functions as a device comprising a control unit 11, an interface unit 12, and a storage unit 13.
[0039] The control unit 11 controls the operation of each functional unit of the information processing device 1. For example, the control unit 11 performs dialogue data acquisition processing.
[0040] The control unit 11 may, for example, perform a first information acquisition process. The first information acquisition process is a process of acquiring first information based on data indicating the words expressed by the first speaker and data indicating the timing of those expressions. The data indicating the words expressed by the first speaker is, for example, the audio data if the first speaker expresses the words by voice, or the image data of the sign language if the first speaker expresses the words by sign language, or the text data indicating the text if the first speaker expresses the words by text.
[0041] The control unit 11 may, for example, execute a second information acquisition process. The second information acquisition process is a process of acquiring second information based on data indicating the words expressed by the second speaker and data indicating the timing of the expression. The data indicating the words expressed by the second speaker is, for example, the audio data of the audio when the second speaker expresses words by voice, and is, for example, the image data of the image indicating the sign language when the second speaker expresses words by sign language. For example, when the second speaker expresses words by text, it is the text data indicating the text.
[0042] The control unit 11 acquires, for example, the information stored in the storage unit 13. The process of acquiring the information stored in the storage unit 13 is specifically reading.
[0043] The interface unit 12 is configured to include a communication interface for connecting the information processing device 1 to an external device. The interface unit 12 communicates with the external device via wired or wireless means.
[0044] The external device is, for example, a device that stores data indicating the words expressed by the first speaker and data indicating the timing of the expression. Such a device may be, for example, a terminal such as a smartphone or a personal computer used by the first speaker, or may be a server used for the video distribution when the first speaker performs video distribution. In such a case, the interface unit 12 acquires data indicating the words expressed by the first speaker and data indicating the timing of the expression by communicating with such an external device.
[0045] The external device is, for example, a device that stores data indicating the words expressed by the second speaker and data indicating the timing of the expression. Such a device may be, for example, a terminal such as a smartphone or a personal computer used by the second speaker, or may be a server used for the video distribution when the second speaker performs video distribution. In such a case, the interface unit 12 acquires data indicating the words expressed by the second speaker and data indicating the timing of the expression by communicating with such an external device.
[0046] The external device may be, for example, a device that is the output destination of the dialogue data. The device that is the output destination of the dialogue data may be, for example, a device that generates an NPC (non player character) that reproduces the personality of the video distributor using, for example, a personality extraction technique or a personality reproduction dialogue technique. In such a case, the interface unit 12 outputs the dialogue data to such an external device through communication with such an external device.
[0047] Note that the interface unit 12 may be configured to include input devices such as a mouse, a keyboard, a touch panel, etc. The interface unit 12 may be configured as an interface that connects these input devices to the information processing device 1. In this way, the input devices of the interface unit 12 receive the input of various information or signals to the information processing device 1 via wired or wireless means. Note that the information or signal does not necessarily have to be input to the communication interface of the interface unit 12 and may be input to the input device of the interface unit 12.
[0048] The interface unit 12 outputs various information, for example. The interface unit 12 is configured to include display devices such as a CRT (Cathode Ray Tube) display, a liquid crystal display, an organic EL (Electro-Luminescence) display, etc., and a speaker. The interface unit 12 may be configured as an interface that connects these display devices or speakers to the information processing device 1. Therefore, the interface unit 12 may output, for example, the information indicated by the information or signal input to the input device of the interface unit 12 as an image or sound.
[0049] The storage unit 13 is configured using a computer-readable recording medium such as a magnetic hard disk drive or a semiconductor memory device. The storage unit 13 stores various information related to the information processing device 1. The storage unit 13 stores various information generated by the operation of the control unit 11, for example. Therefore, the storage unit 13 may store, for example, dialogue data. The storage unit 13 may also store, for example, first information or second information. The storage unit 13 may also store data indicating the words expressed by the first speaker and information indicating the timing at which those words were expressed. The storage unit 13 may also store information indicating the words expressed by the second speaker and the timing at which those words were expressed. The storage unit 13 may reside, for example, on the cloud.
[0050] Figure 3 shows an example of the configuration of the control unit 11 of the embodiment. The control unit 11 includes a first information acquisition unit 111, a second information acquisition unit 112, a dialogue data acquisition unit 113, and a learning unit 114. The first information acquisition unit 111 performs a first information acquisition process. The second information acquisition unit 112 performs a second information acquisition process. The dialogue data acquisition unit 113 performs a dialogue data acquisition process.
[0051] Figure 4 is a flowchart showing an example of the processing flow executed by the information processing device 1 of the embodiment. The control unit 11 acquires data indicating the words expressed by the first speaker and data indicating the timing of those expressions (step S101). Next, the control unit 11 executes the first information acquisition process (step S102). By executing the first information acquisition process, first information is obtained based on the information obtained in step S101.
[0052] Next, the control unit 11 acquires data indicating the words spoken by the second speaker and data indicating the timing of those words (step S103). Next, the control unit 11 executes a second information acquisition process (step S104). The execution of the second information acquisition process yields second information based on the information obtained in step S103.
[0053] Next, the control unit 11 executes the dialogue data acquisition process (step S105). By executing the dialogue data acquisition process, dialogue data is obtained based on the first information obtained in step S102 and the second information obtained in step S104.
[0054] Furthermore, the process in step S102 may be executed at any timing as long as it is performed after the execution of the process in step S101 and before the execution of step S105. Similarly, the process in step S104 may be executed at any timing as long as it is performed after the execution of the process in step S103 and before the execution of step S105. Moreover, the process in step S103 does not necessarily have to be executed after the execution of the process in step S101; it may be executed in parallel with step S101 or before the execution of step S101.
[0055] The information processing device 1 configured in this way includes a control unit 11 that performs dialogue data acquisition processing. Therefore, as described in <Effects of Dialogue Data Acquisition Processing>, the burden required to acquire dialogue data can be reduced.
[0056] (Variation) The dialogue data obtained by executing the dialogue data acquisition process may be used to train the mathematical model to be trained. The mathematical model to be trained may be, for example, the mathematical model in the personality extraction technology or personality reproduction dialogue technology described above. More specifically, such a mathematical model is a mathematical model that infers words that express the personality of the video streamer (hereinafter referred to as the "personality reproduction dialogue model").
[0057] Figure 5 shows an example of the configuration of the control unit 11a in a modified example. The control unit 11a is a control unit provided in the information processing device 1 and differs from the control unit 11 in that it includes a learning unit 114. In other words, the information processing device 1 may include the control unit 11a instead of the control unit 11.
[0058] The learning unit 114 uses dialogue data to train the mathematical model to be trained. The training may be, for example, further training of a mathematical model that has already been trained. The training may be, for example, transfer learning or fine tuning. Here, the mathematical model to be trained may be, for example, a personalized dialogue model.
[0059] The personalization dialogue model obtained by the control unit 11a may be used in technologies for constructing virtual spaces such as the metaverse. An example of the processing performed when constructing such a virtual space will be explained using Figure 6.
[0060] Figure 6 shows an example of the configuration of an information processing system 100 in a modified example. The information processing system 100 comprises an information processing device 1 equipped with a control unit 11a, and an external device 9 that is communicatively connected to the information processing device 1. The external device 9 is an example of an external device.
[0061] The external device 9 includes a control unit 910. The control unit 910 is a control unit that includes a processor (not shown) such as a CPU, GPU, or NPU connected by a bus, and memory (not shown). The control unit 910 controls the operation of each functional part of the external device 9.
[0062] The control unit 910 includes a character selection unit 911, an interaction space generation unit 912, a user input unit 913, a model selection unit 914, and an interaction data generation unit 915. The character selection unit 911 acquires information indicating the person the user wants to interact with in the virtual space. The person the user wants to interact with is, for example, a VTuber.
[0063] The interaction space generation unit 912 generates a virtual space for interaction between the user and the interaction partner selected by the user. The virtual space is generated independently for each user. Note that the same character (such as a VTuber) may appear in the interaction spaces of other users simultaneously.
[0064] The user input unit 913 receives voice or text input data transmitted from the user to the interaction partner. If the transmitted data is voice data, the user input unit 913 converts the input voice data into text and outputs it to the dialogue data generation unit 915.
[0065] The model selection unit 914 selects an individuality-representing dialogue model that has learned the individuality of the interaction partner indicated by the information acquired by the character selection unit 911. An individuality-representing dialogue model that has learned the individuality of the interaction partner is an individuality-representing dialogue model that has been trained using dialogue data that shows the words of that interaction partner.
[0066] The dialogue data generation unit 915 generates dialogue data by inputting the user's input text information obtained from the user input unit 913 into the personality-reproducing dialogue model selected by the model selection unit 914, and outputs this as the dialogue data of the person being interacted with on the interaction space.
[0067] The external device 9 also includes a storage unit for storing multiple personalization dialogue models. In the example in Figure 7, there are three personalization dialogue models: Personalization Dialogue Model A, Personalization Dialogue Model B, and Personalization Dialogue Model C, each of which is stored in a storage unit. In the example in Figure 7, the external device 9 includes storage units 916, 917, and 918. Storage unit 916 stores Personalization Dialogue Model A, storage unit 917 stores Personalization Dialogue Model B, and storage unit 918 stores Personalization Dialogue Model C. Each storage unit is configured using a computer-readable recording medium such as a magnetic hard disk drive or a semiconductor memory device.
[0068] The information processing device 1 may be implemented using multiple information processing devices connected to each other via a network. In this case, each process executed by the control unit 11 or control unit 11a may be performed in a distributed manner by multiple information processing devices.
[0069] Furthermore, all or part of the functions of the information processing device 1 may be implemented using hardware such as ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), or FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Computer-readable recording media include, for example, portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may also be transmitted via a telecommunications line.
[0070] While embodiments of this invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments and includes designs and the like that do not depart from the spirit of this invention.
[0071] 1... Information processing device, 11, 11a... Control unit, 12... Interface unit, 13... Storage unit, 91... Processor, 92... Memory, 100... Information processing system, 9... External device
Claims
1. An information processing apparatus comprising: a control unit that performs a dialogue data acquisition process to acquire dialogue data, which includes:
1. First information including first text data which is text data indicating words expressed by a first speaker and first timing information indicating the timing at which the first speaker expressed the words; and second information including second text data which is text data indicating words expressed by a second speaker and second timing information indicating the timing at which the second speaker expressed the words; and second information including dialogue time series information which shows the words indicated by the first timing information and the words indicated by the second timing information in chronological order, and first speaker identification data which shows which of the words shown in the dialogue time series information is the words of the first speaker.
2. The information processing apparatus according to claim 1, wherein the words expressed by the first speaker are words expressed by voice, and the first text data is text data obtained based on the voice data of the voice.
3. The information processing apparatus according to claim 1, wherein the words expressed by the first speaker are words expressed in sign language, and the first text data is text data obtained based on image data of an image representing the sign language.
4. An information processing method executed by an information processing device comprising: first information including first text data which is text data indicating words expressed by a first speaker and first timing information indicating the timing at which the first speaker expressed the words; second information including second text data which is text data indicating words expressed by a second speaker and second timing information indicating the timing at which the second speaker expressed the words; dialogue data including dialogue time series information which shows the words indicated by the first timing information and the words indicated by the second timing information in chronological order and first speaker identification data which shows which of the words shown in the dialogue time series information is the words of the first speaker, wherein the information processing method comprises: a dialogue data acquisition step in which the control unit performs the dialogue data acquisition process.
Citation Information
Patent Citations
Interactive text summarization apparatus and method
JP2017111190A
Method for displaying received messages, message application program, and mobile terminal
JP2018185816A
Dialogue assistance
JP2019502944A