Response system and response method

The response system generates natural responses by integrating input voice with previously played content, addressing the lack of context consideration in conventional systems.

JP2026066591APending Publication Date: 2026-04-17HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HONDA MOTOR CO LTD
Filing Date
2024-10-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Conventional response systems fail to consider background information when generating responses to input voices, resulting in unnatural interactions.

Method used

A response system and method that includes a content playback unit, microphone, and response generation unit to generate responses based on both input voice and previously played content, using machine learning models to enhance natural dialogue.

Benefits of technology

Enables the generation of natural responses by considering the context of the content being viewed, improving the interaction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026066591000001_ABST
    Figure 2026066591000001_ABST
Patent Text Reader

Abstract

To enable natural responses to input speech. [Solution] The response system (1000) comprises a content playback unit (12, 12A~12D, 13, 13A~13D) that plays content (C), a microphone (11, 11A~11D), an input voice recognition unit (113) that recognizes the input voice to the microphone (11, 11A~11D), and a response generation unit (212) that generates a response sentence (R) to the input voice based on the input voice and the content (C) that the content playback unit (12, 12A~12D, 13, 13A~13D) has played up to the time of input voice to the microphone (11, 11A~11D).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a response system and a response method.

Background Art

[0002] Due to the recent evolution of learning models using machine learning, etc., technologies for appropriately responding to inputs using natural language have been realized. For example, Patent Document 1 discloses a technique for increasing the correct answer probability of answers to question queries composed of natural language texts.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, due to the evolution of speech recognition engines, etc., technologies that use the voice uttered by a person as an input using natural language have been developed. Such technologies are applied to, for example, response systems that output a sentence that is a response to the content uttered by a person in voice or text, etc., and conduct a dialogue with a person. However, conventional response systems have not been able to make responses taking into account background information such as the content being viewed by a person, and there has been room for improvement to make more natural responses. The present invention has been made in view of the above circumstances, and an object thereof is to enable a natural response to an input voice.

Means for Solving the Problems

[0005] One aspect of the present invention is a response system comprising: a content playback unit for playing content; a microphone; an input voice recognition unit for recognizing input voice to the microphone; and a response generation unit for generating a response statement to the input voice based on the input voice and the content played by the content playback unit up to the time the input voice was input to the microphone.

[0006] Another aspect of the present invention is a response method comprising: a content playback unit plays content; an input voice recognition unit recognizes the voice input to the microphone; and a response generation unit generates a response statement to the input voice based on the input voice and the content played by the content playback unit up to the time the input voice was input to the microphone. [Effects of the Invention]

[0007] According to one aspect of the present invention, a response system, and according to another aspect, a response sentence can be generated for input speech by taking into account the content played back by the content playback unit. Therefore, a natural response to input speech can be provided. [Brief explanation of the drawing]

[0008] [Figure 1] A diagram showing the configuration of the response system according to the first embodiment. [Figure 2] A diagram showing the configuration of the response device and response generation server. [Figure 3] A flowchart illustrating the operation of the response device and the response generation server. [Figure 4] A flowchart illustrating the operation of the response generation unit. [Figure 5] A flowchart illustrating the operation of the response generation unit. [Modes for carrying out the invention]

[0009] [1. First Embodiment] First, the first embodiment will be described with reference to the drawings.

[0010] [1-1. Overall Configuration of the Response System] Figure 1 is a diagram showing the configuration of the response system 1000 according to the first embodiment. The response system 1000 is a system that takes into account the content of content C being played inside the vehicle 1 and outputs a response sentence R in response to the text contained in the input voice. The response sentence R output by the response system 1000 may be a sentence that indicates any response to the content of the input voice, such as an answer, agreement, or denial. Furthermore, the response sentence R output by the response system 1000 may be output in any way, such as by displaying text or by voice.

[0011] As shown in Figure 1, the response system 1000 comprises a vehicle 1, a response generation server 200, and a content distribution server 300. The number of vehicles 1, response generation servers 200, and content distribution servers 300 in the response system 1000 can be arbitrarily configured.

[0012] The response generation server 200 is a server device that generates the response statement R that the response device 100 outputs using the display 12 or speaker 13. The response generation server 200 is connected to the communication network NW and communicates with the vehicle 1.

[0013] The content distribution server 300 is a server device that distributes content C. The content distribution server 300 is connected to a communication network NW and communicates with vehicle 1. The communication network NW consists of a public telephone network, a dedicated line, or other communication circuits. The response system 1000 may have a content distribution server 300 for each content distribution source. Alternatively, the content distribution server 300 may be a server device that aggregates content C distributed by multiple content distribution sources and distributes it to vehicle 1.

[0014] The content C distributed to vehicle 1 by the content distribution server 300 includes text in either audio or text format. The text included in content C will be referred to as content text TXC below. Content text TXC may consist of one or more sentences. In this embodiment, each content text TXC consists of one sentence. Content C may include one or more content texts TXC.

[0015] Content C may be, for example, video content such as a movie or video program, or audio content such as an audio program. Content text TXC may be, for example, audio spoken by performers in the video content or audio content, or text such as captions or subtitles displayed in the video content. Note that Content C played back in vehicle 1 is not limited to content distributed from content distribution server 300. For example, Content C may be broadcast via radio or television, or read from a storage medium brought into vehicle 1.

[0016] The vehicle 1 illustrated in Figure 1 is a four-wheeled vehicle. Vehicle 1 is equipped with seats 10 for occupants P, namely a driver's seat 10A, a passenger seat 10B, a rear right seat 10C, and a rear left seat 10D. In Figure 1, vehicle 1 shows occupant P1, who is the driver, seated in the driver's seat 10A. In Figure 1, vehicle 1 also shows occupant P2, who is a passenger, seated in the passenger seat 10B. In Figure 1, vehicle 1 also shows occupant P3, who is a passenger, seated in the rear right seat 10C. In Figure 1, vehicle 1 also shows occupant P4, who is a passenger, seated in the rear left seat 10D.

[0017] Vehicle 1 is equipped with a response device 100. The response device 100 is configured to acquire input audio as audio data using a microphone 11 installed inside Vehicle 1. The response device 100 is also configured to output a response sentence R as text or audio using at least one of a display 12 or a speaker 13 installed inside Vehicle 1.

[0018] The microphone 11 is a device that receives voice input. In the present embodiment, in the vehicle 1, as the microphone 11, a driver's seat microphone 11A, a passenger seat microphone 11B, a rear right seat microphone 11C, and a rear left seat microphone 11D are provided. The driver's seat microphone 11A mainly records the voice spoken by the passenger P1 sitting on the driver's seat 10A. The passenger seat microphone 11B mainly records the voice spoken by the passenger P2 sitting on the passenger seat 10B. The rear right seat microphone 11C mainly records the voice spoken by the passenger P3 sitting on the rear right seat 10C. The rear left seat microphone 11D mainly records the voice spoken by the passenger P4 sitting on the rear left seat 10D. That is, in the present embodiment, the response device 100 can record the voices spoken by the passengers P1 to P4 sitting on each of the seats 10A to 10D, distinguishing the speaking passengers P1 to P4. The driver's seat microphone 11A, the passenger seat microphone 11B, the rear right seat microphone 11C, and the rear left seat microphone 11D correspond to the microphones of the present disclosure.

[0019] The display 12 is a device that outputs characters or images. In the present embodiment, in the vehicle 1, as the display 12, a center display 12A, a passenger seat display 12B, a rear right seat display 12C, and a rear left seat display 12D are provided. The center display 12A mainly displays characters or images to the passenger P1 sitting on the driver's seat 10A. The passenger seat display 12B mainly displays characters or images to the passenger P2 sitting on the passenger seat 10B. The rear right seat display 12C mainly displays characters or images to the passenger P3 sitting on the rear right seat 10C. The rear left seat display 12D mainly displays characters or images to the passenger P4 sitting on the rear left seat 10D. That is, in the present embodiment, the response device 100 is configured to be able to display characters or images to all or any selected part of the passengers P1 to P4 sitting on each of the seats 10A to 10D. The displays 12, 12A to 12D correspond to the content reproduction units of the present disclosure.

[0020] The speaker 13 is a device that outputs sound. In the present embodiment, in the vehicle 1, as the speaker 13, a center speaker 13A, a passenger seat speaker 13B, a rear right seat speaker 13C, and a rear left seat speaker 13D are provided. The center speaker 13A mainly outputs sound to the passenger P1 sitting in the driver's seat 10A. The passenger seat speaker 13B mainly outputs sound to the passenger P2 sitting in the passenger seat 10B. The rear right seat speaker 13C mainly outputs sound to the passenger P3 sitting in the rear right seat 10C. The rear left seat speaker 13D mainly outputs sound to the passenger P4 sitting in the rear left seat 10D. That is, in the present embodiment, the response device 100 is configured to be able to output sound to all or a part of the passengers P1 to P4 sitting in each seat 10A to 10D selected arbitrarily. The speakers 13, 13A to 13D correspond to the content reproduction unit of the present disclosure.

[0021] [1-2. Configuration of Response Device] Next, the configuration of the response device 100 will be described. FIG. 2 is a diagram showing the configuration of the response device 100 and the response generation server 200.

[0022] The response device 100 is connected to the microphones 11A to 11D, the displays 12A to 12D, and the speakers 13A to 13D provided in the vehicle 1. Note that the devices connected to the response device 100 are not limited to these devices, and other types of devices may be connected. Further, the response device 100 may be configured to include a microphone 11, a display 12, a speaker 13, or other types of devices.

[0023] The response device 100 is a control unit comprising a first processor 110, a first memory 120, and a first communication unit 130. The first processor 110 comprises a processor such as a CPU (Central Processing Unit) or an MPU (Micro Processor Unit). The first memory 120 is a storage device for storing programs and data, and comprises, for example, ROM (Read Only Memory) or RAM (Random Access Memory). The first communication unit 130 comprises hardware conforming to a predetermined communication standard, such as a wireless communication circuit. The response device 100 communicates with the response generation server 200 and the content distribution server 300 via a communication network NW using the first communication unit 130.

[0024] The first memory 120 stores the first control program 121, which is a program for controlling the response device 100. The first processor 110 reads and executes the first control program 121 and functions as the first communication control unit 111, input / output control unit 112, input voice recognition unit 113, and content recognition unit 114.

[0025] The first communication control unit 111 communicates with the response generation server 200 and the content distribution server 300 via the communication network NW using the first communication unit 130.

[0026] The input / output control unit 112 uses the microphone 11 as an input device to acquire input audio as audio data. In this embodiment, the input / output control unit 112 identifies which microphone 11A to 11D recorded the input audio recorded by each microphone 11A to 11D and acquires it accordingly. The input / output control unit 112 uses any of the display 12 and speaker 13 as output devices to output the response message received by the first communication control unit 111 from the response generation server 200. The input / output control unit 112 also uses any of the display 12 and speaker 13 as output devices to output content C received by the first communication control unit 111 from the content distribution server 300 in the form of audio or video.

[0027] The input speech recognition unit 113 recognizes the input speech to microphones 11 and 11A-11D. More specifically, the input speech recognition unit 113 converts the text contained in the input speech, which is acquired as audio data by the input / output control unit 112, into text data through speech recognition. The text contained in the input speech is a sentence spoken by one of the crew members P1-P4. Hereinafter, the text contained in the input speech will be referred to as the utterance text TXU. The utterance text TXU converted into text data is sent to the response generation server 200 and used to generate the response sentence R. The input audio input to microphones 11, 11A to 11D may be one or more. Each input audio may contain one or more utterances TXU. Each utterance TXU may contain one or more sentences. In this embodiment, the input audio contains one utterance TXU. In this embodiment, one utterance TXU consists of one sentence.

[0028] In this embodiment, the input speech recognition unit 113 generates speech time information DTU and speaker information IFP along with the utterance text TXU. The speech time information DTU is information indicating the time when the utterance text TXU included in the input speech was spoken and input to the microphone 11, i.e., the speech time.

[0029] The speaker information IFP is information that allows the identification of the crew member P who uttered the spoken sentence TXU from among crew members P1 to P4. For example, the input speech recognition unit 113 may generate the speaker information IFP by determining which of the microphones 11A to 11D was used to record the spoken sentence TXU, and then estimating that the main recording target crew member P of the identified microphones 11A to 11D was the one who uttered the sentence. Alternatively, the input speech recognition unit 113 may generate the speaker information IFP by analyzing the voiceprint of the spoken sentence TXU contained in the input speech.

[0030] The content recognition unit 114 converts the content text TXC contained in content C within the vehicle 1 into text data. The content text TXC converted into text data is sent to the response generation server 200 and used to generate the response text R.

[0031] For example, if content C is video content or audio content, the content recognition unit 114 applies speech recognition to the content text TXC, which is audio data included in content C, and converts it into text data. Alternatively, if content C is video content, the content recognition unit 114 may acquire the content text TXC, which is text data attached to content C as subtitles. Furthermore, the content recognition unit 114 may be configured to apply image recognition or the like to the content text TXC, which is included in the image data of content C as subtitles, captions, etc., and convert it into text data.

[0032] Furthermore, the content recognition unit 114 generates playback time information DTC. The playback time information DTC is information about the time when the content text TXC was played in content C.

[0033] Furthermore, the content recognition unit 114 may be configured to convert content text TXC into text data even when the content C played in the vehicle 1 is played without going through the response device 100. For example, the content recognition unit 114 may acquire audio data of content C, which is video content or audio content, played in the vehicle 1, via the input / output control unit 112 and the microphone 11. Alternatively, the content recognition unit 114 may convert the content text TXC included in the acquired audio data into text data by applying speech recognition.

[0034] Furthermore, the first memory 120 stores content data 122. The content data 122 is a table having records that include content text TXC and playback time information DTC as text data generated by the content recognition unit 114. The content data 122 is updated each time the content recognition unit 114 generates content text TXC and playback time information DTC to include the generated pair of content text TXC and playback time information DTC as a record.

[0035] [1-3. Configuration of the response generation server] Next, we will describe the configuration of the response generation server 200. The response device 100 is a control unit comprising a second processor 210, a second memory 220, and a second communication unit 230. The second processor 210 comprises a processor such as a CPU or MPU. The second memory 220 is a storage device for storing programs and data, and includes, for example, ROM or RAM. The second communication unit 230 comprises hardware conforming to a predetermined communication standard, such as a wireless communication circuit. The response generation server 200 communicates with the response device 100 via the communication network NW using the second communication unit 230.

[0036] The second memory 220 stores the second control program 221, which is a program for controlling the response generation server 200. The second processor 210 functions as the second communication control unit 211 and the response generation unit 212 by reading and executing the second control program 221.

[0037] The second communication control unit 211 communicates with the response device 100 via the communication network NW using the second communication unit 230.

[0038] The response generation unit 212 uses the utterance text TXU and content text TXC received by the second communication unit 230 from the response device 100 to generate a response text R for the utterance text TXU spoken by crew member P. More specifically, the response generation unit 212 further functions as a response text generation unit 213 and an input data generation unit 214.

[0039] The response sentence generation unit 213 inputs input data to the response generation model 222 stored in the second memory 220 and generates a response sentence. The response generation model 222 is a model that takes either a spoken sentence TXU (text data) or a spoken sentence TXU and a content sentence TXC (text data) as input data and outputs a response sentence R for the spoken sentence TXU. The second memory 220 may store both a model that generates a response sentence using only the spoken sentence TXU as input data, and a model that uses both the spoken sentence TXU and the content sentence TXC as input data. The response generation model 222 is, for example, a trained model using machine learning.

[0040] The input data generation unit 214 uses the utterance text TXU and content text TXC received from the response device 100 to generate input data to be input to the response generation model 222. Details of the operation of the input data generation unit 214 will be described later.

[0041] [1-4. Overview of the response system's operation] Next, we will explain the operation of the response system 1000. First, we will provide an overview of the operation of the response system 1000.

[0042] Figure 3 is a flowchart illustrating the operation of the response device 100 and the response generation server 200, showing the process from when the response device 100 responds to a person's speech until it outputs a response sentence R. In Figure 3, flowchart FA shows the operation of the response device 100, and flowchart FB shows the operation of the response generation server 200. The operation in Figure 3 is triggered, for example, by the operation of crew member P, which turns on the power of the response device 100.

[0043] First, in step SA1, the input / output control unit 112 of the response device 100 starts playing the content C received by the first communication control unit 111. At the same time, the input / output control unit 112 also starts acquiring audio via the microphone 11.

[0044] Next, in step SA2, the content recognition unit 114 converts the content text TXC of the content C being played in vehicle 1 into text data. At this time, the content recognition unit 114 generates playback time information DTC corresponding to the converted content text TXC. In addition, each time the content recognition unit 114 generates the content text TXC and playback time information DTC as text data, it adds the generated pair of content text TXC and playback time information DTC to the content data 122. In this way, the content recognition unit 114 updates the content data 122.

[0045] In this embodiment, the content recognition unit 114 is configured to delete records containing playback time information DTCs that correspond to a point in time earlier than a predetermined time from the current time from the content data 122. The predetermined time is, for example, 1 minute. The content recognition unit 114 may also be configured to delete older records in order when updating the content data 122, so that the number of records in the content data 122 does not exceed a predetermined number. The predetermined number is, for example, 10.

[0046] Next, in step SA3, the input speech recognition unit 113 determines whether the input / output control unit 112 has acquired input speech including the spoken text TXU via the microphone 11. In step SA3, if the input speech recognition unit 113 determines that the input / output control unit 112 has not acquired input speech including the spoken text TXU (step SA3: NO), the operation of the response device 100 returns to step SA2. In step SA3, if the input speech recognition unit 113 determines that the input / output control unit 112 has acquired input speech including the spoken text TXU (step SA3: YES), the operation of the response device 100 proceeds to step SA4.

[0047] In step SA4, the input speech recognition unit 113 converts the utterance text TXU of the input speech acquired by the input / output control unit 112 into text data. In addition, along with generating the utterance text TXU as text data, the input speech recognition unit 113 also generates utterance time information DTU and speaker information IFP. The number of utterance time information DTU and speaker information IFP is the same as the number of utterance text TXU generated as text data.

[0048] Next, in step SA5, the first communication control unit 111 transmits the utterance text TXU converted in step SA4 and the content data 122 stored in the first memory 120 to the response generation server 200. At this time, the first communication control unit 111 transmits both the utterance time information DTU and the speaker information IFP to the response generation server 200, associated with each utterance text TXU.

[0049] Next, in step SB1, the second communication control unit 211 of the response generation server 200 receives the transmitted utterance text TXU, utterance time information DTU, speaker information IFP, and content data 122.

[0050] Next, in step SB2, the response generation unit 212 generates a response sentence R based on the received speech sentence TXU and content data 122. In this embodiment, in step SB2, output destination information is also generated to specify which of the crew members P1 to P4 the response sentence R should be output to. The output destination information may be, for example, information that identifies one speaker 13 from speakers 13A to 13D to output the response sentence R, or information that identifies one display 12 from displays 12A to 12D to output the response sentence R. The output destination information is generated based on speaker information IFP corresponding to the speech sentence TXU that is the target of the response sentence R. For example, the output destination information may be information that sets the output destination of the response sentence R to crew member P who spoke the speech sentence TXU indicated by the speaker information IFP. Details of step SB2 will be described later.

[0051] Next, in step SB3, the second communication control unit 211 transmits the generated response message R to the response device 100. At this time, the second communication control unit 211 also transmits the output destination information.

[0052] Next, in step SA6, the first communication control unit 111 of the response device 100 receives the transmitted response message R and the output destination information.

[0053] Next, in step SA7, the input / output control unit 112 outputs the response sentence R received by the first communication control unit 111 via an optional output device such as the display 12 or speaker 13. In this embodiment, in step SA7, the input / output control unit 112 outputs the response sentence R as sound using the speaker 13. With the execution of step SA7, an appropriate response sentence R is given to the spoken sentence TXU uttered by the crew member P, and the operation shown in Figure 3 is completed.

[0054] Furthermore, in this embodiment, the input / output control unit 112 can refer to the received output destination information and output the response message R to one or more targets among the crew members P1 to P4 that are identified by the output destination information. For example, the input / output control unit 112 can refer to the output destination information and output the response message R to crew member P1 by outputting the response message R using the speaker 13A or the display 12A.

[0055] Furthermore, as described later, if multiple response sentences R are generated in step SB2 for multiple occupant P's spoken sentences TXU, in step SA7, the input / output control unit 112 may change the order in which the response sentences R are output according to the output destination information. For example, when a response sentence R is generated for each of occupant P1 to P4, the input / output control unit 112 may refer to the output destination information for each response sentence R and output the response sentences R in the order of occupant P sitting in the driver's seat 10A, passenger seat 10B, rear right seat 10C, and rear left seat 10D. In addition, if multiple response sentences R are generated, the output order of the multiple response sentences R may be arbitrarily determined based on the positional relationship of the seats 10 in which the occupant P who uttered the spoken sentence TXU that is the target of the response sentence R is seated.

[0056] [1-5. Details of the operation of the response generation unit] Next, we will describe the details of the operation of the response generation unit 212 in step SB2. The operation in step SB2 is divided into two patterns depending on whether or not there is only one utterance sentence TXU as text data received in step SB1 in Figure 3. This pattern is determined by the response generation unit 212 based on the number of utterance sentences TXU received in step SB1. As mentioned above, in this embodiment, one input audio contains one utterance sentence TXU. Therefore, the operation in step SB2 can be divided into two patterns: when one input audio is recognized to the microphone 11 by the input audio recognition unit 113, and when multiple input audio is recognized to the microphone 11 by the input audio recognition unit 113.

[0057] [1-5-1. Actions when there is only one spoken sentence] The following describes the operation in step SB2 when only one utterance text TXU is received in step SB1. Figure 4 is a flowchart showing the operation of the response generation unit 212, detailing the operation of step SB2 when there is one received utterance sentence TXU.

[0058] At the beginning of step SB2, in step SB201, the input data generation unit 214 determines whether it can extract content text TXC that is included in the input audio's related content from all content text TXC that were played back after a predetermined time prior to the utterance time of the utterance text TXU. The related content of the input audio refers to the content C that is related to the input audio. In other words, the content text TXC of the related content is related to the utterance text TXU that is included in the input audio. The predetermined time here is, for example, 30 seconds. In other words, in this embodiment, the response generation unit 212 determines whether the content C played by the speaker 13 and the display 12 after a predetermined time prior to the time when the input audio was input to the microphone 11 is related to the input audio. If it is determined that the content C is related to the input audio, the response generation unit 212 determines whether it can extract a content document TXC from the content C.

[0059] Furthermore, for an utterance TXU and a content document TXC to be related, this includes the utterance TXU being a response to the content document TXC. For an utterance TXU to be a response to the content document TXC, this includes, for example, the content of the utterance TXU being about a topic similar to the topic of the content document TXC, or being an opinion, affirmation, denial, or other reaction or response to the content document TXC.

[0060] In this embodiment, in detail, the input data generation unit 214 performs the determination in step SB201 by the following process. The input data generation unit 214 first refers to the speech time information DTU corresponding to the speech text TXU and identifies the speech time of the speech text TXU. Next, the input data generation unit 214 extracts all content texts TXC that were played back before the identified speech time and from a predetermined time before the identified speech time onward, by referring to the playback time information DTC. Then, the input data generation unit 214 determines whether it is possible to further extract content texts TXC related to the speech text TXU from the extracted content texts TXC.

[0061] In contrast to this embodiment, in step SB201, the input data generation unit 214 may be configured to determine whether the content C corresponding to a content document TXC is related content, targeting content documents TXC from a predetermined position onward from the last content document TXC when the speaker 13 and display 12 play content C containing multiple content documents TXC. The predetermined position here is, for example, the 5th document. In this case, more specifically, the input data generation unit 214 performs the determination in step SB201 by the following process. The input data generation unit 214 first refers to the speech time information DTU corresponding to the speech sentence TXU and identifies the speech time of the speech sentence TXU. Next, the input data generation unit 214 extracts all content sentences TXC that were played back before the identified speech time by referring to the playback time information DTC. Furthermore, the input data generation unit 214 extracts content sentences TXC from the extracted content sentences TXC, starting with the most recent playback time up to a predetermined number. Finally, the input data generation unit 214 determines whether one or more content sentences TXC related to the speech sentence TXU can be extracted from the extracted content sentences TXC up to the predetermined number.

[0062] The input data generation unit 214 may determine whether the content text TXC and the spoken text TXU are related by applying natural language processing to the content text TXC and the spoken text TXU. For natural language processing, for example, a trained model using machine learning may be used. Specifically, for example, the input data generation unit 214 vectorizes the words contained in the content text TXC and the words contained in the spoken text TXU using an arbitrary pre-trained model. Then, if there exists a combination in which the cosine similarity between the vectorized words exceeds a predetermined value, it may be determined that the content text TXC and the spoken text TXU are related. Alternatively, any other method may be used to determine whether or not the content text TXC and the spoken text TXU are related.

[0063] If the input data generation unit 214 determines in step SB201 that the content text TXC of the related content can be extracted (step SB201: YES), the response generation unit 212 proceeds to step SB202.

[0064] In step SB202, the input data generation unit 214 determines whether or not there are multiple content documents TXC extracted in step SB201. In step SB202, if the input data generation unit 214 determines that there are multiple content documents TXC extracted in step SB201 (step SB202: YES), the response generation unit 212 proceeds to step SB203. In step SB202, if the input data generation unit 214 determines that there is only one content document TXC extracted in step SB201 (step SB202: NO), the response generation unit 212 proceeds to step SB204.

[0065] In step SB203, the input data generation unit 214 refers to the playback time information DTC and extracts one content document TXC from the multiple content documents TXC extracted in step SB201 whose playback time is closest to the utterance time of the spoken document TXU.

[0066] Next, in step SB204, the input data generation unit 214 generates input data based on one content document TXC and an utterance document TXU extracted in step SB203 or step SB204. The input data is, for example, the content document TXC and utterance document TXU converted from text data into a format that can be input to the response generation model 222.

[0067] Next, in step SB205, the response sentence generation unit 213 inputs the input data generated in step SB204 into the response generation model 222 and generates a response sentence R corresponding to the utterance sentence TXU and the content sentence TXC. After that, the process of step SB2 shown in Figure 4 is completed and the process moves to step SB3 shown in Figure 3. In other words, the response generation unit 212 generates a response statement R for the input audio based on the input audio and the content C that the speaker 13 and display 12 have played up to the time the input audio is received by the microphone 11. Furthermore, the response generation unit 212 determines whether the content C played by the display 12 and speaker 13 up to the time of input audio to the microphone 11 is related content associated with the utterance text TXU contained in the input audio. If the content C played by the display 12 and speaker 13 is related content, the response generation unit 212 generates a response sentence R based on the utterance text TXU contained in the input audio and the content text TXC contained in the related content. Furthermore, in this embodiment, if the content C played by the display 12 and speaker 13 up to the time of input of the audio to the microphone 11 includes a plurality of content sentences TXC related to the utterance sentence TXU included in the input audio, the response generation unit 212 generates a response sentence R based on the utterance sentence TXU included in the input audio and the content sentence TXC played by the display 12 and speaker 13 at the time closest to the time of input of the utterance sentence TXU included in the input audio.

[0068] Furthermore, if the input data generation unit 214 determines in step SB201 that it cannot extract the content text TXC associated with the utterance text TXU (step SB201: NO), the response generation unit 212 proceeds to step SB206.

[0069] In step SB206, the input data generation unit 214 generates input data based on the utterance text TXU, without using the content text TXC. The input data is, for example, the utterance text TXU converted into a format that can be input to the response generation model 222.

[0070] In step SB207, the response sentence generation unit 213 inputs the input data generated in step SB206 into the response generation model 222 and generates a response sentence R corresponding to the utterance sentence TXU. After that, the process of step SB2 shown in Figure 4 is completed and the process moves to step SB3 shown in Figure 3. In other words, the response generation unit 212 determines whether the content C played by the display 12 and speaker 13 up to the time of input audio to the microphone 11 is related content associated with the utterance TXU contained in the input audio. If the content C played by the display 12 and speaker 13 is not related content, the response generation unit 212 generates a response sentence R based on the utterance TXU contained in the input audio.

[0071] As shown in steps SB201 to SB207, the response generation unit 212 uses the utterance text TXU and the content text TXC, which contains the related content of the utterance text TXU, as input data for the response generation model 222 to generate the response sentence R. Therefore, it is possible to generate a response sentence R that takes into account the content of content C in addition to the content of the utterance of crew member P, enabling a natural response. Furthermore, when there is no content text TXC that contains the related content of the utterance text TXU, the utterance text TXU is used as input data for the response generation model 222 to generate the response text R. This allows for the generation of the response text R without considering the content of content C when the content of crew member P's utterance is not related to the content of content C, resulting in a more natural response.

[0072] [1-5-2. Actions when there are multiple spoken sentences] Next, we will explain the operation in step SB2 when there are multiple utterance sentences TXU received in step SB1. Figure 5 is a flowchart showing the operation of the response generation unit 212, and details the operation of step SB2 when there are multiple received speech sentences TXU.

[0073] At the beginning of step SB2, in step SB211, the input data generation unit 214 attempts to extract content text TXCs containing related content for each received speech text TXU from all content text TXCs that were played back before the speech time and after a predetermined time before the speech time. Furthermore, if multiple content text TXCs are extracted for a single speech text TXU, the input data generation unit 214 extracts the content text TXC whose playback time by the speaker 13 and display 12 is closest to the speech time of the speech text TXU. The predetermined time here is, for example, 30 seconds. The detailed method of extraction is as described in steps SB201 and SB203.

[0074] In step SB212, the input data generation unit 214 determines whether a common content sentence TXC was extracted in step SB211 corresponding to multiple utterance sentences TXU. In step SB212, if the input data generation unit 214 determines that a common content sentence TXC has been extracted corresponding to multiple utterance sentences TXU (step SB212: YES), the response generation unit 212 proceeds to step SB213. An example of such a determination is when multiple crew members P utter response utterance sentences TXU for a single content sentence TXC contained in content C.

[0075] In step SB213, the input data generation unit 214 determines whether the multiple utterances TXU, which have been determined to be related to a common content document TXC as a result of steps SB211 and SB212, are similar to each other. Specifically, in this embodiment, in step SB213, the input data generation unit 214 calculates a similarity score indicating the degree to which the multiple utterances TXU related to the common content document TXC are similar. When the calculated similarity score is greater than or equal to a predetermined value, it is determined that the multiple utterances TXU related to the common content document TXC are similar. When the calculated similarity score is less than the predetermined value, it is determined that the multiple utterances TXU related to the common content document TXC are not similar. In other words, in step SB213, the response generation unit 212 determines whether the similarity of the multiple input voices is equal to or greater than a predetermined value.

[0076] Multiple utterances (TXUs) are considered similar if their semantic content is similar or identical. For example, multiple utterances (TXUs) are considered similar if they both express a positive or negative response to the content of a given text (TXC). Conversely, multiple utterances (TXUs) are not considered similar if, for example, one of two utterances (TXUs) expresses a positive response to the content of a given text (TXC) while the other expresses a negative response.

[0077] In detail, the input data generation unit 214 may use an arbitrary pre-trained model utilizing machine learning to determine whether multiple spoken sentences TXU are similar to each other. For example, the input data generation unit 214 may be configured to input multiple spoken sentences TXU into a pre-trained model and calculate the similarity between the spoken sentences TXU.

[0078] In step SB213, if the input data generation unit 214 determines that the similarity is equal to or greater than a predetermined value (step SB213: YES), the response generation unit 212 proceeds to step SB214.

[0079] In step SB214, the input data generation unit 214 generates input data for the response generation model 222 based on multiple utterance sentences TXU contained in multiple input voices and content sentences TXC contained in related content common to the multiple utterance sentences TXU. The input data is, for example, obtained by converting the multiple utterance sentences TXU and content sentences TXC as text data into a format that can be input to the response generation model 222.

[0080] Next, in step SB215, the response sentence generation unit 213 inputs the input data generated in step SB214 into the response generation model 222 and generates response sentences R for multiple utterance sentences TXU and a common content sentence TXC that is commonly related. After that, the processing of step SB2 shown in Figure 5 is completed, and the process moves to step SB3 shown in Figure 3. In other words, in this embodiment, if the similarity of the multiple input voices is greater than or equal to a predetermined value, the response generation unit 212 generates a common response sentence R for the multiple input voices based on the multiple utterance sentences TXU contained in the multiple input voices and the content sentence TXC contained in the content C.

[0081] Furthermore, if the input data generation unit 214 determines in step SB212 that no common content sentence TXC was extracted corresponding to multiple utterance sentences TXU (step SB212: NO), the response generation unit 212 proceeds to step SB216. Examples of cases in which such a determination is made include when multiple crew members P each utter an utterance sentence TXU that is a response to a different content sentence TXC.

[0082] Similarly, if the input data generation unit 214 determines in step SB213 that the similarity is less than a predetermined value (step SB213: NO), the response generation unit 212 proceeds to step SB216. Examples of situations in which such a determination is made include cases where multiple crew members P utter different speech sentences TXU in response to the same content sentence TXC.

[0083] In step SB216, the response generation unit 212 applies the processing from steps SB201 to SB205 in Figure 4 to each of the multiple utterance sentences TXU. Then, for each utterance sentence TXU, the response sentence R is generated using either the utterance sentence TXU alone, or the utterance sentence TXU and the content sentence TXC associated with that utterance sentence TXU, as input data generation unit 214. After that, the processing in step SB2 shown in Figure 5 is completed, and the process moves to step SB3 in Figure 3. In other words, in this embodiment, if the similarity of the multiple input voices is less than a predetermined value, the response generation unit 212 individually generates a response sentence R for each utterance sentence TXU contained in the multiple input voices, based on the utterance sentence TXU contained in each input voice and the content sentence TXC contained in content C.

[0084] As shown in steps SB211 to SB216, when the response generation unit 212 receives multiple utterances TXU that are related to and similar to a common content sentence TXC, it generates a common response sentence R for the multiple utterances TXU. Therefore, when multiple crew members P show similar reactions to content C, outputting a common response sentence R makes it easier to produce a natural response. Furthermore, when multiple utterances TXU are not related to a common content text TXC, or are not similar to each other, the response generation unit 212 generates a response text R individually for each utterance TXU. Therefore, when multiple crew members P show different reactions to content C, a different response text R can be output for each crew member P, making it easier to produce natural responses.

[0085] [2. Other Embodiments] The embodiments described above are merely examples and can be modified and applied as needed.

[0086] In this embodiment, the response system 1000 was configured to output a response sentence R in response to the utterance of an occupant P inside the vehicle 1, but this is just one example. For example, the response device 100 may be placed inside a building or the like, and may be configured to output a response sentence R in response to the utterance of a person inside the room, taking into account the content C played back inside the room. In addition, the response system 1000 may be configured to output a response sentence R in response to the utterance of a person in any space.

[0087] The first processor 110 and the second processor 210 may each be composed of multiple processors or a single processor. Each processor 110, 210 may also be hardware programmed to implement the functional units described above. In this case, these processors may be composed of, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0088] Furthermore, the configuration of each part of the response system 1000 shown in Figure 2 is merely an example, and the specific implementation form is not particularly limited. In other words, it is not necessarily required that hardware corresponding to each part be implemented individually, and it is certainly possible to configure the system so that a single processor executes a program to realize the functions of each part. Also, some of the functions realized by software in the above embodiment may be implemented by hardware, or some of the functions realized by hardware may be implemented by software.

[0089] Furthermore, the operation steps shown in Figures 3 to 5 are divided according to the main processing content, and the present invention is not limited by the way the processing units are divided or their names. Depending on the processing content, it may be further divided into more steps. Alternatively, it may be divided so that one step unit includes even more processing. Also, the order of the steps may be changed as appropriate, as long as it does not impede the spirit of the present invention.

[0090] [3. Configurations supported by the above embodiments] The above embodiment supports the following configuration:

[0091] (Composition 1) A response system comprising: a content playback unit for playing content; a microphone; an input voice recognition unit for recognizing input voice to the microphone; and a response generation unit for generating a response statement to the input voice based on the input voice and the content played by the content playback unit up to the time the input voice was input to the microphone. According to the response system of Configuration 1, the content playback unit can take into account the content played back and generate a response sentence to the input speech. Therefore, it can provide a natural response to the input speech.

[0092] (Configuration 2) The response system according to Configuration 1, wherein the response generation unit determines whether the content played by the content playback unit up to the time of input of the input sound to the microphone is related content associated with the input sound, and if the content played by the content playback unit is related content, generates the response statement based on the input sound and the related content. According to the response system of configuration 2, when the input audio relates to content played back by the content playback unit, a response sentence can be generated for the input audio while taking the content into account. Therefore, a natural response to the input audio can be provided.

[0093] (Composition 3) The response system according to configuration 1 or 2, wherein the response generation unit determines whether the content is related content by targeting content played by the content playback unit, or, in the case of content containing multiple sentences played by the content playback unit, the sentences from the last predetermined number onward, at a predetermined time prior to the time when the input sound is input to the microphone. According to the response system of Configuration 3, it is possible to determine whether the content played at a time close to the time the input audio was received is related to the input audio, and to generate a response sentence to the input audio. Therefore, it is possible to provide a natural response to the input audio.

[0094] (Composition 4) The response system according to any one of configurations 1 to 3, wherein the response generation unit determines whether the content played by the content playback unit up to the time of input of the input sound to the microphone is related content associated with the input sound, and if the content played by the content playback unit is not related content, the response system generates the response statement based on the input sound. According to the response system of configuration 4, when the input audio is not related to the content played back by the content playback unit, a response sentence can be generated to the input audio without taking the content into consideration. Therefore, a natural response to the input audio can be provided.

[0095] (Composition 5) The response system according to any one of configurations 1 to 4, wherein the response generation unit generates the response statement based on the input audio and the sentence played by the content playback unit at the time closest to the time of input audio input, if the content played by the content playback unit up to the time of input audio input to the microphone includes a plurality of sentences related to the input audio. According to the response system of configuration 5, it is possible to determine whether the content played at a time close to the time the input voice was input is related to the input voice, and to generate a response sentence to the input voice. Therefore, it is possible to provide a natural response to the input voice.

[0096] (Composition 6) A response system according to any one of configurations 1 to 5, wherein, when the input voice recognition unit recognizes multiple input voices to the microphone, the response generation unit generates a common response statement for the multiple input voices based on the multiple input voices and the content if the similarity of the multiple input voices is greater than or equal to a predetermined value, and generates a response statement individually for each of the multiple input voices based on the input voice and the content if the similarity of the multiple input voices is less than the predetermined value. According to the response system of configuration 6, when multiple sentences in the input speech are similar, a common response sentence can be generated for all of them, and when multiple sentences in the input speech are not similar, a response sentence can be generated for each of the sentences. Therefore, it is possible to provide a natural response to the input speech.

[0097] (Composition 7) A response method comprising: a content playback unit plays content; an input voice recognition unit recognizes the voice input to the microphone; and a response generation unit generates a response statement to the input voice based on the input voice and the content played by the content playback unit up to the time the input voice was input to the microphone. According to the response method of configuration 7, the content playback unit can take into account the content played back and generate a response sentence to the input audio. Therefore, it is possible to provide a natural response to the input audio. [Explanation of Symbols]

[0098] 1...Vehicle, 10...Seat, 10A...Driver's seat, 10B...Passenger seat, 10C...Rear right seat, 10D...Rear left seat, 11...Microphone, 11A...Driver's seat microphone (microphone), 11B...Passenger seat microphone (microphone), 11C...Rear right seat microphone (microphone), 11D...Rear left seat microphone (microphone), 12...Display (content playback unit), 12A...Center display (content playback unit), 12B...Passenger seat display (content playback unit), 12C...Rear right seat display (content playback unit), 12D...Rear left seat display (content playback unit), 13...Speaker (content playback unit), 13A...Center speaker (content playback unit), 13B...Passenger seat speaker (content playback unit), 13C...Rear right seat speaker (content playback unit), 13D...Rear left seat speaker (content playback unit), 13D...Speaker (content playback unit) ), 100...Response device, 110...First processor, 111...First communication control unit, 112...Input / output control unit, 113...Input speech recognition unit, 114...Content recognition unit, 120...First memory, 121...First control program, 122...Content data, 130...First communication unit, 200...Response generation server, 210...Second processor, 211...Second communication control unit, 212...Response generation unit, 213...Response sentence generation unit, 214...Input Data generation unit, 220... Second memory, 221... Second control program, 222... Response generation model, 230... Second communication unit, 300... Content distribution server, 1000... Response system, C... Content, DTC... Playback time information, DTU... Utterance time information, DUC... Playback time information, IFP... Speaker information, NW... Communication network, P, P1~P4... Crew, R... Response sentence, TXC... Content text, TXU... Utterance text.

Claims

1. A content playback unit that plays content, Mike and, The input voice recognition unit recognizes the audio input to the microphone, A response generation unit generates a response statement to the input audio based on the input audio and the content played by the content playback unit up to the time the input audio is input to the microphone, A response system equipped with the following features.

2. The response generation unit, The system determines whether the content played by the content playback unit up to the time of input audio to the microphone is related content associated with the input audio, and if the content played by the content playback unit is related content, it generates the response statement based on the input audio and the related content. The response system according to claim 1.

3. The response generation unit, The system determines whether the content is related content by considering the content played by the content playback unit after a predetermined time prior to the time when the input audio was input to the microphone, or, in the case where the content playback unit plays the content containing multiple sentences, the sentences from the last predetermined number onward. The response system according to claim 2.

4. The response generation unit, The system determines whether the content played by the content playback unit up to the time of input audio input to the microphone is related content associated with the input audio, and if the content played by the content playback unit is not related content, it generates the response statement based on the input audio. The response system according to claim 1.

5. The response generation unit, If the content played by the content playback unit up to the time of input audio input to the microphone includes multiple sentences related to the input audio, the response sentence is generated based on the input audio and the sentence played by the content playback unit at the time closest to the time of input audio input. The response system according to claim 2.

6. When the input voice recognition unit recognizes multiple input voices to the microphone, The response generation unit, If the similarity of the multiple input voices is greater than or equal to a predetermined value, a common response statement for the multiple input voices is generated based on the multiple input voices and the content. If the similarity of the multiple input voices is less than the predetermined value, the response sentence is generated individually for each of the multiple input voices based on the input voice and the content. The response system according to claim 2.

7. The content playback unit plays the content, The input voice recognition unit recognizes the voice input to the microphone. The response generation unit generates a response statement for the input audio based on the input audio and the content that the content playback unit has played up to the time the input audio was input to the microphone. How to respond.

Citation Information

Patent Citations

  • Information processing device and information processing method

    WO2022050060A1