Information processing device

The information processing apparatus enhances voice assistant activation by using preset speech conditions and user attributes to ensure timely activation without relying on detected keywords, addressing the urgency in user speech content.

JP2025094674AInactive Publication Date: 2025-06-25TOYOTA JIDOSHA KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023210378
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing voice assistant activation techniques fail to promptly activate the function when there is a high urgency in the content of the user's speech, even if the wake-up word is not included.

Method used

An information processing apparatus that activates the voice assistant function based on preset speech conditions, including user attributes and speech content urgency, without requiring a detected keyword.

Benefits of technology

Enables prompt activation of the voice assistant function in response to high-urgency speech content, improving the responsiveness and effectiveness of voice assistant activation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094674000001_ABST
    Figure 2025094674000001_ABST
Patent Text Reader

Abstract

To provide an information processing device that improves a technology for activating a voice assistant function.SOLUTION: An information processing device 10 includes a control unit 14. The control unit 14 activates, when detecting an utterance that satisfies a preset utterance condition, a voice assistant function and outputs a response to contents of the utterance, even if a keyword for activating the voice assistant function is not detected.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus.

Background Art

[0002] Conventionally, there has been known a technique of activating a voice assistant function when a wake-up word included in a user's spoken voice is detected (Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is room for improvement in the technique of activating the voice assistant function. For example, even if the wake-up word is not included in the user's spoken voice, depending on the content of the speech, etc., there may be a high urgency to activate the voice assistant function.

[0005] In view of this point, an object of the present disclosure is to improve the technique of activating the voice assistant function.

Means for Solving the Problems

[0006] An information processing apparatus according to an embodiment of the present disclosure includes a control unit that, when detecting a speech that satisfies a preset speech condition, activates the voice assistant function and outputs a response to the content of the speech even if a keyword for activating the voice assistant function is not detected.

Effects of the Invention

[0007] According to an embodiment of the present disclosure, the technique of activating the voice assistant function can be improved.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings.

[0010] (Configuration of the Vehicle) As shown in FIG. 1, the vehicle 1 according to the present embodiment includes an information processing apparatus 10, microphone units 21, 22, 23, 24, a speaker unit 30, a display device 40, and a plurality of vehicle devices 50. Hereinafter, when the microphone units 21, 22, 23, 24 are not particularly distinguished, they are described as "microphone unit 20". The vehicle 1 may further include an in-vehicle camera 60.

[0011] As shown in FIG. 2, the vehicle 1 has seats 2, 3, 4, 5. The seat 2 is the driver's seat. The seat 3 is the passenger seat. The seat 4 is the seat on the right side in the traveling direction among the two rear seats. The seat 5 is the seat on the left side in the traveling direction among the two rear seats. Adults are sitting on the seats 2 and 3. Children are sitting on the seats 4 and 5. In the present embodiment, the vehicle 1 has four seats. However, the vehicle 1 may have any number of seats.

[0012] The information processing apparatus 10 as shown in FIG. 1 has a voice assistant function. The information processing apparatus 10 may be any device as long as it has a voice assistant function.

[0013] The microphone unit 20 collects sound based on the control of the information processing apparatus 10. The microphone unit 20 transmits the data of the collected sound to the information processing apparatus 10.

[0014] Each microphone unit 20 is located at a position where it can collect the sound of each user riding in the vehicle 1. For example, as shown in FIG. 2, the microphone unit 21 is located at the seat 2. By being located at the seat 2, the microphone unit 21 can collect the sound of the user sitting on the seat 2. The microphone unit 22 is located at the seat 3. By being located at the seat 3, the microphone unit 22 can collect the sound of the user sitting on the seat 3. The microphone unit 23 is located at the seat 4. By being located at the seat 4, the microphone unit 23 can collect the sound of the user sitting on the seat 4. The microphone unit 24 is located at the seat 5. By being located at the seat 5, the microphone unit 24 can collect the sound of the user sitting on the seat 5.

[0015] The speaker unit 30 as shown in FIG. 1 outputs sound based on the control of the information processing apparatus. The speaker unit 30 may be located at any location in the passenger compartment of the vehicle 1.

[0016] The display device 40 is, for example, a car navigation device. The display device 40 is configured to include a display. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display or the like.

[0017] The vehicle device 50 may be any device mounted on the vehicle 1. The vehicle device 50 is, for example, an air conditioner, an actuator for opening and closing a door, an actuator for opening and closing a window, or an audio device or the like. The vehicle device 50 executes processing based on the control of the information processing apparatus 10.

[0018] The in-vehicle camera 60 generates video data by performing imaging under the control of the information processing device 10. The in-vehicle camera 60 transmits the video data to the information processing device 10. The in-vehicle camera 60 is located at a position in the vehicle interior of the vehicle 1 where the face of each user can be imaged.

[0019] (Configuration of Information Processing Device) As shown in FIG. 1, the information processing device 10 includes a communication unit 11, a positioning unit 12, a storage unit 13, and a control unit 14.

[0020] The communication unit 11 is configured to include at least one communication module capable of communicating with various components of the vehicle 1. The communication module is a communication module compatible with the in-vehicle network standard such as CAN (Controller Area Network).

[0021] The communication unit 11 may be configured to include at least one communication module connectable to a network. The network may be any network including a mobile communication network and the Internet, etc. The communication module is, for example, a communication module compatible with a mobile communication standard such as LTE (Long Term Evolution), 4G (4th Generation) or 5G (5th Generation).

[0022] The positioning unit 12 can acquire the position information of the vehicle 1. The positioning unit 12 is configured to include at least one receiving module compatible with a satellite positioning system. The receiving module is, for example, a receiving module compatible with GPS (Global Positioning System).

[0023] The storage unit 13 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a RAM (Random Access Memory) or a ROM (Read Only Memory), etc. The RAM is, for example, an SRAM (Static Random Access Memory) or a DRAM (Dynamic Random Access Memory), etc. The ROM is, for example, an EEPROM (Electrically Erasable Programmable Read Only Memory), etc. The storage unit 13 may function as a main memory device, an auxiliary memory device, or a cache memory. The storage unit 13 may store a system program, an application program, embedded software, etc. The storage unit 13 stores data used for the operation of the information processing apparatus 10 and data obtained by the operation of the information processing apparatus 10. For example, the storage unit 13 stores information associating the microphone units 21, 22, 23, 24 as shown in FIG. 2 with the seats 2, 3, 4, 5 where the microphone units 21, 22, 23, 24 are respectively located. Also, map information may be stored in the storage unit 13.

[0024] The control unit 14 is configured to include at least one processor, at least one dedicated circuit, or a combination of these. The processor is, for example, a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), etc. The control unit 14 executes processing related to the operation of the information processing apparatus 10 while controlling each part of the information processing apparatus 10. For example, the control unit 14 controls the display device 40 or the vehicle equipment 50 by transmitting a control signal to the display device 40 or the vehicle equipment 50 through the communication unit 11.

[0025] The control unit 14 receives, via the communication unit 11, the voice data collected by the microphone unit 20 from the microphone unit 20. The control unit 14 executes voice recognition processing on the received voice data. The voice recognition processing is, for example, processing for converting voice data into character data. When the control unit 14 executes the voice recognition processing and detects a first keyword (keyword) for activating the voice assistant from the received voice data, the control unit 14 activates the voice assistant function. The voice assistant function is a function for controlling the display device 40 or the vehicle device 50 based on the content of the user's speech. When the control unit 14 activates the voice assistant function, the control unit 14 controls the display device 40 or the vehicle device 50 based on the voice data received from the microphone unit 20. That is, the control unit 14 executes the voice assistant function. After executing the voice assistant function, the control unit 14 sets the voice assistant function to the standby state.

[0026] Here, even if the first keyword is not detected from the received voice data, the control unit 14 activates the voice assistant function when detecting a speech that satisfies a preset speech condition. The speech condition may be arbitrarily set based on the content of a speech with a high urgency for activating the voice assistant function, the state of the user, or the like. In the present embodiment, the speech condition is a condition that the user who made the speech has a preset predetermined attribute and the content of the user's speech is preset predetermined content.

[0027] The preset predetermined attribute in the speech condition may be set in consideration of the attributes of users who may become vulnerable in the vehicle interior of the vehicle 1. The user attribute is the nature or characteristic of the user. Here, a child or an elderly person may become vulnerable in the vehicle interior of the vehicle 1. Therefore, the preset predetermined attribute in the speech condition may be a child or an elderly person. In the present embodiment, the predetermined attribute is assumed to be a child below a predetermined age. The predetermined age is, for example, 6 years old.

[0028] The preset predetermined content under the speaking conditions may be content with a high urgency to activate the voice assistant function in the passenger compartment of the vehicle 1. When the control unit 14 detects a preset second keyword from the data of the user's voice, it may determine that the content of the user's speech is the preset predetermined content. The second keyword may be set in consideration of the content of speech with a high urgency to activate the voice assistant function. For example, when a child says "I want to go to the toilet", "I feel sick", or "Help me", the urgency to activate the voice assistant function is higher than when an adult says "I want to go to the toilet", etc. In this case, the second keyword may be "I want to go to the toilet", "I feel sick", or "Help me".

[0029] Figure 3 is a flowchart showing an example of the processing procedure executed by the information processing apparatus 10 shown in FIG. 1. Before the processing of S1, the voice assistant function is in a standby state. When the ignition of the vehicle 1 is turned on, for example, the control unit 14 starts the processing of S1.

[0030] In the processing of S1, the control unit 14 receives, by the communication unit 11, the voice data of the users of the respective seats 2, 3, 4, and 5 from the respective microphone units 21, 22, 23, and 24. By analyzing the voice data of the users of the respective seats 2, 3, 4, and 5, the control unit 14 identifies the attributes of the users of the respective seats 2, 3, 4, and 5. In the present embodiment, the control unit 14 identifies whether the user is a child under a predetermined age or the user is a child or an adult over the predetermined age as the user's attribute. For example, the control unit 14 identifies that the attributes of the users of the respective seats 2 and 3 are adults by analyzing the voice data respectively acquired from the microphone units 21 and 22. The control unit 14 identifies that the attributes of the users of the respective seats 4 and 5 are children under a predetermined age by analyzing the voice data respectively acquired from the microphone units 23 and 24.

[0031] In the processing of S2, the control unit 14 receives voice data by the communication unit 11 from any one of the microphone units 21, 22, 23, and 24.

[0032] In the process of S3, the control unit 14 determines whether the user's speech has been detected by performing speech recognition processing on the voice data received in the process of S2. If the control unit 14 determines that the user's speech has been detected (S3: YES), it proceeds to the process of S4. On the other hand, if the control unit 14 determines that the user's speech has not been detected (S3: NO), it returns to the process of S2.

[0033] In the process of S4, the control unit 14 determines whether the user who made the speech detected in the process of S3 is a child under a predetermined age. As an example of this process, first, the control unit 14 identifies the seat where the microphone unit 20 that received the voice data in the process of S2 is located according to the above-mentioned association information stored in the storage unit 13. Next, the control unit 14 identifies the attribute of the user in the identified seat based on the result of the process of S1. Further, the control unit 14 determines whether the attribute of the identified user is a child under a predetermined age. For example, assume that the microphone unit 20 that received the voice data in the process of S2 is the microphone unit 23. In this case, first, the control unit 14 identifies the seat 4 where the microphone unit 23 is located according to the above-mentioned association information stored in the storage unit 13. Next, based on the result of the process of S1, the control unit 14 identifies that the attribute of the user in the identified seat 4 is a child under a predetermined age. Further, the control unit 14 determines that the user who made the speech detected in the process of S3 is a child under a predetermined age.

[0034] If the control unit 14 determines that the user who made the speech detected in the process of S3 is not a child under a predetermined age (S4: NO), it proceeds to the process of S5. On the other hand, if the control unit 14 determines that the user who made the speech detected in the process of S3 is a child under a predetermined age (S4: YES), it proceeds to the process of S6.

[0035] In the process of S5, the control unit 14 maintains the voice assistant function in a standby state. After the process of S5, the control unit 14 returns to the process of S2.

[0036] In the process of S6, the control unit 14 determines whether the content of the utterance detected in the process of S3 is a content with a high urgency to activate the voice assistant function. When the control unit 14 detects a second keyword from the user's voice data received in the process of S2, it determines that the content of the utterance detected in the process of S3 is a content with a high urgency to activate the voice assistant function.

[0037] When the control unit 14 determines that the content of the utterance detected in the process of S3 is not a content with a high urgency (step S6: NO), it proceeds to the process of step S5. On the other hand, when the control unit 14 determines that the content of the utterance detected in the process of S3 is a content with a high urgency (step S6: YES), it proceeds to the process of S7.

[0038] In the process of S7, the control unit 14 activates the voice assistant function.

[0039] In the process of S8, the control unit 14 estimates the intention of the utterance detected in the process of S3. For example, the control unit 14 estimates the intention of the utterance detected in the process of S3 by performing natural language processing on the voice data received in the process of S2.

[0040] In the process of S9, the control unit 14 outputs a response to the content of the utterance based on the intention of the utterance estimated in the process of S8.

[0041] As an example of the process of S9, when the control unit 14 estimates that the intention of the speech in the process of S8 is the intention of "wanting to go to the toilet", the control unit 14 acquires the position information of the vehicle 1 by the positioning unit 12. The control unit 14 refers to the map information stored in the storage unit 13 and acquires the information of the toilet within a predetermined range from the position of the vehicle 1. The predetermined range is, for example, a range that the vehicle 1 can reach in a relatively short time such as 10 minutes. When the control unit 14 acquires the information of the toilet, the control unit 14 transmits a control signal to the display device 40 through the communication unit 11, and causes the display device 40 to display a display asking whether to guide to the toilet as a response to the intention of the speech estimated in the process of S8. For example, when there is a convenience store with a toilet within a predetermined range from the position of the vehicle 1, the control unit 14 causes the display device 40 to display a response such as "There is a toilet in the nearby convenience store. Do you want to be guided?" The control unit 14 may output the voice data of the response "There is a toilet in the nearby convenience store. Do you want to be guided?" to the speaker unit 30 by transmitting a control signal to the speaker unit 30 through the communication unit 11.

[0042] As another example of the process of S9, when the control unit 14 estimates that the intention of the speech in the process of S8 is the intention of "feeling sick", the control unit 14 outputs an inquiry of whether to open the window to the speaker unit 30 by transmitting a control signal to the speaker unit 30 through the communication unit 11. For example, the control unit 14 outputs the voice data of the response "Do you want to open the window?" to the speaker unit 30.

[0043] In the process of S10, the control unit 14 outputs an execution instruction of a function according to the user's answer to the response output in the process of S9 to the display device 40 or the vehicle device 50. The control unit 14 may receive the user's answer to the response as voice data from the microphone unit 20.

[0044] As an example of the process of S10, it is assumed that the control unit 14 has obtained an affirmative answer of "Yes" from the user in response to the question "There is a toilet in a nearby convenience store. Do you want to be guided there?" in the process of S9. In this case, as an instruction to execute a function according to the user's answer, the control unit 14 outputs an instruction to the display device 40, which is a car navigation device, to execute guidance to the toilet. The control unit 14 outputs an instruction to execute guidance to the toilet to the display device 40 by transmitting a control signal to the display device 40 through the communication unit 11. When receiving this instruction, the display device 40 superimposes and displays a list of toilets near the vehicle 1 on the map data. The display device 40 executes guidance to the toilet selected by the user from the displayed list of toilets.

[0045] As another example of the process of S10, it is assumed that the control unit 14 has obtained an affirmative answer of "Yes" from the user in response to the question "Do you want to open the window?" in the process of S9. In this case, as an instruction to execute a function according to the user's answer, the control unit 14 outputs an instruction to open the window to the vehicle device 50, which is an actuator for opening and closing the window. The control unit 14 outputs an instruction to open the window to the vehicle device 50 by transmitting a control signal to the vehicle device 50 through the communication unit 11.

[0046] In the process of S11, after executing the voice assistant function, the control unit 14 sets the voice assistant function to the standby state. After the process of S11, the control unit 14 returns to the process of S2.

[0047] Here, in the processes from S1 to S11, when the ignition of the vehicle 1 is turned off, the control unit 14 may end the processing procedure as shown in FIG. 3.

[0048] As described above, in the information processing apparatus 10 according to the present embodiment, when the control unit 14 detects a speech that satisfies a preset speech condition, the control unit 14 activates the voice assistant function even if a first keyword (keyword) for activating the voice assistant function is not detected. Further, the control unit 14 outputs a response to the content of the speech. With such a configuration, even if the speech uttered by the user does not include the first keyword, when the content of the speech has a high urgency to activate the voice assistant function, the voice assistant function can be activated promptly. Therefore, according to the present embodiment, the technology for activating the voice assistant function can be improved.

[0049] Furthermore, in the present embodiment, the speech condition may be a condition that the user who uttered the speech has a preset predetermined attribute and the content of the user's speech is a preset predetermined content. Further, the predetermined attribute may be a child. The predetermined content may be content that has a high urgency to activate the voice assistant function in the vehicle interior. With such a configuration, for example, even if a child does not utter the first keyword, if the child utters a speech with high urgency, the voice assistant function can be activated promptly. Further, a response to the content of the speech can be output.

[0050] Although the present disclosure has been described based on the drawings and examples, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions included in each component or each step, etc. can be rearranged so as not to be logically contradictory, and it is possible to combine or divide a plurality of components or steps, etc. into one.

[0051] For example, in the process of S1, it was described that the control unit 14 identifies the attributes of the users in seats 2, 3, 4, and 5 based on the voice data of each of the users in seats 2, 3, 4, and 5 received from the microphone units 21, 22, 23, and 24. However, the control unit 14 may identify the attributes of the users in seats 2, 3, 4, and 5 by analyzing the video data received from the in-vehicle camera 60 by the communication unit 11.

[0052] For example, the control unit 14 may not execute the process of S1. In this case, in the process of S4, the control unit 14 may identify the attributes of the user by analyzing the voice data of the user received in the process of S2. The control unit 14 may determine whether the user who made the speech detected in the process of S3 is a child under a predetermined age based on the identified attributes of the user (S4).

[0053] For example, in the process of S6, it was described that when the control unit 14 detects a second keyword from the voice data of the user received in the process of S2, the control unit 14 determines that the content of the speech detected in the process of S3 is content with a high urgency to activate the voice assistant function. However, the control unit 14 may determine whether the content of the speech detected in the process of S3 is content with a high urgency to activate the voice assistant function by any method other than the second keyword. As another example, the control unit 14 may estimate the emotion of the user by analyzing the voice data of the user received in the process of S2. The control unit 14 may determine whether the content of the speech detected in the process of S3 is content with a high urgency to activate the voice assistant function based on the estimated emotion of the user.

[0054] For example, in the above-described embodiment, the predetermined attribute preset in the speech condition was described as being a child. However, the predetermined attribute is not limited to a child. As another example, the predetermined attribute may be an elderly person. In this case, the elderly person may be a person who is equal to or older than a preset age. The preset age may be set in consideration of the age of a person who may become a vulnerable person due to being elderly in the passenger compartment of the vehicle 1.

[0055] For example, in the above-described embodiment, the information processing apparatus 10 has been described as being mounted on the vehicle 1. However, the information processing apparatus 10 does not necessarily have to be mounted on the vehicle 1. As another example, the information processing apparatus 10 may be a cloud server. In this case, the vehicle 1 may include a communication device capable of communicating with the information processing apparatus 10 that is a cloud server.

Explanation of Reference Numerals

[0056] 1: Vehicle, 2, 3, 4, 5: Seats, 10: Information processing apparatus, 11: Communication unit, 12: Positioning unit, 13: Storage unit, 14: Control unit, 20, 21, 22, 23, 24: Microphone units, 30: Speaker unit, 40: Display device, 50: Vehicle equipment, 60: In-vehicle camera

Claims

1. When a speech that satisfies a preset speech condition is detected, even if a keyword for activating the voice assistant function is not detected, the information processing apparatus includes a control unit that activates the voice assistant function and outputs a response to the content of the speech.

2. The information processing apparatus according to claim 1, wherein the speech condition is a condition that the user who uttered the speech has a preset predetermined attribute and the content of the user's speech is a preset predetermined content.

3. The information processing apparatus according to claim 2, wherein the predetermined attribute is being a child.

4. The information processing apparatus according to claim 2 or 3, wherein the predetermined content is content with a high urgency to activate the voice assistant function in the vehicle interior.

5. The information processing apparatus according to claim 4, wherein the predetermined content is content such as wanting to go to the toilet.

Citation Information

Patent Citations

  • Apparatus controller and apparatus control method

    JP2018207169A

  • Electronic apparatus and control method thereof

    JP2020122819A

  • Agent device, control method of agent device, and program

    JP2020152298A

  • Automated assistants that address multiple age groups and / or vocabulary levels

    JP2021513119A

  • Adapting automated assistants without hotwords

    JP2021520590A