Speech recognition method and speech recognition device

The speech recognition method enhances the accuracy of identifying vehicle components by using a voice recognition device to locate and reference component positions, addressing the challenge of inaccurate component identification in voice input systems.

JP7720420B2Active Publication Date: 2025-08-07NISSAN MOTOR CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023576248
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-08-07
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

Existing voice input systems struggle to accurately identify vehicle components mentioned in user utterances, especially when the user is engaged in other tasks like driving, leading to challenges in speech recognition accuracy.

Method used

A speech recognition method that involves a controller to identify a reference position in the utterance and reference a storage device or learning model to estimate the target component based on its position, using a voice recognition device with components like lamps, displays, and sensors to enhance accuracy.

Benefits of technology

Improves the accuracy of estimating vehicle components mentioned in user speech, enabling precise identification of visual and auditory devices within the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007720420000001
    Figure 0007720420000001
  • Figure 0007720420000002
    Figure 0007720420000002
  • Figure 0007720420000003
    Figure 0007720420000003
Patent Text Reader

Abstract

In a voice recognition method for acquiring an utterance content of a user of a vehicle, and estimating an object constituent that is a constituent mentioned in the utterance content among a plurality of constituents constituting the vehicle, a mentioned position that is a position mentioned in the utterance content is specified on the basis of the utterance content (S4), and with reference to a storage device storing constituent positions that are positions at which the plurality of constituents are respectively provided or a learning model that has learned the constituent positions, a constituent provided at a constituent position that matches the specified mentioned position is estimated as the object constituent (S5).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a speech recognition method and a speech recognition device. [Background technology]

[0002] In recent years, voice input systems have been proposed that use voice recognition to respond to questions from users and operate devices. For example, Patent Document 1 listed below describes a vehicle lighting device that, when it detects that a user has asked a question about how to operate an air conditioner, illuminates the air conditioner switch and moves a pointer displayed in the illuminated area along the operation direction of the switch. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6668296 specification Summary of the Invention [Problem to be solved by the invention]

[0004] The voice input system can respond to a vehicle user's voice question by providing the meaning of components that make up the vehicle (for example, the meaning of a lamp) and how to operate the vehicle (how to operate a switch). In such a voice input system, it is necessary to accurately identify the component mentioned by the user. However, it may be difficult to produce sufficient speech to accurately identify the component (e.g., a lamp or a switch). For example, if the user is engaged in other tasks, such as driving a vehicle, it may be difficult to produce appropriate speech. The present invention aims to improve the accuracy of estimating components mentioned in a user's utterance in speech recognition that estimates components mentioned in a user's utterance from among multiple components that make up a vehicle. [Means for solving the problem]

[0005] According to one aspect of the present invention, there is provided a speech recognition method for acquiring an utterance content of a vehicle user and estimating a target component that is a component mentioned in the utterance content among multiple components that make up the vehicle. The speech recognition method causes a controller to execute a process of identifying a reference position that is a position mentioned in the utterance content based on the utterance content, and a process of referencing a storage device that stores component positions that are positions where multiple components are located or a learning model that has learned the component positions, and estimating, as the target component, a component that is located at a component position that matches the identified reference position. [Effects of the Invention]

[0006] According to the present invention, in speech recognition that estimates a component mentioned in a user's speech content out of multiple components that make up a vehicle, the accuracy of estimating the component mentioned in the speech content can be improved. The objects and advantages of the invention will be realized and attained by means of the elements and combinations set forth in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention as claimed. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a schematic diagram illustrating an example of a vehicle equipped with a voice recognition device according to an embodiment. [Figure 2] 1 is a block diagram illustrating an example of a functional configuration of a voice recognition device according to an embodiment. [Figure 3] FIG. 1 is a schematic diagram of an example of components that configure a vehicle. [Figure 4] FIG. 2 is a schematic diagram of another example of components that configure a vehicle. [Figure 5] FIG. 10 is a schematic diagram of an example of component location information. [Figure 6] 1 is a flowchart illustrating an example of a speech recognition method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] (composition) 1 is a schematic diagram of an example of a vehicle equipped with a voice recognition device according to an embodiment. The vehicle 1 includes components 2 constituting the vehicle 1, an in-vehicle device controller 3, in-vehicle sensors 4, a human-machine interface (hereinafter referred to as "HMI") 6, and a voice recognition device 7. Components 2 are various in-vehicle devices that are installed in the vehicle 1 and make up the vehicle 1.

[0009] For example, the component 2 may be a lamp such as a warning lamp or an indicator lamp arranged in the meter cluster of the instrument panel at the driver's seat or near the A-pillar of the vehicle 1. For example, the component 2 may be a display device provided in the center cluster or center console. The lamp or display device is an example of a device provided inside the vehicle 1 to present visual information to the user. Furthermore, for example, the component 2 may be an alarm device that outputs an alarm sound to a user of the vehicle 1. The alarm device is an example of a device that is provided inside the vehicle and presents auditory information to the user. For example, component 2 may be a navigation system that sets a driving route based on the current position of vehicle 1 measured by a positioning device (such as a Global Navigation System (GNSS) receiver) and map information, and provides route guidance to the occupants according to this driving route. Furthermore, for example, the component 2 may be a window provided in a door of the vehicle 1.

[0010] The in-vehicle device controller 3 is an electronic control unit (ECU) that controls the operation of the component 2, which is an in-vehicle device, and generates control signals for controlling the component 2. The in-vehicle device controller 3 includes, for example, a processor and peripheral components such as a storage device. The processor may be, for example, a CPU (Central Processing Unit) or an MPU (Micro-Processing Unit). The storage device may include a semiconductor storage device, a magnetic storage device, an optical storage device, etc. The storage device may include a register, a cache memory, a memory such as a ROM (Read Only Memory) and a RAM (Random Access Memory) used as a main memory device.

[0011] The in-vehicle device controller 3 may be formed of dedicated hardware for executing each of the information processes described below. For example, the in-vehicle device controller 3 may include a functional logic circuit configured in a general-purpose semiconductor integrated circuit. For example, the in-vehicle device controller 3 may include a programmable logic device (PLD) such as a field-programmable gate array (FPGA).

[0012] The interior sensor 4 is a sensor that detects the state inside the vehicle 1. For example, the interior sensor 4 may be an interior camera that takes pictures of the interior of the vehicle, a pressure sensor or a seat belt sensor that is provided in a seat to determine whether an occupant is seated, a biosensor that detects biometric information of an occupant, or a microphone that detects sounds generated from the vehicle 1. The HMI 6 is an interface device that exchanges information between the voice recognition device 7 and the user. The HMI 6 includes a display device that can be seen by the user of the vehicle 1 (for example, a display screen of a navigation system), as well as a speaker and a buzzer for outputting warning sounds, notification sounds, and voice information. The HMI 6 also includes a voice input device (for example, a microphone) that acquires voice input from the user.

[0013] The voice recognition device 7 is an electronic control unit that operates as a controller that performs voice recognition to recognize the content of an utterance made by a user of the vehicle 1. The voice recognition device 7 estimates a component 2 mentioned in the content of the user's utterance and outputs information related to the mentioned component 2 from the HMI 6 to provide the information to the user. Alternatively, the voice recognition device 7 operates the component 2 mentioned in the content of the user's utterance.

[0014] The speech recognition device 7 includes a processor 8 and peripheral components such as a storage device 9. The processor 8 may be, for example, a CPU or an MPU. The storage device 9 may include a semiconductor storage device, a magnetic storage device, an optical storage device, etc. The storage device 9 may include memories such as a register, a cache memory, and a ROM and a RAM used as a main storage device. The functions of the speech recognition device 7 described below are realized by, for example, the processor 8 executing a computer program stored in the storage device 9. The speech recognition device 7 may be formed of dedicated hardware for executing each of the information processes described below. For example, the speech recognition device 7 may include a functional logic circuit configured in a general-purpose semiconductor integrated circuit. For example, the speech recognition device 7 may include a programmable logic device such as a field programmable gate array.

[0015] 2 is a block diagram showing an example of the functional configuration of the speech recognition device 7. The speech recognition device 7 operates as a speech recognition unit 10, a natural language understanding unit 11, an input signal acquisition unit 12, a constituent identification unit 13, and a control unit 14. The speech recognition unit 10 recognizes speech input from a user acquired by the HMI 6 and converts it into linguistic information such as text. The speech recognition unit 10 outputs the linguistic information generated by converting the speech input to the natural language understanding unit 11.

[0016] The natural language understanding unit 11 analyzes the linguistic information output from the speech recognition unit 10 by natural language processing, and extracts keywords related to the user's speech intention and the component 2 mentioned by the user. For example, the natural language understanding unit 11 extracts keywords indicating the location of the component 2 mentioned in the utterance content as keywords related to the component 2. The location indicated by the keyword extracted from the utterance content (i.e., the location indicating the location of the component 2) is an example of the "location of mention" described in the claims. For example, keywords and their synonyms may be defined in advance, and synonyms contained in the user's utterance may be converted into keywords.

[0017] For example, when a user utters "What's that light above the meter?" to ask about the meaning of a lit light, the natural language understanding unit 11 extracts the utterance intention "inquiry of meaning" that inquires about the meaning of the construct, and extracts "meter," "above," and "lamp" as keywords. In this case, for example, synonyms for the keyword "meter" may be predefined as "instrument," "indicator," "meter," etc., synonyms for the keyword "above" may be predefined as "above," "directly above," "above," etc., and synonyms for the keyword "lamp" may be predefined as "warning light," "indicator light," etc. In addition, the user's speech intention extracted by the natural language understanding unit 11 includes various speech intentions, such as ``inquiry about meaning'' and ``operation instructions'' (e.g., ``open the window'') that instruct the operation of the in-vehicle equipment, which is component 2.

[0018] Furthermore, for example, the natural language understanding unit 11 extracts, as keywords indicating the position of the component 2 mentioned in the utterance content, keywords indicating a reference position for indicating the position of the component 2 (hereinafter, sometimes referred to as a "reference position") and keywords indicating the relative position of the component 2 with respect to the reference position. For example, in the above example, "meter" is extracted as the keyword indicating the reference position, and "up" is extracted as the keyword indicating the relative position of the component 2 with respect to the reference position (position of the meter). For example, the keyword indicating the reference position may be a component other than the component 2 mentioned in the utterance content. For example, in the above utterance example, the component 2 mentioned in the utterance content is one of the lamps, and the component 2 serving as the reference position is the meter, which is a component 2 other than the lamp. 3 is a schematic diagram of an example of the arrangement of lamps and meters, which are an example of the component 2. A meter cluster 20 has a plurality of lamps and meters arranged therein, together with a tachometer 21 and a speedometer 22. Hereinafter, the meter cluster 20 will be simply referred to as "meter 20."

[0019] In the example of Fig. 3, a tire pressure warning light 30 is disposed as a lamp above the meter 20. Additionally, a fog lamp indicator light 31a, a high beam warning light 31b, a headlight indicator light 31c, and an engine warning light 31d are disposed as lamps at the upper left of the meter 20 and above the tachometer 21. Additionally, a VDC (Vehicle Dynamics Control) warning light 32 is disposed as a lamp at the upper right of the meter 20 and above the speedometer 22. Additionally, an auto brake hold indicator light 33 is disposed as a lamp at the lower left of the speedometer. Furthermore, an idling stop indicator light 34, a brake warning light 35a, an oil pressure warning light 35b, and a low water temperature indicator light 36a are arranged as lamps within the tachometer 21. An HEV power meter 36b is arranged as a meter within the tachometer 21. Additionally, lamps arranged within the speedometer 22 include a seat belt warning light 37, a hill descent control indicator light 38, a pedal misapplication collision prevention assist OFF indicator light 39a, and a remaining fuel warning light 39b. A fuel gauge 39c is also arranged within the speedometer 22.

[0020] For example, if a user utters "What's that light above the meter?" to ask about the meaning of the tire pressure warning light 30, the natural language understanding unit 11 may extract the keywords "meter," "above," and "light." The keyword "meter" is a keyword indicating a reference position for indicating the position of the component 2, and "above" is a keyword indicating the relative position of the component 2 with respect to the reference position (the position of the meter 20). For example, if a user utters "What is that light on the bottom left of the speedometer?" to ask about the meaning of the auto brake hold indicator light 33, the natural language understanding unit 11 may extract the keywords "speedometer," "bottom left," and "light." The keyword "speedometer" is a keyword indicating a reference position for indicating the position of the component 2, and "bottom left" is a keyword indicating the relative position of the component 2 with respect to the reference position (the position of the speedometer 22).

[0021] For example, if a user utters "What's the light below the exclamation mark light?" to ask about the meaning of the oil pressure warning light 35b, the natural language understanding unit 11 may extract the keywords "exclamation mark," "lamp," and "below." The keywords "exclamation mark" and "lamp" are keywords that indicate a reference position for indicating the position of the component 2, and in the example of FIG. 3, they indicate the brake warning light 35a. "under " is a keyword indicating the relative position of the component 2 with respect to the reference position (the position of the brake warning light 35a). For example, if a user utters "What is the rightmost lamp in the row of headlight indicator lights?" to inquire about the meaning of engine warning light 31d, the natural language understanding unit 11 may extract the keywords "light," "lamp," "row," and "rightmost." The keywords "light," "lamp," and "row" are keywords indicating a reference position for indicating the position of component 2, and in the example of FIG. 3, they indicate the arrangement of fog lamp indicator light 31a, high beam warning light 31b, and headlight indicator light 31c. "Rightmost" is a keyword indicating the relative position of component 2 with respect to the reference position (the position of the row of lamps 31a to 31c).

[0022] 4 is a schematic diagram showing an example of the arrangement of steering wheel switches, which is another example of the component 2. The steering wheel switches 41 to are switches provided on a steering wheel . For example, the steering wheel switches 41 to 43 are a group of switches for utilizing the autonomous driving control function of the vehicle 1. For example, the rightmost switch 41 on the right side of the steering wheel 40 is a main switch for turning on / off the autonomous driving control function of the vehicle 1. On the right The middle switch 42 is a set / coast switch that starts the autonomous driving control function. The leftmost switch 43 on the right side of the steering wheel 40 is a cancel switch 43 that cancels the autonomous driving control function. The cancel switch 43 is a switch that is located near the thumb of the user's right hand when the user is holding the steering wheel 40 (i.e., when the user places their fingers on the steering wheel 40).

[0023] Also, for example, steering wheel switches 44-4 6 are a group of switches for using the audio function of the vehicle 1. For example, the switch 44 at the bottom on the left side of the steering wheel 40 is a play / stop switch that instructs playback / stop of music by the audio function of the vehicle 1. The rightmost switch 45 on the left side of the steering wheel 40 is a volume switch that increases the volume of the audio function of the vehicle 1. The volume switch 45 is a switch located near the thumb of the user's left hand when the user holds the steering wheel 40. The leftmost switch 46 on the left side of the steering wheel 40 is a volume switch that decreases the volume of the audio function of the vehicle 1.

[0024] For example, if a user utters "What's that switch under the left side of the steering wheel?" to ask about the meaning of the playback / stop switch 44, the natural language understanding unit 11 may extract the keywords "steering wheel," "left side," "down," and "switch." The keyword "steering wheel" is a keyword indicating a reference position for indicating the position of the component 2, and "left side" and "down" are keywords indicating the relative position of the component 2 with respect to the reference position (the position of the steering wheel 40). Furthermore, for example, the keyword indicating the reference position may be the user's fingers when the user grips the steering wheel 40. For example, when the user utters "What switch is that around the thumb of the right hand?" to ask about the meaning of the cancel switch 43, the natural language understanding unit 11 may extract the keywords "right hand," "thumb," and "switch." The keywords "right hand" and "thumb" are keywords indicating the reference position for indicating the position of the component 2, and are shown in FIG. 4 In the example shown, the thumb of the user's right hand is placed on the steering wheel 40.

[0025] For example, if a user utters "What's the switch to the right of the cancel switch?" to ask about the meaning of the set / coast switch 42, the natural language understanding unit 11 may extract the keywords "cancel," "switch," and "right." The keywords "cancel" and "switch" are keywords indicating the reference position for indicating the position of the component 2, and are shown in FIG. 4 In this example, the cancel switch 43 is shown. "Right" is a keyword that indicates the relative position of the component 2 with respect to the reference position (the position of the cancel switch 43).

[0026] Furthermore, the natural language understanding unit 11 may auxiliary extract a keyword indicating the state of the component 2. For example, when a user utters "What light just came on?" to inquire about the meaning of a warning light, the natural language understanding unit 11 may extract "on" as a keyword indicating the state of the component 2. Also, when a user utters "What was that beep in the front left?" to inquire about the meaning of an alarm sound output by an alarm device, the natural language understanding unit 11 may extract a keyword "beeping" indicating the state of the component 2. The natural language understanding unit 11 outputs the extracted information on the intention of the utterance and the extracted information on the keywords to the constituent identification unit 13 .

[0027] See Fig. 2. The input signal acquisition unit 12 acquires, as an input signal, a control signal for the component 2 (on-vehicle equipment) generated by the on-vehicle equipment controller 3. For example, the control signal may be a lamp on / off signal. Alternatively, the control signal may be a signal instructing an alarm device to output or stop an alarm sound. Alternatively, the control signal may be a drive signal for opening or closing a window provided in a door of the vehicle 1, or a status signal indicating the open or closed state of the window. Furthermore, the input signal acquisition unit 12 acquires the output signal of the in-vehicle sensor 4 as an input signal. The input signal acquisition unit 12 converts the acquired control signals of the components 2 and the output signals of the in-vehicle sensors 4 into a specific data format determined in advance to represent the detected situation.

[0028] For example, the input signal acquiring section 12 may convert the control signal into flag information and set the value of the flag according to the control state of the component 2 . For example, the flag information may be converted to flag information that is set to the value "True" when the target lamp is on and set to the value "False" when the target lamp is off. Alternatively, the flag information may be converted to flag information that is set to the value "True" when an alarm device has been activated and an alarm sound has been output, and set to the value "False" when the alarm device is not activated. Alternatively, the flag information may be converted to flag information that is set to the value "True" when a window is open and set to the value "False" when the window is closed.

[0029] The input signal acquiring unit 12 may also convert the output signal of the in-vehicle sensor 4 into flag information and set the value of the flag according to the state and position of the object detected by the in-vehicle sensor 4. For example, a flag may be set according to the seating position of the user inside the vehicle detected based on the output signal of the in-vehicle sensor 4, such as an in-vehicle camera, a pressure sensor, a seat belt sensor, a biological sensor, etc. For example, the value of the flag may be set to "True" when the user is sitting in the driver's seat, and "False" when the user is sitting in the passenger seat. The input signal acquisition unit 12 outputs the converted input signal (hereinafter simply referred to as “input signal”) to the component identification unit 13.

[0030] The constituent identification unit 13 receives the information on the intention of the utterance and the information on the keywords output from the natural language understanding unit 11. The constituent identification unit 13 estimates the constituent 2 mentioned in the utterance content based on the keywords indicating the location of the constituent 2 output from the natural language understanding unit 11. Hereinafter, the constituent 2 mentioned in the utterance content will be referred to as the "target constituent." For example, the component identification unit 13 may estimate the target component by referring to component position information, which is the position where each component 2 is provided. For example, the storage device 9 of the voice recognition device 7 may store component position information 15, which is information on the component position.

[0031] 5 is a schematic diagram of an example of the component location information 15. The component location information 15 stores multiple rows of records. Each record stores information about a component and keywords related to the component. That is, the component location information 15 stores the information about the component and keywords related to the component in association with each other. The keywords stored in the component location information 15 include at least a keyword indicating the location of the component as component location information. The component identification unit 13 estimates, as the target component, the component 2 stored in the component location information 15 in association with a keyword that matches (e.g., matches) the keyword output from the natural language understanding unit 11. That is, the component 2 provided at a component location that matches (e.g., matches) the reference location mentioned in the utterance content is estimated to be the target component.

[0032] For example, assume that a user utters "What's that light above the meter?" and the natural language understanding unit 11 extracts the keywords "meter," "above," and "lamp." "Meter" and "above" are keywords that indicate the location of the mention. The component identification unit 13 refers to the component location information 15, selects the record in the first row that contains the same keywords as the keywords "meter," "above," and "lamp" extracted by the natural language understanding unit 11, and estimates that the tire pressure warning light 30 in the record in the first row is the target component. Also, for example, suppose that a user utters "What is the light on the bottom left of the speedometer?" and the natural language understanding unit 11 extracts the keywords "speedometer," "bottom left," and "lamp." "Speedometer" and "bottom left" are keywords that indicate the location of the mention. The component identification unit 13 refers to the component location information 15, selects a record in the second row that includes the same keywords as the keywords "speedometer," "bottom left," and "lamp" extracted by the natural language understanding unit 11, and estimates that the auto brake hold indicator light 33 in the record in the second row is the target component.

[0033] Also, for example, assume that a user utters "What's the lamp below the exclamation mark lamp?" and the natural language understanding unit 11 extracts the keywords "exclamation mark," "lamp," and "down." "Exclamation mark," "lamp," and "down" are keywords that indicate the mention location. The component identification unit 13 refers to the component location information 15 and selects the record in the third row that contains the same keywords as the keywords "exclamation mark," "lamp," and "down" extracted by the natural language understanding unit 11, and estimates that the oil pressure warning light 35b in the record in the third row is the target component. Also, for example, suppose that a user utters, "What switch is under the left side of the steering wheel?" and the natural language understanding unit 11 extracts the keywords "steering wheel," "left side," "down," and "switch." "Steering wheel," "left side," and "down" are keywords that indicate the location of the mention. The constituent identification unit 13 refers to the constituent location information 15, selects the record in the fourth row that contains the same keywords as the keywords "steering wheel," "left side," "down," and "switch" extracted by the natural language understanding unit 11, and estimates that the play / stop switch 44 in the record in the fourth row is the target constituent.

[0034] Also, assume that the user utters, for example, "What switch is around the thumb of the right hand?" and the natural language understanding unit 11 extracts the keywords "right hand," "thumb," and "switch." "Right hand" and "thumb" are keywords indicating the reference position. The component identification unit 13 refers to the component position information 15, selects the record in the fifth row that includes the same keywords as the keywords "right hand," "thumb," and "switch" extracted by the natural language understanding unit 11, and estimates the cancel switch 43 in the record in the fifth row as the target component. Note that in this case, the natural language understanding unit 11 extracts the keywords "right hand" and "thumb" that indicate the reference position for indicating the position of component 2, but does not extract any keywords that indicate the relative position with respect to the reference position "right thumb." In this case, the relative position of the reference position with respect to the component position is "nearby," and there is no need to indicate the relative position using keywords.

[0035] Also, for example, suppose that a user utters, "What's the switch to the right of the cancel switch?" and the natural language understanding unit 11 extracts the keywords "cancel," "switch," and "right." "Cancel," "switch," and "right" are keywords that indicate the location of the mention. The constituent identification unit 13 refers to the constituent location information 15 and estimates that the target constituent is the set-coast switch 42, which is a constituent of the record in the sixth line that contains the same keywords as the keywords "cancel," "switch," and "right" extracted by the natural language understanding unit 11.

[0036] In addition, as in the case where a user utters "What light above the tachometer just came on?", the keywords indicating location, "tachometer" and "above", may correspond to multiple components 2 (in this example, fog lamp indicator light 31a, high beam warning light 31b, headlight indicator light 31c, and engine warning light 31d). In this case, the component identification unit 13 may estimate which of the multiple components 2 is the target component based on the input signal output from the input signal acquisition unit 12. For example, a keyword indicating the state of the component 2 may be extracted from the content of the utterance, and the component 2 in the same control state as the state indicated by the extracted keyword may be estimated to be the target component. In the above example, if the keyword "lit" indicating the state of component 2 is extracted from the utterance content "just turned on," the lamp that is in the lit state based on the input signal may be selected from among the fog lamp indicator light 31a, high beam warning light 31b, headlight indicator light 31c, and engine warning light 31d that correspond to the keywords "tachometer" and "up," and estimated to be the target component.

[0037] Furthermore, the constituent identification unit 13 may identify the reference position mentioned in the utterance content based on the time series of keywords acquired from the natural language understanding unit 11. For example, the constituent identification unit 13 may identify the reference position based on the time series of the target constituent estimated by the keywords acquired from the natural language understanding unit 11. For example, if the cancel switch 43 is estimated as the target constituent based on the utterance content "What switch is around the thumb of your right hand?" and then the user utters "What switch is on the right of that?", the position of the target constituent (cancel switch 43) estimated from the previous utterance content based on the keyword "that" indicating a demonstrative may be identified as the reference position, and the set-coast switch 42 located to the right of the reference position (cancel switch 43) may be estimated as the target constituent.

[0038] The component identification unit 13 may estimate the target component that presents the auditory information based on the position where the auditory signal of the component 2 that presents the auditory information to the user can be heard. For example, if the user utters "What sound did I hear from the speaker on the right?" and the natural language understanding unit 11 extracts the keywords "right side," "speaker," and "sound," the target component may be estimated to be a navigation system that presents route guidance as auditory information. For example, if a user utters, "What was that sound I heard from in front?" and the natural language understanding unit 11 extracts the keywords "front" and "sound," it may be estimated that the target component is an alarm device that presents an alarm sound to the user. For example, if a user utters, "What is that sound coming from the rear right?" and the natural language understanding unit 11 extracts the keywords "rear right" and "sound," the rear side vehicle approach warning device may be estimated to be the target component.

[0039] Furthermore, the component identification unit 13 may identify the relative position of the component 2 with respect to the user's seating position as the reference position. In this case, the component identification unit 13 determines the user's seating position based on the input signal output from the input signal acquisition unit 12. The component identification unit 13 may identify the reference position mentioned in the utterance content based on the determination result of the user's seating position and keywords indicating the relative position extracted from the user's utterance content. For example, if it is determined based on the input signal that the user is sitting in the driver's seat, the user utters "open this," and the keyword "here" indicating a relative position is extracted, the reference position is near the driver's seat. Therefore, the component identification unit 13 may estimate that the window on the driver's seat side is the target component.

[0040] For example, if it is determined that the user is seated in the driver's seat, the user utters "open the other side," and the keyword "opposite side" indicating the relative position is extracted, the reference position is near the passenger seat on the opposite side of the driver's seat in the vehicle width direction. Therefore, the component identification unit 13 may estimate that the passenger seat window is the target component. The constituent identification unit 13 may also identify the mention position based on the time series of the target constituent estimated by the keywords acquired from the natural language understanding unit 11. For example, assume that the window on the driver's seat side is estimated as the target constituent based on the utterance content "Open this," and then the user utters "Open the back too." In this case, the constituent identification unit 13 may identify the position of the target constituent (the window on the driver's seat side) estimated from the previous utterance content as the reference position, and may estimate the window behind the window on the driver's seat side as the target constituent based on the keyword "behind" indicating a relative position from the current utterance content.

[0041] For example, if it is determined that a user of a right-hand drive vehicle is sitting in the driver's seat and the user utters "the display on the left" or "the switch on the left," the reference position is the left side of the right seat, i.e., near the center in the vehicle width direction. Therefore, the component identification unit 13 may estimate that the switch or display located on the center console is the target component. In addition, when estimating the target component based on the relative position of the component 2 with respect to the user's seating position, for example, multiple records that differ depending on the user's seating position can be stored for each component 2, and a keyword indicating the relative position depending on the seating position can be stored in each record. The constituent identification unit 13 outputs information on the estimated target constituent and information on the utterance intention output from the natural language understanding unit 11 to the control unit 14.

[0042] The constituent identification unit 13 may estimate the target constituent by referring to a learning model 16 that has learned the constituent positions instead of the constituent position information 15. As the learning model 16, various classifiers can be used, such as a neural network or a rule-based (tree structure) inference model. When the learning model 16 is made to learn the position of a component, keywords indicating the component position (for example, a keyword for a reference position and a keyword for a relative position) are used as example data, and training data that combines the example data and a correct label (i.e., the target component) is provided to the learning model 16, and the model is trained to output a correct label for the example data.

[0043] When the input signal output by the input signal acquisition unit 12 is used to estimate the target component, the keyword and the input signal may be used as example data. 2 shows both the constituent location information 15 and the learning model 16, the speech recognition device 7 does not need to have both the constituent location information 15 and the learning model 16. When the constituent location information 15 is provided, the learning model 16 may be omitted, and when the learning model 16 is provided, the constituent location information 15 may be omitted.

[0044] The control unit 14 generates a response to the user's utterance based on the target constituent identified by the constituent identification unit 13 and the information on the speech intention extracted by the natural language understanding unit 11 and input via the constituent identification unit 13. For example, when the utterance intention extracted by the natural language understanding unit 11 is "inquiry about meaning," the control unit 14 may control the HMI 6 to output information about the estimated target constituent. For example, the control unit 14 outputs a response message notifying the information about the target constituent and a command signal to cause the HMI 6 to output the response message. The HMI 6 may output the audio information and text information of the response message from a speaker or display them on a display device.

[0045] The information about the target component may be, for example, functional information about the function of the target component. For example, if the target component is the main switch 41, a response message "This is a switch that turns the autonomous driving control function on / off" may be output as functional information. The information about the target component may be, for example, operation information about an operation for utilizing the function of the target component. For example, if the target component is the set / coast switch 42, a response message saying "To start the autonomous driving control function, turn on the main switch and then press the set / coast switch" may be output as operation information.

[0046] Furthermore, for example, when the utterance intention extracted by the natural language understanding unit 11 is an operation instruction to operate an in-vehicle device (e.g., "open a window"), the control unit 14 may operate the estimated target component. For example, when the utterance content is "open this," the control unit 14 outputs a command signal to the in-vehicle device controller 3 to open the window on the driver's seat where the user is seated. The in-vehicle device controller 3 opens the window on the driver's seat in accordance with the command signal. When activating the target component in response to the speech intention, the control unit 14 may output a notification from the HMI 16 prompting the user to input whether or not to activate the target component. For example, when the component identification unit 13 cannot uniquely determine the target component from the content of the user's utterance and estimates multiple candidates for the target component, a notification prompting the user to input whether or not to activate the estimated candidate may be output from the HMI 16. For example, when the user's utterance intention is to open a window and it is not possible to distinguish whether the target component is the driver's side window or the passenger's side window, a notification prompting the user to input whether or not to activate the target component may be output, such as "Do you want to open the driver's side window?"

[0047] (operation) FIG. 6 is a flowchart of an example of a speech recognition method according to an embodiment. In step S1, the HMI 6 receives a voice input from the user. In step S2, the speech recognition unit 10 recognizes speech input from the user and converts it into linguistic information such as text. The natural language understanding unit 11 analyzes the linguistic information output from the speech recognition unit 10 using natural language processing and extracts the user's intention in speaking. In step S3, the natural language understanding unit 11 extracts keywords indicating the location of the component 2 from the linguistic information output from the speech recognition unit 10.

[0048] In step S4, the natural language understanding unit 11 identifies the mention position, which is the position mentioned in the utterance content. In step S5, the constituent identification unit 13 estimates the target constituent mentioned in the utterance content based on the research location mentioned in the utterance content. In step S6, the control unit 14 generates a response to the user's utterance based on the target constituent identified by the constituent identification unit 13 and the information on the utterance intention extracted by the natural language understanding unit 11. Then, the process ends.

[0049] (Effects of the embodiment) (1) The voice recognition device 7 acquires the content of an utterance by a vehicle user and estimates a target component, which is a component mentioned in the utterance, from among multiple components that make up the vehicle. The voice recognition device 7 performs a process of identifying a reference position, which is a position mentioned in the utterance, based on the utterance content, and a process of referencing a storage device that stores component positions, which are positions where multiple components are located, or a learning model that has learned component positions, and estimating, as the target component, a component located at a component position that matches the identified reference position. This makes it possible to improve the accuracy of estimating the components mentioned in the user's speech content in speech recognition that estimates the components mentioned in the user's speech content out of the multiple components that make up the vehicle.

[0050] (2) The component may be a device installed inside the vehicle that presents visual information to the user. This makes it possible to estimate whether the device that presents visual information is mentioned in the content of the utterance. (3) The component may be a device installed inside the vehicle that presents auditory information to the user, thereby making it possible to estimate whether the device that presents auditory information is mentioned in the content of the utterance.

[0051] (4) The voice recognition device 7 may identify, as the mentioned position, the relative position of a component with respect to a meter provided on an instrument panel of the vehicle. The voice recognition device 7 may identify, as the mentioned position, the relative position of a component with respect to a lamp provided on an instrument panel of the vehicle. The voice recognition device 7 may identify, as the mentioned position, the relative position of a component with respect to a steering wheel of the vehicle. The voice recognition device 7 may identify, as the mentioned position, the relative position of a component with respect to a switch provided on the steering wheel of the vehicle. The voice recognition device 7 may identify, as the mentioned position, the relative position of a component with respect to the position of a user's finger when the user places the finger on the steering wheel of the vehicle. The voice recognition device 7 may also detect the user's position, which is the position of the user in the vehicle, and identify the mention position based on the speech content and the user's position. This allows the target constituent to be estimated using relative position keywords contained in the user's speech.

[0052] (5) The speech recognition device 7 may output information about the estimated target component. For example, the speech recognition device 7 may output function information about the function of the estimated target component. For example, the speech recognition device 7 may output operation information about an operation for utilizing the function of the estimated target component. This allows providing information about the constructs mentioned in the user's speech.

[0053] (6) The voice recognition device 7 may activate the estimated target component based on the estimation result of the target component. This allows the components of the vehicle to be activated by voice input. (7) The speech recognition device 7 may output a notification prompting the user to input whether or not to activate the target component. This allows the user to confirm the estimation result of the target component, for example, when the target component cannot be uniquely determined from the user's utterance content and multiple candidates for the target component are estimated.

[0054] All examples and conditional terms described herein are intended for educational purposes to aid the reader in understanding the present invention and the concepts provided by the inventor for the advancement of technology, and should be construed without limitation to the specifically described examples and conditions above, and the configuration of examples herein for illustrating the advantages and disadvantages of the present invention. Although the embodiments of the present invention have been described in detail, it should be understood that various changes, substitutions, and alterations can be made thereto without departing from the spirit and scope of the present invention. [Explanation of symbols]

[0055] 1...vehicle, 2...component, 3...in-vehicle equipment controller, 4...in-vehicle sensor, 6...human-machine interface, 7...speech recognition device, 8...processor, 9...storage device, 10...speech recognition unit, 11...natural language understanding unit, 12...input signal acquisition unit, 13...component identification unit, 14...control unit, 15...component position information, 16...learning model

Claims

1. A speech recognition method for acquiring a speech content of a vehicle user and estimating a target component that is a component mentioned in the speech content among a plurality of components that configure the vehicle, comprising: A process of identifying a mention position, which is a position mentioned in the utterance content, based on the utterance content; a process of referring to a storage device that stores component positions, which are positions where the plurality of components are respectively provided, or a learning model that has learned the component positions, and estimating, as the target component, a component that is provided at the component position that matches the identified reference position; A speech recognition method characterized by causing a controller to execute the above.

2. 2. The speech recognition method according to claim 1, wherein the component is a device provided inside the vehicle for presenting visual information to the user.

3. 2. The speech recognition method according to claim 1, wherein the component is a device provided in the vehicle for presenting auditory information to the user.

4. 3. The speech recognition method according to claim 1, wherein a relative position of the component with respect to a meter provided on an instrument panel of the vehicle is identified as the mention position.

5. 3. The speech recognition method according to claim 1, wherein a relative position of the component with respect to a lamp provided on an instrument panel of the vehicle is identified as the mentioned position.

6. 3. The speech recognition method according to claim 1, wherein a relative position of the component with respect to a steering wheel of the vehicle is identified as the mention position.

7. 3. The speech recognition method according to claim 1, wherein a relative position of the component with respect to a switch provided on a steering wheel of the vehicle is identified as the mention position.

8. 3. The speech recognition method according to claim 1, wherein the position of the component relative to the position of the user's finger when the user places the finger on a steering wheel of the vehicle is identified as the remark position.

9. 9. The speech recognition method according to claim 1, wherein the controller outputs information about the estimated target constituent.

10. The speech recognition method according to claim 9 , wherein the controller outputs functional information relating to the estimated function of the target constituent.

11. 10. The speech recognition method according to claim 9, wherein the controller outputs operation information relating to an operation for utilizing the estimated function of the target component.

12. 2. The speech recognition method according to claim 1, wherein the controller activates the estimated target component based on the result of the target component estimation.

13. 13. The speech recognition method according to claim 12, wherein the controller outputs a notification prompting the user to input whether or not to activate the target component.

14. The controller Detecting a user position, which is a position of the user in the vehicle; 14. The speech recognition method according to claim 1, wherein the mention position is identified based on the speech content and the user position.

15. A speech recognition device that acquires a speech content of a vehicle user and estimates a target component that is a component mentioned in the speech content among a plurality of components that configure the vehicle, A process of identifying a mention position, which is a position mentioned in the utterance content, based on the utterance content; a process of referring to a storage device that stores component positions, which are positions where the plurality of components are respectively provided, or a learning model that has learned the component positions, and estimating, as the target component, a component that is provided at the component position that matches the identified reference position; A speech recognition device comprising a controller that executes the above.

Citation Information

Patent Citations

  • Device and method for photographing around vehicle

    JP2006343829A

  • Vehicular voice recognition apparatus

    JP2015089697A

  • On-vehicle device

    JP2019127192A

  • Vehicle door control device

    JP2019183504A

  • Agent system, agent method, and program

    JP2020060861A