Voice recognition system based on intention action of passenger

By adopting multi-category machine learning algorithms and image data analysis in the speech recognition system, combining eye and body tracking algorithms, the accuracy of occupant identity recognition and instruction interpretation in the prior art is solved, and the accuracy of hands-free tasks and the continuous dialogue capabilities of the system are improved.

CN120236568APending Publication Date: 2025-07-01GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311836062.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing voice recognition system has problems in accurately identifying the occupant's identity and interpreting instructions, especially in the influence of background noise, which makes it difficult to identify wake-up instructions, resulting in the occupant having to repeat the instructions multiple times and unable to achieve continuous dialogue.

Method used

Multi-category machine learning algorithms are used to combine image data to determine occupant intention factors through eye and body tracking algorithms, regression machine learning algorithms determine intention actions, and hands-free tasks are determined based on context and pattern recognition algorithms.

Benefits of technology

The accuracy of the speech recognition system in determining hands-free tasks is improved, the need for repeated instructions for occupants is reduced, and continuous dialogue with the speech recognition system is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236568A_ABST
    Figure CN120236568A_ABST
Patent Text Reader

Abstract

A speech recognition system includes one or more controllers that receive electrical signals representative of speech signals generated by an occupant and image data representative of a head and an upper body of the occupant. The controller converts electrical signals representing words spoken by the occupant into a sequence of markers based on a supervised multi-category machine learning algorithm, generates one or more statements based on the sequence of markers, and executes one or more eye and body tracking algorithms to determine one or more occupant intent factors. The controller determines an intent action of the occupant based on the occupant intent factor and a context of the speech signal generated by the occupant. The controller determines a hands-free task based on a context of the occupant-generated speech signal, the intent action, and the one or more statements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a speech recognition system that determines a hands-free task based at least on a speech signal created by a vehicle occupant and an intended action of the occupant, where the intended action is determined based on image data captured by an occupant monitoring system. Background Art

[0002] Many vehicles include an in-vehicle speech recognition system that allows a driver or occupant of the vehicle to interact with various vehicle technologies based on voice commands. Although speech recognition systems allow hands-free operation of various vehicle technologies, it should be understood that speech recognition systems also have some drawbacks. For example, some speech recognition systems may have problems accurately identifying the occupant's identity. As another example, some speech recognition systems may also have problems accurately interpreting the commands issued by the occupant. In addition, when one of multiple occupants issues a wake-up command, some speech recognition systems may have difficulty recognizing it due to background noise. As a result, the occupant may have to issue the wake-up command multiple times, thus preventing a continuous conversation with the speech recognition system.

[0003] Therefore, although current speech recognition systems achieve their intended purpose, there is a need in the art to improve the accuracy when a speech recognition system determines a hands-free task. Summary of the Invention

[0004] According to several aspects, a speech recognition system is disclosed that includes one or more controllers, each controller including one or more processors that execute instructions to receive an electrical signal representative of a speech signal generated by an occupant and image data representative of the occupant's head and upper body. The one or more controllers convert the electrical signal representative of the words spoken by the occupant into a sequence of tokens based on a supervised multi-class machine learning algorithm, where the sequence of tokens includes two or more tokens. The one or more controllers generate one or more statements based on the sequence of tokens. The one or more controllers execute one or more eye and body tracking algorithms to determine one or more occupant intent factors based on the image data representative of the occupant's head and upper body. The one or more controllers execute one or more regression machine learning algorithms to determine the intended action of the occupant based on one or more of the occupant intent factors. The one or more controllers determine the context of the speech signal generated by the occupant based on the intended action, the one or more statements, and the mood of the occupant. The one or more controllers execute one or more pattern recognition algorithms to determine a hands-free task based on the context of the speech signal generated by the occupant, the intended action, and the one or more statements.

[0005] In another aspect, a speech recognition system includes one or more peripheral systems in electronic communication with one or more controllers, wherein one or more processors of the one or more controllers direct one of the peripheral systems to perform a hands-free task.

[0006] In yet another aspect, an occupant is located within an interior passenger compartment of a vehicle.

[0007] In one aspect, the one or more peripheral systems include one or more of the following: a heating, ventilation, and air conditioning (HVAC) system, a radio, an autonomous driving system, a navigation system, an infotainment system, a lighting system, a personal electronic device, and a smart seat system that communicates with the occupant based on haptic feedback.

[0008] In another aspect, the speech recognition system further includes a microphone in electronic communication with the one or more controllers, the microphone converting a speech signal generated by the occupant into an electrical signal representative of the speech signal.

[0009] In yet another aspect, one or more processors of the one or more controllers execute instructions to continuously monitor the microphone to obtain an electrical signal representative of a speech signal generated by the occupant.

[0010] In one aspect, the speech recognition system further includes an occupant monitoring system, the occupant monitoring system including an occupant monitoring system camera in electronic communication with the one or more controllers, wherein the occupant monitoring system camera is positioned to capture image data representative of the occupant's head and upper body.

[0011] In another aspect, a confidence level is assigned to each token.

[0012] In yet another aspect, one or more processors of the one or more controllers execute instructions to compare the confidence level of each token that is part of a token sequence with a threshold confidence level; in response to determining that the confidence level of a particular token of the token sequence is less than the threshold confidence level, perform a mask on the particular token to create a missing token; execute one or more large language models to predict the content of the missing token based on the context of adjacent tokens that are part of the token sequence; and determine the content of the missing token based on one or more machine learning algorithms to complete one or more statements.

[0013] In one aspect, the large language model is a Bidirectional Encoder Representations from Transformers (BERT) model.

[0014] In another aspect, the one or more machine learning algorithms are Long Short-Term Memory (LSTM) models.

[0015] In yet another aspect, the occupant intent factors include one or more of the following: the occupant's gaze point, touch point, one or more gestures, and body position.

[0016] In one aspect, the context of a voice signal generated by an occupant is determined based on one or more of the following: current traffic conditions, current date, current time, and session history.

[0017] In another aspect, one or more processors of one or more controllers execute instructions to determine the mood of an occupant by analyzing a voice signal generated by the occupant based on a trained regression model.

[0018] In yet another aspect, one or more processors of one or more controllers execute instructions to execute one or more history-based large language models to predict upcoming voice commands issued by the occupant based on the occupant's session history.

[0019] In one aspect, a method for determining a hands-free task via a voice recognition system is disclosed. The method includes receiving, by one or more controllers, an electrical signal representative of a voice signal generated by an occupant and image data representative of the occupant's head and upper body. The method includes converting, by one or more controllers, the electrical signal representative of the words spoken by the occupant into a token sequence based on a supervised multi-class machine learning algorithm, wherein the token sequence includes two or more tokens. The method includes generating, by one or more controllers, one or more utterances based on the token sequence. The method further includes executing, by one or more controllers, one or more eye and body tracking algorithms to determine one or more occupant intent factors based on the image data representative of the occupant's head and upper body. The method further includes executing, by one or more controllers, one or more regression machine learning algorithms to determine the intended action of the occupant based on one or more of the occupant intent factors. The method includes determining the context of the voice signal generated by the occupant based on the intended action, one or more utterances, and the mood of the occupant. Finally, the method includes executing one or more pattern recognition algorithms to determine the hands-free task based on the context of the voice signal generated by the occupant, the intended action, and one or more utterances.

[0020] In another aspect, the method includes instructing a peripheral system to perform the hands-free task.

[0021] In yet another aspect, a speech recognition system for a vehicle is disclosed. The speech recognition system includes a microphone that converts a speech signal generated by a vehicle occupant into an electrical signal representative of the speech signal, an occupant monitoring system including an occupant monitoring system camera positioned to capture image data representative of the head and upper body of the occupant, and one or more controllers in electronic communication with the microphone and the occupant monitoring system camera. Each of the one or more controllers includes one or more processors that execute instructions to convert the electrical signal representative of the words spoken by the occupant into a sequence of tokens based on a supervised multi-class machine learning algorithm, wherein the sequence of tokens includes two or more tokens. The one or more controllers generate one or more statements based on the sequence of tokens. The one or more controllers execute one or more eye and body tracking algorithms to determine one or more occupant intent factors based on the image data representative of the head and upper body of the occupant. The one or more controllers execute one or more regression machine learning algorithms to determine an intended action of the occupant based on one or more of the occupant intent factors. The one or more controllers determine the context of the speech signal generated by the occupant based on the intended action, the one or more statements, and the mood of the occupant, and execute one or more pattern recognition algorithms to determine a hands-free task based on the context of the speech signal generated by the occupant, the intended action, and the one or more statements.

[0022] In another aspect, the speech recognition system further includes one or more peripheral systems in electronic communication with the one or more controllers, wherein one or more processors of the one or more controllers direct one of the peripheral systems to perform a hands-free task.

[0023] In yet another aspect, the one or more peripheral systems include one or more of the following: a heating, ventilation, and air conditioning (HVAC) system, a radio, an autonomous driving system, a navigation system, an infotainment system, a lighting system, a personal electronic device, and an intelligent seat system that communicates with the occupant based on tactile feedback.

[0024] From the description provided herein, further application areas will become apparent. It should be understood that the specification and specific examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way.

[0026] Figure 1 A schematic illustration of a vehicle including the disclosed speech recognition system according to an exemplary embodiment is shown, the speech recognition system including one or more controllers in electronic communication with a microphone and an occupant monitoring system;

[0027] Figure 2is a block diagram showing the software architecture of one or more controllers according to an exemplary embodiment; and Figure 1 as shown; and

[0028] Figure 3 is a process flow diagram showing a method for determining a hands - free task by the disclosed speech recognition system according to an exemplary embodiment; Figure 1 as shown. DETAILED DESCRIPTION

[0029] The following description is merely exemplary in nature and is not intended to limit the disclosure, application, or use.

[0030] Referring to Figure 1 , a vehicle 10 including the disclosed speech recognition system 12 is shown. It should be understood that the vehicle 10 can be any type of vehicle, such as but not limited to a sedan, a truck, a sport utility vehicle, a van, or a recreational vehicle. In a non - limiting embodiment as Figure 1 shown, the speech recognition system 12 includes one or more controllers 20 that are in electronic communication with a plurality of sensing sensors 22, a microphone 24, an occupant monitoring system 26, one or more peripheral systems 28 for performing hands - free tasks, and a speaker 30. It should be understood that although Figure 1 the speech recognition system 12 is shown as part of a vehicle, the speech recognition system 12 is not limited to vehicles and can also be used in various other applications. For example, in another embodiment, the speech recognition system 12 can be used in a building such as a home or an office.

[0031] As described below, one or more controllers 20 of the speech recognition system 12 direct one or more peripheral systems 28 to perform hands - free tasks indicated by one or more individuals or occupants 38 located within the interior cabin 16 of the vehicle 10. In one embodiment, one or more peripheral systems 28 include any vehicle system or subsystem, such as but not limited to a heating, ventilation, and air conditioning (HVAC) system, a radio, an autonomous driving system, a navigation system, an infotainment system, a lighting system, and a smart seat system that communicates with the occupant 38 based on haptic feedback. In the case where the speech recognition system 12 is part of a building such as a residence, the occupant 38 is instead located within a room or other enclosed space within the building, and one or more peripheral systems 28 can include a lighting system, household appliances such as a television or a refrigerator, and an HVAC system. In one embodiment, one or more peripheral systems 28 can include personal electronic devices of the occupant 38 of the vehicle 10, where the personal electronic devices are wirelessly connected to one or more controllers 20. The portable electronic device can be, for example, a smart phone, a smart watch, or a tablet computer.

[0032] A hands-free task is any type of operation that an occupant 38 would traditionally perform using his or her hands, but now the voice recognition system 12 instructs one or more peripheral systems 28 to perform the hands-free task without the occupant 38 manually performing the operation. For example, if the peripheral system 28 is a radio, the hands-free task may include turning on the radio, selecting a specific audio file of the radio to play, or selecting a specific radio channel or station. In another example, if the peripheral system 28 is a smart phone, the hands-free task is to send a text message or make a phone call. The disclosed voice recognition system 12 determines the hands-free task based at least on a voice signal created by an occupant 38 of the vehicle 10 and the intended actions of the occupant 38 determined by the occupant monitoring system 26. As described below, the hands-free task may also be determined based on other inputs as well as traffic conditions, date and time, and conversation history. The voice signal is captured by a microphone 24, and the intended actions of the occupant 38 are determined based on image data captured by an occupant monitoring system camera 54 that is part of the occupant monitoring system 26.

[0033] A plurality of sensing sensors 22 are configured to collect sensing data representative of an external environment 14 around the vehicle 10. In Figure 1 the illustrated non-limiting embodiment, the plurality of sensing sensors 22 includes one or more cameras 40 that capture image data representative of the external environment 14, an inertial measurement unit (IMU) 42, a global positioning system (GPS) 44, a radar 46, and a lidar 48. However, it should be understood that additional sensors may also be used. The microphone 24 represents a device that converts sound waves into an electrical signal, where the electrical signal is received by one or more controllers 20. Specifically, the microphone 24 converts a voice signal generated by an occupant 38 of the vehicle 10 into an electrical signal representative of the voice signal. The occupant monitoring system 26 includes an occupant monitoring system camera 54 that is positioned to capture image data representative of the head and upper body of an occupant 38 of the vehicle 10.

[0034] Figure 2 is shown Figure 1 The software architecture of one or more controllers 20 shown. One or more controllers 20 of the voice recognition system 12 include a voice box 70 and an intent box 72. The voice box 70 of one or more controllers 20 includes a noise reduction module 80, a voice recognition module 82, a tag generation module 84, a mask module 86, a prediction module 88, and a statement generation module 90. The intent box 72 of one or more controllers 20 includes a behavior detection module 92, an intent module 94, a context module 96, a response generation module 98, and a prediction module 100.

[0035] The voice box 70 of one or more controllers 20 receives, as an input, an electrical signal representative of a voice signal from the microphone 24, where the voice signal indicates an occupant 38 ( Figure 1)One or more words spoken. As described below, the speech box 70 determines, based on the speech signal, one or more utterances representative of the words spoken by the occupant 38. The noise reduction module 80 of the speech box 70 continuously monitors the microphone 24 to obtain an electrical signal representative of the speech signal generated by the occupant 38. Thus, it can be understood that the speech recognition system 12 does not require an individual to issue an activation or wake-up command. The noise reduction module 80 performs one or more noise reduction algorithms that reduce background noise from the electrical signal representative of the speech signal generated by the occupant 38. An example of a noise reduction algorithm that can be used is Fourier analysis.

[0036] The speech recognition module 82 of the speech box 70 receives the electrical signal representative of the speech signal generated by the occupant 38 from the noise reduction module 80. The speech recognition module 82 performs one or more background noise recognition algorithms that extract the background noise from the electrical signal representative of the speech signal generated by the occupant 38. Some examples of background noise include, but are not limited to, engine noise, road noise based on a particular type of road material, ambient noise, or music or other sound files transmitted by a radio. Ambient noise can include background noise from sources such as highways, airports, shopping areas, and urban areas. An example of a background noise recognition algorithm is a machine learning-based model that is trained to identify and extract background noise from the electrical signal representative of the speech signal generated by the occupant 38.

[0037] The speech recognition module 82 also performs one or more speaker recognition algorithms that determine when more than one individual or occupant of the vehicle 10 generates a speech signal. Then, in response to determining that more than one individual generates a speech signal, one or more speaker recognition algorithms identify the different individuals by their respective identities 102. In the example shown, there is a first individual A, a second individual B, and a third individual C.

[0038] The token generation module 84 of the speech box 70 converts the electrical signal representative of the words spoken by the occupant 38 received from the speech recognition module 82 into a token sequence based on a supervised multi-class machine learning algorithm, where the sequence includes two or more tokens. Each token represents a word or part of a word, or punctuation. In another embodiment, the token is an index number mapped to a word database. It should be understood that each token in the token sequence is assigned a confidence level, where a higher confidence level represents that the token accurately represents the word spoken by the occupant 38( Figure 1 )spoken words.

[0039] Next, the masking module 86 of the speech box 70 compares the confidence level of each token that is part of the token sequence with a threshold confidence level. The threshold confidence level is based on the target accuracy of the speech recognition system 12. In response to determining that a particular token in the token sequence includes a confidence level less than the threshold confidence level, the masking module 86 of the speech box 70 performs masking on the particular token to create a missing token.

[0040] Next, the prediction module 88 of the speech box 70 executes one or more large language models to predict the content of the missing token based on the context of neighboring tokens that are part of the token sequence. An example of a large language model that can be used is a Bidirectional Encoder Representations from Transformers (BERT) model. However, it should be understood that other large language models can also be used. It should be understood that in some embodiments, the content of the missing token may not be accurately predicted based on the large language model.

[0041] The statement generation module 90 of the speech box 70 generates one or more statements based on the token sequence received from the large language model. In the case where the token sequence includes a missing token, the statement generation module 90 can determine the content of the missing token based on one or more machine learning algorithms to complete one or more statements. Specifically, in one embodiment, the statement generation module 90 determines the content of the missing token based on a Long Short-Term Memory (LSTM) model to complete one or more statements representing the words spoken by the occupant 38.

[0042] The intent box 72 of one or more controllers 20 receives as input one or more statements determined by the speech box 70 and image data representing the head and upper body of the occupant 38 of the vehicle 10 ( Figure 1 ) from the occupant monitoring system camera 54. The intent box 72 determines a hands-free command based at least on the speech signal created by the occupant 38 of the vehicle 10 and the image data captured by the occupant monitoring system camera 54 of the occupant monitoring system 26. As described below, the hands-free task can also be determined based on other inputs such as traffic conditions, date and time, and conversation history. In one embodiment, the intent box 72 can determine the hands-free task without a voice-based input from the occupant 38. That is, in one embodiment, the intent box 72 can determine the hands-free task based on the image data captured by the occupant monitoring system camera 54 of the occupant monitoring system 26 without a speech signal created by the occupant 38 of the vehicle 10.

[0043] The behavior detection module 92 of the intent box 72 receives, as input, image data representing the head and upper body of the occupant 38 captured by the occupant monitoring system camera 54 of the occupant monitoring system 26. The behavior detection module 92 of the intent box 72 executes one or more eye and body tracking algorithms to determine one or more occupant intent factors based on the image data representing the head and upper body of the occupant 38. The occupant intent factors can include one or more of the following: the gaze point of the occupant 38, the touch point, one or more gestures, and the body position. The gaze point of the occupant 38 indicates the movement of the eyes relative to the head and represents the position where the occupant 38 is looking. The touch point indicates the component that the occupant 38 is touching. For example, the occupant 38 can use his or her hand to manipulate the knob of the HVAC system to change the temperature inside the vehicle cabin. The gesture represents the movement of the head and hands of the occupant 38 that expresses an idea. The body position of the occupant 38 indicates the mental state of the occupant 38. For example, the body position can indicate when the occupant 38 is relaxed or anxious.

[0044] The intent module 94 of the intent box 72 determines the intent action of the occupant 38 based on one or more of the occupant intent factors (the gaze point, touch point, one or more gestures, and body position of the occupant 38 received from the behavior detection module 92) by executing one or more regression machine learning algorithms. The intent action of the occupant 38 can be expressed as an intent set, where the intent set indicates the intent action and at least one of the following: the touch point of the occupant 38, one or more gestures, and the body position, and is represented as: {intent action|gaze point|touch point|one or more gestures|body position}. For example, if the occupant 38 is anxious because it is too hot and wants to adjust the temperature inside the vehicle cabin, the intent set can be represented as: {adjust the temperature inside the vehicle cabin|gaze at the HVAC button|the occupant's body is tense}.

[0045] The context module 96 of the intent box 72 receives at least the intent action from the intent module 94, one or more statements from the speech box 70, and an electrical signal representing the speech signal generated by the occupant 38 from the speech recognition module 82. As Figure 2 shown, in one embodiment, the context module 96 of the intent box 72 also receives the current traffic condition, date, and time from one or more remaining controllers 104 that are part of the vehicle 10. The traffic condition indicates the current traffic that the vehicle 10 is experiencing, and the date and time indicate the current date and the current time. In one embodiment, the context module 96 communicates electronically with one or more historical databases 106, where the historical databases 106 store the session history of the occupant 38. The session history of the occupant 38 indicates previous sessions that have been captured by the microphone 24 and analyzed by one or more controllers 20 to determine hands-free tasks.

[0046] The context module 96 of the intent box 72 executes one or more machine learning algorithms to determine the context of the electrical signal representing the voice signal generated by the occupant 38 based on the intent actions from the intent module 94, one or more statements from the voice box 70, the current traffic conditions (if applicable), the current date (if applicable), the current time (if applicable), the conversation history of the occupant 38 (if applicable), and the mood of the occupant. The machine learning algorithms can include, but are not limited to, the LSTM model or a prediction-based machine learning model. The context module 96 determines the mood of the occupant 38 by analyzing the electrical signal representing the voice signal generated by the occupant 38 based on a trained regression model. It should be understood that the trained regression model is trained based on the voice signals created by the occupant 38.

[0047] The response generation module 98 of the intent box 72 receives as inputs the context of the electrical signal representing the voice signal generated by the occupant 38 from the context module 96, the intent actions from the intent module 94, one or more statements from the voice box 70, and the conversation history of the occupant 38 (if applicable). The response generation module 98 of the intent box 72 executes one or more pattern recognition algorithms that determine a hands-free task based on the inputs within a restricted duration. In one embodiment, the restricted duration is approximately 10 milliseconds. Specifically, the pattern recognition algorithm compares the current values of the context of the electrical signal representing the voice signal generated by the occupant 38, the intent actions, one or more statements, and the conversation history of the occupant 38 with the previously determined hands-free tasks stored in one or more historical hands-free databases 108. The one or more historical hands-free databases 108 indicate, for each previously determined hands-free task, the corresponding context of the electrical signal representing the voice signal generated by the occupant 38, the corresponding intent actions, the corresponding one or more statements, and the corresponding conversation history of the occupant 38.

[0048] Then, the response generation module 98 instructs one or more peripheral systems 28 to execute the hands-free task. In one embodiment, the response generation module 98 can also instruct the speaker 30 to announce the hands-free task based on a synthetic or computer-generated audio output representing human speech.

[0049] The prediction module 100 of the intent box 72 executes one or more history-based large language models to predict upcoming voice commands issued by the occupant 38 based on the session history of the occupant 38 stored in one or more history databases 106. Thereafter, in one embodiment, the prediction module 100 of the intent box 72 may instruct the speaker 30 to announce the upcoming command. In the case where the upcoming voice command indicates that the occupant 38 is requesting a hands-free task, the prediction module 100 instructs a human-machine interface (HMI), such as a touch screen, to generate an instruction requesting the occupant 38 to confirm the hands-free task. In response to receiving the confirmation from the occupant 38, the prediction module 100 also instructs one or more peripheral systems 28 to execute the hands-free task.

[0050] Figure 3 is a process flow diagram showing a method 300 for a voice recognition system 12 to determine and execute a hands-free task. Generally referring Figures 1-3 to, method 300 may start at decision box 302. In box 302, the noise reduction module 80 of the voice box 70 continuously monitors the microphone 24 to obtain an electrical signal representative of the voice signal generated by the occupant 38. In response to receiving the electrical signal representative of the voice signal generated by the occupant 38, method 300 proceeds to box 304.

[0051] In box 304, the noise reduction module 80 executes one or more noise reduction algorithms that reduce background noise in the electrical signal representative of the voice signal generated by the occupant 38. Then, method 300 may proceed to box 306.

[0052] In box 306, the speech recognition module 82 of the voice box 70 executes one or more background noise recognition algorithms that extract the background noise in the electrical signal representative of the voice signal generated by the occupant 38. The speech recognition module 82 also executes one or more speaker recognition algorithms that determine when multiple speakers generate the voice signal. Then, method 300 may proceed to box 308.

[0053] In box 308, the tag generation module 84 of the voice box 70 converts the electrical signal representative of the words spoken by the occupant 38 received from the speech recognition module 82 into a tag sequence based on a supervised multi-class machine learning algorithm, where the sequence includes two or more tags and each tag is assigned a confidence level. Then, method 300 may proceed to box 310.

[0054] In box 310, the masking module 86 of the voice box 70 compares the confidence level of each tag that is part of the tag sequence with a threshold confidence level. Then, method 300 may proceed to decision box 312.

[0055] In decision block 312, in response to determining that the confidence level of a particular token in the token sequence is less than a threshold confidence level, method 300 proceeds to block 314. In block 314, the masking module 86 of the speech box 70 masks the particular token to create a missing token. Otherwise, method 300 proceeds to block 320.

[0056] Then, in block 316, the prediction module 88 of the speech box 70 executes one or more large language models to predict the content of the missing token based on the context of neighboring tokens that are part of the token sequence. Then, method 300 can proceed to block 318.

[0057] In block 318, the statement generation module 90 of the speech box 70 determines the content of the missing token based on one or more machine learning algorithms to complete one or more statements. Then, method 300 can proceed to block 320.

[0058] In block 320, the statement generation module 90 of the speech box 70 generates one or more statements based on the token sequence. Then, method 300 can proceed to block 322.

[0059] In block 322, the behavior detection module 92 of the intent box 72 executes one or more eye and body tracking algorithms to determine one or more occupant intent factors based on image data representing the head and upper body of the occupant 38 from the occupant monitoring system camera 54. The occupant intent factors can include one or more of the following: the gaze point of the occupant 38, the touch point, one or more gestures, and the body position. Then, method 300 can proceed to block 324.

[0060] In block 324, the intent module 94 of the intent box 72 executes one or more regression machine learning algorithms to determine the intent action of the occupant 38 based on one or more of the occupant intent factors. Then, method 300 can proceed to block 326.

[0061] In block 326, the context module 96 of the intent box 72 determines the context of the speech signal generated by the occupant 38 based on the intent action from the intent module 94, one or more statements from the speech box 70, the current traffic conditions (if applicable), the current date (if applicable), the current time (if applicable), the session history of the occupant 38 (if applicable), and the mood of the occupant. Then, method 300 can proceed to block 328.

[0062] In block 328, the response generation module 98 of intent box 72 executes one or more pattern recognition algorithms to determine a hands-free task based on the context of the electrical signal representative of the voice signal generated by occupant 38 from context module 96, the intent action from intent module 94, one or more statements from speech box 70, and the conversation history of occupant 38 if applicable. Then, method 300 can proceed to block 330.

[0063] In block 330, the response generation module 98 of intent box 72 instructs one or more peripheral systems 28 to perform the hands-free task. In one embodiment, the response generation module 98 can also instruct speaker 30 to announce the hands-free task based on computer-generated audio output representative of human speech. Then, method 300 can proceed to block 332.

[0064] In block 332, the prediction module 100 of intent box 72 executes one or more history-based large language models to predict upcoming voice commands issued by occupant 38 based on the conversation history of occupant 38 stored in one or more history databases 106. In an embodiment, the prediction module 100 of intent box 72 instructs speaker 30 to announce the upcoming command. In the case where the upcoming voice command indicates that occupant 38 is requesting a hands-free task, the prediction module 100 instructs the HMI to generate an instruction requesting occupant 38 to confirm the hands-free task. In response to receiving the confirmation from occupant 38, the prediction module 100 also instructs one or more peripheral systems 28 to perform the hands-free task. Then, method 300 can terminate.

[0065] Referring generally to the drawings, the disclosed speech recognition system has various technical effects and benefits. Specifically, the speech recognition system provides a method for determining a hands-free task based on an occupant's utterance in combination with the occupant's intent determined based on non-verbal input. In particular, the intent of the occupant is determined based on image data representative of the occupant's head and upper body. It should also be understood that the speech recognition system continuously monitors the speech of the occupant, so the disclosed speech recognition system does not require an individual to issue an activation or wake-up command. Instead, the speech recognition system can naturally intervene and assist an occupant who is driving or performing another task related to vehicle operation. When determining a hands-free task, the speech recognition system can also consider other inputs, such as traffic conditions, the current date and time, and the conversation history of the occupant.

[0066] The controller can refer to an electronic circuit, combinational logic circuit, field programmable gate array (FPGA), a processor (shared, dedicated, or group) that executes code, or some or all combinations of the above, or as part of them, such as in a system-on-chip. Additionally, the controller can be microprocessor-based, such as a computer having at least one processor, memory (RAM and / or ROM), and associated input and output buses. The processor can operate under the control of an operating system resident in the memory. The operating system can manage computer resources such that computer program code instantiated as one or more computer software applications (such as applications resident in the memory) can have instructions executed by the processor. In an alternative embodiment, the processor can directly execute the application, in which case the operating system can be omitted.

[0067] The description of the present disclosure is merely exemplary in nature, and variations that do not depart from the gist of the present disclosure are intended to be within the scope of the present disclosure. Such variations should not be regarded as a departure from the spirit and scope of the present disclosure.

Claims

1. A voice recognition system, comprising: One or more controllers, each of the controllers including one or more processors, the processors executing instructions to: Receive an electrical signal representing a voice signal generated by an occupant and image data representing the head and upper body of the occupant; Convert the electrical signal representing the word spoken by the occupant into a sequence of tokens based on a supervised multi-class machine learning algorithm, wherein the sequence of tokens includes two or more tokens; Generate one or more statements based on the sequence of tokens; Execute one or more eye and body tracking algorithms to determine one or more occupant intent factors based on the image data representing the head and upper body of the occupant; Execute one or more regression machine learning algorithms to determine the intended action of the occupant based on one or more of the occupant intent factors; Determine the context of the voice signal generated by the occupant based on the intended action, the one or more statements, and the mood of the occupant; and Execute one or more pattern recognition algorithms to determine a hands-free task based on the context of the voice signal generated by the occupant, the intended action, and the one or more statements.

2. The voice recognition system according to claim 1, further comprising: One or more peripheral systems in electronic communication with the one or more controllers, wherein the one or more processors of the one or more controllers direct one of the peripheral systems to perform the hands-free task.

3. The speech recognition system according to claim 2, wherein, The occupant is located in the interior compartment of a vehicle.

4. The speech recognition system according to claim 3, wherein, The one or more peripheral systems include one or more of the following: a heating, ventilation, and air conditioning (HVAC) system, a radio, an autonomous driving system, a navigation system, an infotainment system, a lighting system, a personal electronic device, and a smart seat system that communicates with the occupant based on haptic feedback.

5. The voice recognition system according to claim 1, further comprising: A microphone in electronic communication with the one or more controllers, the microphone converting the voice signal generated by the occupant into the electrical signal representing the voice signal.

6. The speech recognition system according to claim 5, wherein, The one or more processors of the one or more controllers execute instructions to: Continuously monitor the microphone to obtain the electrical signal representing the voice signal generated by the occupant.

7. The voice recognition system according to claim 1, further comprising: An occupant monitoring system including an occupant monitoring system camera in electronic communication with the one or more controllers, wherein the occupant monitoring system camera is positioned to capture image data representing the head and upper body of the occupant.

8. The speech recognition system according to claim 1, wherein, Assign a confidence level to each token.

9. The speech recognition system according to claim 8, wherein, The one or more processors of the one or more controllers execute instructions to: Compare the confidence level of each token that is part of the sequence of tokens with a threshold confidence level; In response to determining that the confidence level of a particular token of the sequence of tokens is below the threshold confidence level, mask the particular token to create a missing token; Execute one or more large language models to predict the content of the missing token based on the context of adjacent tokens that are part of the sequence of tokens. and determining the content of the missing token based on one or more machine learning algorithms to complete the one or more statements.

10. The speech recognition system according to claim 9, wherein, The large language model is a Transformer-based Bidirectional Encoder Representations from Transformers (BERT) model.