Adaptive speech recognition system and method for integrating auditory and non-auditory inputs

CN116741158BActive Publication Date: 2026-09-08GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211302154.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-03
Filing Date
2022-10-24
Publication Date
2026-09-08
Estimated Expiration
2042-10-24

Smart Images

  • Figure CN116741158B_ABST
    Figure CN116741158B_ABST
Patent Text Reader

Abstract

A voice recognition method includes receiving audible data and user data. The audible data includes information about an utterance made by a user. The user data includes information about a movement made by the user. The method also includes fusing the audible data and the user data to obtain fused data, and determining at least one spoken word of the utterance based on the fused data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to adaptive speech recognition systems and methods for integrating auditory and non-auditory inputs. Background Technology

[0002] This introduction provides a general overview of the context of this disclosure. Within the scope described in this introduction, the work of the currently named inventor, and aspects of the description that may not conform to the prior art at the time of filing, are neither expressly nor impliedly acknowledged as prior art to this disclosure.

[0003] Some vehicles include voice recognition systems. In these vehicles, users can speak words or phrases, and the voice recognition system will recognize them accordingly. Once the spoken words are recognized, the vehicle or another device can perform a specific task.

[0004] Public content

[0005] While speech recognition systems are highly useful in vehicles and other devices, recognizing speech can be challenging when users have speech impairments, accents, or other speech-related special needs. Therefore, it is highly valuable to develop speech recognition systems and methods that take into account the specific speech-related needs of users to improve the quality of speech recognition for each individual user.

[0006] This disclosure describes a speech recognition system and method that uses auditory input (i.e., audible data) and non-auditory input (i.e., user data) for speech recognition. By fusing auditory input (i.e., audible data) and non-auditory input (i.e., user data), the speech recognition system improves the quality of speech recognition. In this way, the speech recognition system can recognize the auditory and non-auditory speech patterns of a specific user. For example, the currently disclosed system can use the user's utterances (i.e., auditory input) and images of that specific user's facial cues (i.e., non-auditory input) to determine the words spoken by the user, thereby improving the speech recognition quality of the system when the user has specific speech-related needs.

[0007] In one aspect of this disclosure, a speech recognition method includes receiving audible data and user data. The audible data includes information about utterances produced by the user. The user data includes information about movements made by the user. The method also includes fusing the audible data and user data to obtain fused data, and determining at least one spoken word of the utterance based on the fused data. Different techniques can be used for data fusion, such as data Bayesian networks, Dempster-Shafer theory, Bayesian filters, and / or neural networks.

[0008] In one aspect of this disclosure, the speech recognition method also includes storing a fine-tuned user profile in a cloud-based system. The user profile includes specific parameters for a personalized speech recognition system tailored to the user. The speech recognition method may further include sending the user profile to the vehicle once the user enters the vehicle, and utilizing the personalized speech recognition system to process audible data and user data.

[0009] In one aspect of this disclosure, the step of determining at least one spoken word of a utterance includes using a trained neural network to determine one or more spoken words of the utterance.

[0010] In one aspect of this disclosure, the method also includes fine-tuning a trained neural network based on fused data to adapt to the user's speech patterns.

[0011] In one aspect of this disclosure, audible data is received via a microphone in a vehicle.

[0012] In one aspect of this disclosure, audible data is received via a microphone on a mobile device.

[0013] In one aspect of this disclosure, user data is received via a camera on a vehicle.

[0014] In one aspect of this disclosure, user data is received via a camera on a mobile device.

[0015] In one aspect of this disclosure, user data includes at least one image of a user's facial expression.

[0016] In one aspect of this disclosure, user data includes at least one image of the user's lips.

[0017] In one aspect of this disclosure, user data and audible data are received via a camera and microphone of a mobile device, respectively. The method also includes transmitting the user data and audible data from the mobile device to a vehicle.

[0018] In one aspect of this disclosure, audible data and user data are received via a microphone and a camera of a vehicle, respectively. The method also includes transmitting the audible data and user data from the vehicle to a mobile device.

[0019] In one aspect of this disclosure, the method further includes transmitting user data and audible data from a mobile device to a vehicle. Additionally, the method includes transmitting user profiles to a fleet of vehicles via a cloud-based system. The user profiles include specific parameters of the voice recognition system for a particular user. Therefore, the vehicle fleet can personalize the voice recognition system.

[0020] This disclosure also describes a speech recognition system. In one aspect of this disclosure, the speech recognition system includes a first sensor configured to detect utterances from a user, a second sensor configured to detect movements made by the user, and a controller communicating with the first and second sensors. The controller is configured to receive audible data from the first sensor. The audible data includes information about utterances from the user. The controller is configured to receive user data from the second sensor. The user data includes information about movements made by the user. The controller is configured to fuse the audible data and the user data to obtain fused data, and to determine at least one spoken word of the utterance based on the fused data.

[0021] In one aspect of this disclosure, the controller is configured to determine at least one spoken word of a utterance by using a trained neural network.

[0022] In one aspect of this disclosure, the controller is configured to fine-tune a trained neural network based on fused data to adapt to the user's speech patterns.

[0023] In one aspect of this disclosure, the first sensor is a microphone, and the second sensor is a camera. Both the microphone and the camera are located in the vehicle.

[0024] In one aspect of this disclosure, user data includes at least one image of a user's facial expression.

[0025] In one aspect of this disclosure, user data includes at least one image of the user's lips.

[0026] In one aspect of this disclosure, the controller is configured to transmit user data and audible data from a mobile device to a vehicle. Furthermore, the controller is configured to transmit user profiles to the fleet via a cloud-based system. The user profiles include specific parameters of the voice recognition system for a particular user. Therefore, the vehicle fleet can personalize the voice recognition system.

[0027] In one aspect of this disclosure, the first sensor is a microphone, and the second sensor is a camera. Both the microphone and the camera are located in the mobile device. The controller is configured to command the mobile device to transmit user data and audible data to the vehicle.

[0028] In one aspect of this disclosure, the controller is also configured to store a fine-tuned user profile in a cloud-based system. The user profile includes specific parameters for a personalized voice recognition system tailored to the user. The controller is also configured to send the user profile to the vehicle once the user enters the vehicle and utilize the personalized voice recognition system to process audible and user data.

[0029] In one aspect of this disclosure, the controller is configured to use at least one of a Bayesian network, Dempster-Shafer theory, a Bayesian filter, or a neural network to fuse audible data and user data.

[0030] Other areas of application of this disclosure will become apparent from the detailed description provided below. It should be understood that the detailed description and specific examples are intended for illustrative purposes only and are not intended to limit the scope of this disclosure.

[0031] The above-described features and advantages, as well as other features and advantages, of the systems and methods of this disclosure will become apparent when considered in conjunction with the accompanying drawings and from the detailed description including the claims and exemplary embodiments. Attached Figure Description

[0032] This disclosure will be more fully understood from the detailed specifications and accompanying drawings, in which:

[0033] Figure 1 This is a block diagram depicting an embodiment of a transportation vehicle including a voice recognition system;

[0034] Figure 2 It describes communication with mobile devices and other transportation vehicles. Figure 1 A block diagram of an embodiment of a transportation vehicle;

[0035] Figure 3 This is a flowchart of a method for creating and training a speech recognition system; and

[0036] Figure 4 This is a flowchart of a speech recognition method according to an embodiment of the present disclosure. Detailed Implementation

[0037] Reference will now be made in detail to several examples of this disclosure shown in the accompanying drawings. Wherever possible, the same or similar reference numerals are used in the drawings and description to refer to the same or similar parts or steps.

[0038] Reference Figure 1The vehicle 10 typically includes a chassis 12, a body 14, and front and rear wheels 17, and may be referred to as a vehicle system. In the illustrated embodiment, the vehicle 10 includes two front wheels 17a and two rear wheels 17b. The body 14 is disposed on the chassis 12 and substantially surrounds the components of the vehicle 10. The body 14 and the chassis 12 may together form a frame. Each wheel 17 is rotatably coupled to the chassis 12 near a corresponding corner of the body 14. The vehicle 10 includes a front axle 19 coupled to the front wheels 17a and a rear axle 25 coupled to the rear wheels 17b.

[0039] In various embodiments, the vehicle 10 may be an autonomous vehicle, and a control system 98 is incorporated into the vehicle 10. The control system 98 may be referred to as a system or a voice recognition system. The vehicle 10, for example, is an automatically controlled vehicle capable of transporting passengers from one location to another. The vehicle 10 is depicted as a pickup truck in the illustrated embodiment, but it should be understood that other vehicles, including motorcycles, trucks, cars, sports cars, sports utility vehicles (SUVs), recreational vehicles (RVs), etc., may also be used. In one embodiment, the vehicle 10 is a so-called Level 4 or Level 5 automation system. A Level 4 system signifies “high automation,” referring to the driving mode-specific performance of the autonomous driving system in all aspects of a dynamic driving task, even if the human driver does not respond appropriately to intervention requests. A Level 5 system signifies “full automation,” referring to the full-time performance of the autonomous driving system in all aspects of a dynamic driving task under a variety of road and environmental conditions that can be managed by a human driver.

[0040] As shown in the figure, a vehicle 10 typically includes a propulsion system 20, a transmission system 22, a steering system 24, a braking system 26, a sensor system 28, an actuator system 30, at least one data storage device 32, at least one controller 34, and a communication system 36. In various embodiments, the propulsion system 20 may include an electric motor, such as a traction motor, and / or a fuel cell propulsion system. The vehicle 10 may also include a battery (or battery pack) 21 electrically connected to the propulsion system 20. Thus, the battery 21 is configured to store electrical energy and supply electrical energy to the propulsion system 20. In some embodiments, the propulsion system 20 may include an internal combustion engine. The transmission system 22 is configured to send power from the propulsion system 20 to the wheels 17 according to a selectable speed ratio. According to various embodiments, the transmission system 22 may include a step-ratio automatic transmission, a continuously variable transmission (CVT), or other suitable transmission. The braking system 26 is configured to provide braking torque to the wheels 17. In various embodiments, the braking system 26 may include a friction brake, a brake-by-wire brake, a regenerative braking system such as an electric motor, and / or other suitable braking systems. Steering system 24 affects the position of vehicle wheels 17 and may include steering wheels 33. Although depicted for illustrative purposes as including steering wheel 33, in some embodiments contemplated within the scope of this disclosure, steering system 24 may not include steering wheel 33.

[0041] Sensor system 28 includes one or more sensors 40 (i.e., sensing devices) that sense observable conditions of the external and / or internal environment of the vehicle 10. Sensors 40 communicate with controller 34 and may include, but are not limited to, one or more radars, one or more light detection and ranging (LiDAR) sensors, one or more odometers, one or more ground-penetrating radar (GPR) sensors, one or more steering angle sensors, one or more Global Positioning System (GPS) transceivers, one or more tire pressure sensors, one or more cameras 41 (e.g., optical cameras and / or infrared cameras), one or more gyroscopes, one or more accelerometers, one or more speed sensors, one or more steering angle sensors, one or more ultrasonic sensors, one or more inertial measurement units (IMUs), and / or other sensors. Each sensor 40 is configured to generate a signal indicating the sensed observable conditions of the external and / or internal environment of the vehicle 10. Because sensor system 28 provides data to controller 34, sensor system 28 and its sensors 40 are considered a source of information (or simply a source).

[0042] Sensor system 28 includes one or more Global Navigation Satellite System (GNSS) transceivers (e.g., Global Positioning System (GPS) transceivers) configured to detect and monitor route data (i.e., route information). The GNSS transceivers are configured to communicate with GNSS to locate the position of vehicle 10 globally. The GNSS transceivers communicate electronically with controller 34.

[0043] The actuator system 30 includes one or more actuator devices 42 that control one or more vehicle features, such as, but not limited to, a propulsion system 20, a transmission system 22, a steering system 24, and a braking system 26. In various embodiments, the vehicle features may further include internal and / or external vehicle features, such as, but not limited to, doors, trunk, and cabin features (e.g., air, music, lighting, etc.).

[0044] Data storage device 32 stores data for automatically controlling the vehicle 10. In various embodiments, data storage device 32 stores a defined map of the navigable environment. In various embodiments, the defined map may be predefined by and obtained from a remote system. For example, the defined map may be assembled by a remote system, communicate with the vehicle 10 (wirelessly and / or via a wired connection), and stored in data storage device 32. Data storage device 32 may be part of controller 34, separate from controller 34, or part of controller 34 and a separate system.

[0045] The vehicle 10 may also include one or more airbags 35 communicating with a controller 34 or another controller of the vehicle 10. The airbags 35 include inflatable airbags and are configured to switch between a loading configuration and a deployment configuration to cushion the effects of external forces applied to the vehicle 10. Sensors 40 may include airbag sensors, such as an IMU, configured to detect external forces and generate signals indicating the magnitude of those forces. The controller 34 is configured to command the airbags 35 to deploy based on signals from one or more sensors 40 (e.g., airbag sensors). Therefore, the controller 34 is configured to determine when the airbags 35 deploy.

[0046] The controller 34 includes at least one processor 44 and a non-transitory computer-readable storage device or medium 46. The processor 44 may be a custom or commercially available processor, a central processing unit (CPU), a graphics processing unit (GPU), an auxiliary processor among several processors associated with the controller 34, a semiconductor-based microprocessor (in the form of a microchip or chipset), a macroprocessor, a combination thereof, or a means generally used for executing instructions. The computer-readable storage device or medium 46 may include, for example, volatile and non-volatile memory in the form of read-only memory (ROM), random access memory (RAM), and power-on-retaining memory (KAM). KAM is a persistent or non-volatile memory that can be used to store various operational variables when the processor 44 is powered off. The computer-readable storage device or medium 46 may be implemented using multiple storage devices, such as PROM (programmable read-only memory), EPROM (electrical PROM), EEPROM (electrically erasable PROM), flash memory, or other electrical, magnetic, optical, or combined storage devices capable of storing data, some of which represent executable instructions used by the controller 34 to control the vehicle 10. The controller 34 of the vehicle 10 may be referred to as a vehicle controller and may be programmed to perform the speech recognition method 300 described in detail below. Figure 4 ).

[0047] The instructions may include one or more separate programs, each of which includes an ordered list of executable instructions for implementing logical functions. When executed by processor 44, the instructions receive and process signals from sensor system 28, execute logic, calculations, methods, and / or algorithms for automatically controlling components of traffic vehicle 10, and generate control signals to actuator system 30 based on the logic, calculations, methods, and / or algorithms to automatically control components of traffic vehicle 10. Although Figure 1 A single controller 34 is shown, but embodiments of the vehicle 10 may include multiple controllers 34 that communicate via a suitable communication medium or a combination of communication media and collaboratively process sensor signals, perform logic, calculations, methods and / or algorithms, and generate control signals to automatically control the features of the vehicle 10.

[0048] In various embodiments, one or more instructions of controller 34 are included in control system 98. Vehicle 10 includes user interface 23, which may be a touchscreen in a dashboard. User interface 23 may include, but is not limited to, alarms, such as one or more speakers 27 for providing audible sound, haptic feedback in vehicle seats or other objects, one or more displays 29, one or more microphones 31, and / or other means adapted to provide notifications to vehicle users of vehicle 10. Microphone 31 may be considered sensor 40 and configured to detect utterances made by the user. Specifically, microphone 31 is configured to convert audible sounds, such as utterances made by the user, into electrical signals. These electrical signals represent the user's utterances. In some embodiments, microphone 31 may be referred to as a first sensor, and camera 41 may be referred to as a second sensor. The second sensor (e.g., a camera) is configured to detect movements made by the user. User interface 23 communicates electronically with controller 34 and is configured to receive input from a user (e.g., a vehicle operator). For example, user interface 23 may include a touchscreen and / or buttons configured to receive input from a vehicle user. Therefore, controller 34 is configured to receive input from a user via user interface 23. User interface 23 includes one or more displays 29 configured to display information to a user (e.g., a vehicle operator or passenger), such as a head-up display (HUD), information cluster display, and / or infotainment center display.

[0049] Communication system 36 communicates with controller 34 and is configured to wirelessly transmit information to and from other entities 48, such as, but not limited to, other vehicles (“V2V” communication), infrastructure (“V2I” communication), remote systems at remote call centers (e.g., General Motors’ ON-STAR), and / or personal devices. In some embodiments, communication system 36 is a wireless communication system configured to communicate using the IEEE 802.11 standard or via a wireless local area network (WLAN) using cellular data communication. However, additional or alternative communication methods, such as Dedicated Short Range Communication (DSRC) channels, are also considered within the scope of this disclosure. A DSRC channel refers to a one-way or two-way short- to medium-range wireless communication channel designed specifically for automotive use, along with the corresponding set of protocols and standards. Therefore, communication system 36 may include one or more antennas and / or transceivers for receiving and / or transmitting signals such as Cooperative Sensing Messages (CSM). Communication system 36 is configured to wirelessly transmit information between vehicle 10 and another vehicle. In addition, the communication system 36 is configured to wirelessly transmit information between the vehicle 10 and infrastructure or other vehicles.

[0050] Reference Figure 2When the controller 34 as shown above is in the vehicle 10, the voice recognition method 300 ( Figure 4 The voice recognition method 300 can be executed by a controller 34 in the mobile device 50. The inputs and / or outputs of the voice recognition method 300 are then sent to one or more vehicles 10. As a non-limiting example, the mobile device 50 can be a mobile phone, tablet, or laptop. In some embodiments, the mobile device 50 includes a controller 34 and sensors 40 communicating with the controller 34. The controller 34 of the mobile device 50 can be referred to as a device controller. The sensors 40 of the mobile device 50 can be referred to as device sensors and include, for example, one or more microphones 31, one or more cameras 41 (e.g., optical cameras and / or infrared cameras), and one or more lidar sensors. The mobile device 50 can communicate with the fleet 10. Each vehicle 10 in the fleet includes the aforementioned components and can communicate with the mobile device 50 using a communication system 36. Thus, a user can use the voice recognition method 300 to provide verbal instructions or messages to all or part of the vehicles 100 in the fleet using one or more mobile devices 50.

[0051] Continue to refer to Figure 2 In some embodiments, the speech recognition method 300 ( Figure 4 The speech recognition process can be performed by a cloud-based system 52, with input collected by mobile devices 50 and / or vehicles 10 communicating with the cloud-based system 52. In this case, vehicles 10 and / or mobile devices 50 can collect auditory and non-auditory input and send these inputs to the cloud-based system 52. In the speech recognition method 300, the speech recognition system 98 can send user profiles to the fleet 10 via the cloud-based system 52. The user profile includes specific parameters of the speech recognition system for a particular user. Therefore, the user profile is stored in the cloud-based system 52. Once a user enters the vehicle 10, the cloud-based system 52 sends the user profile (i.e., the personalized speech recognition system 98) to the vehicle 10. Subsequently, the vehicle 10 uses the personalized speech recognition system 98 to process onboard data (i.e., data collected by the vehicle 10). Therefore, the vehicles in the fleet can personalize the speech recognition system 98.

[0052] Figure 3 It is used to create and train speech recognition methods 300 ( Figure 4The flowchart of method 200 is shown below. Method 200 begins at block 202. At block 202, one or more controllers 34 collect public data using sensors 40. The controllers 34 that create and train the speech recognition method 400 are not necessarily part of the vehicle 10 and / or the mobile device 50. As used herein, the term "public data" refers to a sufficient amount of audible data and user data from community users to develop a reliable speech recognition model. In this disclosure, the term "audible data" refers to data about utterances made by users, such as audio data. Thus, audible data includes information about utterances made by one or more users. Audible data may be referred to as audible input because it includes audible sounds from utterances made by users. Audible data may be collected by one or more microphones 31. Microphone 31 may be a transducer that converts audible sounds into electrical signals representing audible sounds. As mentioned above, public data also includes user data. In this disclosure, the term "user data" refers to data about movements made by users. Thus, user data includes information about movements made by users. As a non-limiting example, user data may include spatial data of various parts of the user's body (e.g., mouth, face, and / or hands) relative to a reference point. Thus, user data may include data about user actions, such as posture (e.g., gestures and / or facial movements), facial expressions, head movements, and mouth movements (e.g., lip movements). User data may, for example, include one or more images of the user's mouth and / or facial expressions. User data may be collected by one or more cameras 41 (e.g., optical cameras and / or infrared cameras) and / or lidar sensors. Optical cameras may be used for optical imaging, and infrared cameras may be used for imaging in low light or through sunglasses. In other words, optical cameras can capture images of user movement, and infrared cameras can capture thermal images of the user's face even when the user is wearing sunglasses. Lidar sensors may perform depth measurements. Lidar sensors can measure distances from the user's face or other parts of the user's body to a reference point and thus can detect user movement. Regardless of the specific sensor 40 used to acquire user data, these sensors 40 are used to detect head / face positioning, facial expressions, gesture mouth shapes, and / or gaze / eye tracking. Additionally, sensor 40 for collecting user data may include a pressure sensor for collecting pressure data and data from wearable devices such as smartwatches. Pressure data can be used to detect the user's posture and / or arm movements in the seat. Data from the wearable device can be used to monitor user movement and posture. Once public data is collected, method 200 proceeds to block 204.

[0053] At box 204, a speech recognition system (e.g., control system 98) is created and trained using public data. A neural network can be used to create and train the speech recognition system. As a non-limiting example, a recurrent neural network (RNN) or a transformer architecture can be used to create and train the speech recognition system. The RNN is used to analyze sequence-like data. As a non-limiting example, the RNN may include a long short-term memory (LSTM) and gated recurrent unit (GRU) architecture. At this point, the speech recognition system is general and not personalized for any particular user. Following box 204, method 200 continues to box 206.

[0054] In box 206, a voice recognition system (e.g., control system 98) is deployed. For this purpose, the voice recognition system can be deployed on cloud-based system 52, mobile device 50, and / or one or more vehicles 10. The voice recognition system can be configured as a voice assistant.

[0055] Figure 4 This is a flowchart of a speech recognition method 300. The speech recognition method 300 begins at block 302. At block 302, the controller 34 receives audible data from one or more sensors 40. As described above, the audible data includes information about utterances made by a user. The audible data may be referred to as audible input because it includes audible sounds from utterances made by the user. Furthermore, the audible data is collected by one or more sensors 40 (e.g., microphone 31) that may be located in the vehicle 10 and / or the mobile device 50. In some embodiments, if the sensor 40 (e.g., microphone 31) is located in the mobile device 50, the audible data may be transmitted to the vehicle 10 and / or the cloud-based system 52. The audible data may be shared with other vehicles 10 via the cloud-based system 52. If the sensor 40 is located in the vehicle 10, the audible data collected by the sensor 40 may be transmitted to the mobile device 50 and / or the cloud-based system 52. Therefore, the controller 34 receives the audible data. The method 300 also includes block 304.

[0056] At box 304, controller 34 (which may be in vehicle 10 and / or mobile phone 50) receives user data from one or more sensors 40. As described above, user data includes information about movements made by the user, such as facial expressions, mouth movements, and postures (e.g., gestures and / or facial movements). User data may be referred to as non-auditory input because it does not include audible sounds. User data is collected by one or more sensors 40 (e.g., one or more cameras 41 and / or lidar sensors). Collecting user data is desirable because such collection can facilitate speech recognition. For example, a user may say the word "yes" while nodding. In this case, sensor 40 detects that the user is nodding and associates this user movement with the word "yes." For this purpose, sensor 40 may detect (e.g., capture an image) the user's head movement (e.g., nodding). If sensor 40 (e.g., one or more cameras 41 and / or lidar sensors) is located in mobile device 50, user data may be transmitted to vehicle 10 and / or cloud-based system 52. User data can be shared with other vehicles 10 via cloud-based system 52. If sensor 40 is located in the vehicle 10, the user data collected by sensor 40 can be sent to mobile device 50 and / or cloud-based system 52. Therefore, controller 34 receives the user data. Blocks 302 and 304 can be executed simultaneously because some users coordinate movements such as gestures or lip movements with their speech. Therefore, by executing blocks 302 and 304 simultaneously, speech recognition method 300 can enhance its accuracy. After executing blocks 302 and 304, method 300 proceeds to block 306.

[0057] At block 306, controller 34 fuses audible data and user data to obtain fused data. Different techniques can be used for data fusion, such as data Bayesian networks, Dempster-Shafer theory, Bayesian filters, and / or neural networks. Regardless of the technique used, controller 34 integrates (i.e., fuses) audible data (i.e., auditory input) and user data (i.e., non-auditory input) to improve speech recognition. In this disclosure, the terms "fusion" or "integration" refer to linking individual utterances or specific audible data with individual motion or specific camera data. Fusing user data and audible data is desirable because, in many cases, users coordinate utterances with body movements. For example, a user might utter the word "no" while moving their head from side to side. In this case, sensor 40 detects that the user is moving their head from side to side, and controller 34 associates this head movement with the word "no" by fusing user data and audible data. After fusing audible data and user data, method 300 proceeds to block 308.

[0058] At block 308, controller 34 determines at least one spoken word based on user speech based on integrated user data and audible data (i.e., fused data). Furthermore, at block 308, controller 34 may command vehicle 10 to perform a specific task. For example, user speech may include a command such as “turn on the seat heaters.” In response to receiving this speech, controller 34 commands vehicle 10 to turn one or more actuator devices 42 to turn on the seat heaters. In another example, user speech may include a command such as “go home.” In this case, controller 34 commands the GNSS transceiver to retrieve the location of “home,” and the navigation system provides instructions to reach “home.” Following block 308, controller 34 proceeds to block 310.

[0059] At box 310, controller 34 fine-tunes the speech recognition system based on fused data to adapt to the speech patterns of a specific user. In this disclosure, "fine-tuning" or "adjustment" refers to training a neural network based on recognizing spoken words using actions associated with them and classifying those actions as meanings of the spoken words. As a non-limiting example, controller 34 fine-tunes the trained neural network to adapt to the speech patterns of a specific user. Therefore, fused data is stored relative to the specific user. For example, a user with a speech impairment may move their lips 55 in a specific manner when saying the word "navigate." While fine-tuning the trained neural network, when the specific user says the word "navigate," controller 34 recognizes that specific lip movement (i.e., the speech pattern) and thus associates that specific lip movement with the word "navigate" for that specific user. As a result, speech recognition method 300 allows the speech recognition system to be adapted to a specific user. After fine-tuning, speech recognition system 98 can send a user profile to a vehicle 10 or fleet 10 via cloud-based system 52. The user profile includes specific parameters of the speech recognition system for the specific user. Therefore, the user profile is stored in cloud-based system 52. Once a user enters vehicle 10, cloud-based system 52 sends a user profile (i.e., personalized voice recognition system 98) to vehicle 10. Vehicle 10 then uses personalized voice recognition system 98 to process onboard data (i.e., audible data and user data collected by vehicle 10). Therefore, the voice recognition system 98 can be personalized across a fleet of vehicles.

[0060] While exemplary embodiments have been described above, these embodiments are not intended to describe all possible forms covered by the claims. The words used in this specification are descriptive rather than restrictive, and it should be understood that various changes may be made without departing from the spirit and scope of this disclosure. As previously stated, features of various embodiments may be combined to form further embodiments of the systems and methods of this disclosure, which may not be explicitly described or shown. While various embodiments may be described as providing an advantage or superiority over other embodiments or prior art implementations with respect to one or more desired characteristics, those skilled in the art will recognize that one or more features or characteristics may be compromised to achieve desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, availability, weight, manufacturability, ease of assembly, etc. Therefore, embodiments described as less desirable than other embodiments or prior art implementations with respect to one or more characteristics do not fall outside the scope of this disclosure and may be desirable for a particular application.

[0061] The accompanying drawings are simplified and not to precise scale. For convenience and clarity only, directional terms such as top, bottom, left, right, upper, above, above, down, below, back, and front may be used in the drawings. These and similar directional terms should not be construed as limiting the scope of this disclosure in any way.

[0062] This document describes embodiments of the present disclosure. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The drawings are not necessarily drawn to scale; some features may be enlarged or minimized to show details of specific components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but only as a representative basis for teaching those skilled in the art to use the currently disclosed systems and methods in various ways. As will be understood by those skilled in the art, various features illustrated and described with reference to any of the drawings may be combined with features shown in one or more other drawings to produce embodiments not explicitly illustrated or described. The combinations of features shown provide representative embodiments for typical applications. However, for a particular application or implementation, various combinations and modifications of features conforming to the teachings of this disclosure may be required.

[0063] This document describes embodiments of the present disclosure based on functional and / or logical block components and various processing steps. It should be understood that such block components can be implemented by multiple hardware, software, and / or firmware components configured to perform specified functions. For example, embodiments of the present disclosure may employ various integrated circuit components, such as memory elements, digital signal processing elements, logic elements, lookup tables, etc., which can perform various functions under the control of one or more microprocessors or other control devices. Furthermore, those skilled in the art will understand that embodiments of the present disclosure can be practiced in conjunction with multiple systems, and the systems described herein are merely exemplary embodiments of the present disclosure.

[0064] For the sake of brevity, techniques related to signal processing, data fusion, signaling, control, and other functional aspects of the system (and its various operational components) are not described in detail herein. Furthermore, the connecting lines shown in the various figures included herein are intended to represent exemplary functional relationships and / or physical connections between various elements. It should be noted that alternative or additional functional relationships or physical connections may exist in the embodiments of this disclosure.

[0065] This specification is illustrative in nature and is in no way intended to limit the scope of this disclosure, its application, or its uses. The broad teachings of this disclosure can be implemented in various forms. Therefore, while this disclosure contains specific examples, its true scope should not be limited thereto, as other modifications will become apparent upon examination of the drawings, specification, and appended claims.

Claims

1. A speech recognition method applied to a transportation vehicle, comprising: Receive audible data, wherein the audible data includes information about speech produced by a user, including a person with a speech impairment; Receive user data, wherein the user data includes information about movements made by the user, the movement information includes data about user actions, and the action data includes user gestures or lip movements; The audible data and the user data are fused to obtain fused data; and At least one spoken word of the utterance is determined based on the fused data; Wherein, determining at least one spoken word of the utterance includes: The at least one spoken word of the utterance is determined using a trained neural network; The trained neural network is fine-tuned based on the fused data to adapt to the user's voice pattern; Fine-tuning the trained neural network includes: The user action is categorized into the meaning of the spoken word based on the identified words associated with the user action.

2. The method according to claim 1, wherein, The audible data is received via the microphone of the vehicle.

3. The method according to claim 1, wherein, The audible data is received via the microphone of the mobile device.

4. The method according to claim 1, wherein, The user data is received through the camera of the vehicle.

5. The method according to claim 1, wherein, The user data is received via the camera of the mobile device.

6. The method according to claim 1, wherein, The user data includes at least one image of the user's facial expression or at least one image of the user's lips.

7. The method according to claim 3, further comprising: The fine-tuned user profile is stored in a cloud-based system, wherein the user profile includes specific parameters of the personalized speech recognition system for the user. Once the user enters the vehicle, the user profile is sent to the vehicle; and The personalized voice recognition system is used to process the audible data and the user data.

8. The method according to claim 1, wherein, The method receives the user data and the audible data respectively through the camera and microphone of the mobile device, and further includes transmitting the user data and the audible data from the mobile device to the vehicle.

Citation Information

Patent Citations

  • Vehicle-mounted terminal equipment, vehicle-mounted interaction system and interaction method

    CN109941231A

  • Sensor enhanced speech recognition

    US20150364139A1