Voice recognition technology, particularly in a motor vehicle

The method enhances speech recognition in motor vehicles by using a voice sensor, a predetermined question list, and a transition matrix to improve accuracy and user experience in noisy environments, addressing the limitations of existing systems.

FR3164311A1Active Publication Date: 2026-01-09AMPERE SAS
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
FR2024007385
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2026-01-09
Estimated Expiration
2044-07-05

AI Technical Summary

Technical Problem

Existing speech recognition systems in motor vehicles are not adequately adapted to noisy environments and specific professional needs, leading to decreased accuracy and user experience, particularly for delivery drivers.

Method used

A method utilizing a voice sensor, a predetermined list of questions, and a transition matrix to enhance speech recognition by identifying probable states and keywords, incorporating driving and vehicle context information, and implementing a probabilistic model to optimize voice assistant responses.

Benefits of technology

Improves speech recognition accuracy and user experience by focusing on specific professional needs, reducing performance losses due to connectivity issues and enhancing the efficiency of voice interactions in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000027_0000
    Figure 00000027_0000
  • Figure 00000027_0001
    Figure 00000027_0001
  • Figure 00000027_0002
    Figure 00000027_0002
Patent Text Reader

Abstract

A speech recognition method, particularly in a motor vehicle. A speech recognition assistance method, particularly in a motor vehicle, adapted to a predetermined user objective and using: - at least one voice sensor located in said vehicle, in particular a microphone, connected to a voice assistant implementing at least natural speech understanding processing, - a predetermined list of a number N of questions Qi=1…N previously asked and learned according to the predetermined objective and / or the driving and vehicle cabin context, each question corresponding to a state, said list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of said system of states, the transitions having been previously learned. Figure for the abstract: Fig. 5
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Voice recognition method, particularly in a motor vehicle. Technical field

[0001] The present invention relates to a speech recognition method, particularly in a motor vehicle. Previous technique

[0002] Automotive voice recognition offers numerous practical and safety advantages to drivers. First, it allows drivers to control various vehicle functions, such as the air conditioning, audio system, and navigation, without taking their eyes off the road. This can help reduce distractions, limit accidents, and improve road safety. Furthermore, voice recognition can also facilitate communication between drivers and passengers or people outside the vehicle, for example, during hands-free phone calls or when dictating text messages. Finally, in a professional context, voice recognition could be useful for delivery drivers, who can read delivery instructions and interact with the delivery interface without having to take their eyes off the road.

[0003] It is known, particularly in the vehicles of the Applicant Company, to use voice recognition. This technology allows drivers to control various functions of their car, such as the audio system, with voice commands. The latter allows drivers to interact with their vehicle via voice commands to perform tasks such as selecting music and making hands-free phone calls. The use of a voice-activated personal assistant system provides drivers with a smoother and more intuitive experience.

[0004] This advanced voice recognition system is integrated into certain vehicles. This system allows the driver to control various car functions, such as navigation, music, and climate control, using simple and natural voice commands. It uses voice recognition technology to understand the driver's instructions and requests and provide clear and precise feedback. It is also compatible with connected vehicle controls, allowing drivers to control their favorite applications with voice commands, such as playing music and replying to messages. The voice personal assistant system is integrated into the system vehicle infotainment (IVI), offering an intuitive and easy-to-use interface for the driver.

[0005] Parcel delivery has also become an increasingly important activity, particularly in the context of online sales or e-commerce. Commercial vehicles, such as vans, are used to transport and deliver parcels. Shippers face complex tasks in organizing and executing their own deliveries efficiently. In this context, voice recognition has become a valuable tool for delivery drivers, as it allows them to control their vehicles and manage their work more effectively. As shown in [Fig. 1], connected commerce vehicles are now equipped with voice interfaces, allowing shippers to interact with their vehicles and manage functions related to delivery, navigation, mailing lists, and, where appropriate, communicate with customers using simple voice commands.This makes the delivery driver's job easier and faster, thus contributing to improved efficiency and productivity. As a result, delivery companies and carriers using commercial vehicles equipped with voice recognition can provide a faster and more reliable service to their customers.

[0006] By using voice commands, drivers can interact with their vehicle without taking their eyes off the road or using their work tablet, thus reducing distractions and improving road safety. Furthermore, voice recognition can also help drivers navigate unfamiliar or complex environments by providing clear and precise voice guidance that keeps them focused on the road. In short, voice recognition can significantly improve road safety and make driving a commercial vehicle less accident-prone by allowing the driver to concentrate on driving without being distracted by other tasks.

[0007] As shown in [Fig. 2], in step 1, the delivery person pronounces the keyword and then asks the system their question Q related to their delivery. In step 2, the keyword is detected by the process called "WuW" ("Wake up Word"). In step 3, the Automatic Speech Recognition (ASR) system reports the detected words agnostically. Natural Voice Understanding (NLU) reports the speaker's understood intent. Step 4 corresponds to the B2B speech recognition processing for the voice command based on question Q. The system queries the customer database and receives the associated response. In step 5, the system formats the response in bimodal mode via audio to communicate the response, using Text-to-Speech, and via the screen for navigation if necessary.

[0008] CN 115330290 discloses a distribution logistics management system based on intelligent voice-activated human-machine interaction. The system comprises several steps, including the configuration of basic information and the linking of a driver and a vehicle, the generation of packaging information, the generation of delivery order information, the generation of dispatch order information, the loading of goods by the driver, the transport of goods to the target delivery point, and the confirmation of delivery. The voice interaction module identifies and analyzes the voice via the mobile application, generates corresponding text, and performs semantic analysis to generate a voice instruction.

[0009] CN application 110059996 describes a vehicle loading system using a voice guidance system. The system includes several constraints and steps to optimize the loading of goods into a given loading space. The system uses a voice guidance system as the basis for managing the goods. It can give voice instructions to the operator on the order in which the goods should be loaded and the target positions within the loading space.

[0010] CN 108537480 describes a logistics management method based on speech recognition. A user ID corresponding to the goods order number is assigned, and input information for the user ID is obtained. If the input information is voice data, this data is identified to obtain an identification text corresponding to the voice data. A semantic analysis is then performed on the identification text, and if the semantic analysis matches cargo information, the cargo information is recorded in the cargo logistics information. Speech recognition makes it possible to identify users' voice data and convert it into text to perform semantic analysis. Thus, instructions, requests, and reports issued orally can be processed by the logistics management system.

[0011] US patent application 2021 / 0358496 describes a system for a motor vehicle that uses speech recognition. The system includes a multimedia device that recognizes the user's voice, converts speech to text, and analyzes the text to understand the requested action. Depending on the requested action, the system uses specific capabilities to perform that action, either locally in the vehicle or via a connection to a cloud service provider. The system then generates a voice response that is played through the vehicle's loudspeaker.

[0012] It should be noted that known speech recognition systems are not perfectly suited to professional needs. Indeed, it should be noted that this system is currently designed to be very general, which may not meet the specific needs of a particular application domain. In fact, general speech recognition is not very contextual and is designed to understand many different commands and queries in different contexts, which can be useful for more general uses. Furthermore, this system is mostly network-based, whereas having embedded speech recognition would be more advantageous.

[0013] Furthermore, noisy environments and regional accents can make communication difficult, which can negatively impact driver productivity, as accuracy is essential, and the quality of customer service. As shown in [Fig. 2], in step 3 of the ASR, the noisy environment of the commercial vehicle cab can significantly impact the ASR and NLU, resulting in decreased performance of the entire system and therefore a poor user experience.

[0014] In the automotive sector, noisy environments can include ambient noise such as engine noise, cabin noise, traffic noise, and conversations of people present, and potentially any other source of unwanted noise. Background noise can mask spoken words, making word recognition difficult. Speech recognition techniques use sophisticated algorithms to identify a voice and distinguish it from background noise. This requires precise analysis of various speech characteristics, such as frequency, amplitude, and duration, to determine the phonemes, syllables, and spoken words.

[0015] Another problem is speech comprehension in noisy environments. Ambient noise can alter the meaning of a sentence, making it difficult to understand the delivery person's request or instructions.

[0016] There is therefore a need to further improve speech recognition processes, particularly in motor vehicles, in order to offer speech recognition that is robust to the noise of the motor vehicle cabin and adapted to the context considered. Summary of the invention

[0017] The present invention addresses this need by means of, according to one of its aspects, a method for assisting speech recognition, particularly in a motor vehicle, adapted to a predetermined user objective and using:

[0018] - at least one voice sensor disposed in said vehicle, in particular a microphone, connected to a voice assistant that at least implements natural speech understanding processing,

[0019] - a predetermined list of a number N of questions Q^ia previously asked and learned according to the predetermined objective and / or the driving and cab context vehicle, each question corresponding to a state, said list being translated into a system of states, and

[0020] - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned,

[0021] the process comprising at least the following steps: a. After the voice sensor detects and recognizes a question posed to the voice assistant, search for the next question by choosing the two most probable states based on the transition matrix. b. identify a pair of keywords associated with the first and second states chosen in step a) from among predefined keywords for each question in the predetermined list, c. When a new question is detected by the voice sensor and recognized by the voice assistant, search for the presence in said question of the keywords of the pairs identified in step b) and:

[0022] - if no keyword is detected, recognition by the voice assistant continues by natural voice comprehension processing without additional information, or

[0023] - if the pair of keywords associated with the first state chosen in step a) is detected, the Natural voice comprehension processing continues with the indication of the question associated with this first state, or

[0024] - if the pair of keywords associated with the second state chosen in step a) is detected, The natural voice comprehension process continues with the indication of the question associated with this second state.

[0025] d) continue the process with the selected question in order to facilitate future speech recognitions.

[0026] The invention improves the performance of voice recognition in the vehicle cabin. The invention offers a professional-focused solution, meaning it addresses more specific and precise voice recognition needs. The invention is tailored to these specific requirements to ensure maximum accuracy and an optimal user experience.

[0027] In addition, the invention is placed in an embedded model, which makes it possible to avoid performance losses due to lack of connectivity when speech recognition is on the network.

[0028] The method according to the invention makes it possible to fully exploit the processing, interpretation, and speech recognition capabilities of the voice assistant. Sequencing logic is taken into account, for example, if the user is a delivery driver, questions about the first delivery will be asked before those concerning the last.

[0029] The invention implements a probabilistic model that focuses more specifically on intention domains related to the user's objective, aiming to optimize the efficiency and accuracy of the voice assistant. This means that the voice assistant can be adapted to respond more accurately and quickly to the user's questions. For example, if the user is a delivery driver, the voice assistant can efficiently answer questions related to the delivery route, thus improving the overall efficiency of the delivery process.

[0030] The invention assists natural voice understanding processing by providing it with prior information about what it is supposed to interpret. Thus, natural voice understanding processing is no longer alone in this interpretation process and therefore solely responsible for the correct or incorrect interpretation of the user's intention.

[0031] In a preferred embodiment, vehicle driving and cabin context information is used, including information about the user, the predetermined objective, the environment, previous actions, or geographical information. This allows for the estimation of a prediction of the user's questions based on the context. State system and transition matrix

[0032] A transition matrix is ​​a known concept used in graph theory and probability to represent the change of state in a system.

[0033] In the invention, the transition matrix represents the possible transitions between each state of said state system, learned during past occurrences, according to the predetermined objective, and updated after each occurrence, for example after each delivery in the case where the user is a delivery person.

[0034] The transition matrix is ​​advantageously a square matrix of size N x N, where each element of the matrix represents the probability of transition from one state to another, i.e. from one question to another.

[0035] In a preferred embodiment, the transition matrix T contains an element in the i-th row and j-th column, denoted T[i, j] or q^, which represents the probability of moving from question Q; to question Qj. The elements of the transition matrix T advantageously satisfy the known properties of a transition matrix, namely that each element of the matrix must be between 0 and 1, since it represents a probability. The sum of the elements in each row of the matrix must be equal to 1 or 100%.

[0036] [Math.l]

[0037] This matrix advantageously models the sequencing of questions asked during the different stages of the predetermined objective, particularly during a delivery route. For the mth stage, the matrix can be presented as follows:

[0038] [Math.2]

[0039] For the m+l-th step, the matrix can be presented as follows:

[0040] [Math.3] / Q1 (wire + 1) y Ci. i (îîîj \ / Q1 (m)

[0041] Maximum likelihood can be used to candidate the next two most probable states, represented by the two highest probabilities.

[0042] As is well known, the maximum likelihood principle consists of estimating the parameters of a probabilistic model by finding the value that makes the observed data most probable. This allows for statistical decisions to be made based on real data and for obtaining accurate estimates of unknown parameters.

[0043] At step m, for question Qk (m) for k :1..N, the following two most probable questions can be chosen:

[0044] [Math.4] = Ckr MI # j)

[0045] When a new question is detected by the voice sensor and recognized by the voice assistant, the presence of keywords is searched in said question. pairs identified in step b). This will allow the selection of the most probable question from the two questions retained.

[0046] For the question Qi for i: 1..N:

[0047] WuW(i,l): wake word #1

[0048] WuW(i,2) : wake word #2.

[0049] To achieve this, a 'wake-up word' algorithm is used. This type of algorithm is more robust at detecting words without necessarily recognizing them, as it is not a natural speech comprehension process. This principle is used in speech recognition systems to activate the system based on a specific short speech keyword, before triggering automatic speech recognition and natural speech comprehension processing. This mechanism allows devices to listen continuously but only react when the wake-up word is detected.

[0050] Known voice assistants use a wake word. These assistants are in standby mode and do not actively listen to all conversations. However, as soon as the specific wake word is spoken, the assistant wakes up and is ready to respond to commands.

[0051] This principle relies on signal feature recognition technology (mel-frequency cepstral coefficients, MFCCs, or neural networks). When the device is in listening mode, it continuously records ambient audio, but without actively processing it. However, it performs a rapid and continuous analysis of the audio signal to detect whether the wake word is spoken. Wake word detection is generally performed using machine learning models, such as neural networks, which have been trained on large amounts of speech data to specifically recognize the wake word.

[0052] The detection of the wake word must be precise enough to prevent accidental activations. These systems are therefore qualified for their robustness and are continuously improved to reduce false triggers and react only to genuine wake words.

[0053] The keyword pairs are preferably distinct in pairs so that no overlap and therefore no double detection is possible.

[0054] This "wake-up word" detection principle thus makes it possible to identify the most probable question based on signal characteristics. The wake-up words present in the candidate question indicate the likelihood of that question, which can be communicated to the natural speech understanding processing. Depending on their detection, the task of this processing will then be facilitated.

[0055] When the new question is finally determined, it is advantageously transmitted to a processing module associated with the predetermined objective.

[0056] During step d), the transition matrix can be updated with the selected question.

[0057] The updating of the coefficients of the transition matrix is ​​carried out by the method based on the data related to the occurrences of the questions during an occurrence according to the predetermined objective.

[0058] The coefficients can be updated according to the following equation:

[0059] [Math.5] Ht Jl - qÿ - Ml D /

[0060] with n(i, j) the number of occurrences of the transition from state i to state j observed in the empirical data. The total number of occurrences of all transitions starting from state i is given by n_total(i).

[0061] This means that the probability of going from state i to state j is equal to the number of occurrences of this transition divided by the total number of occurrences of all transitions starting from state i.

[0062] The updating of the coefficients of the transition matrix between time t (question asked at t equivalent to step m) and time t+dt (question which will be asked at t+dt equivalent to step m+1) is preferably based on empirical observations and questions of transitions between the states of the system.

[0063] The updating of the coefficients of the transition matrix between time t and time t + 1 can be done according to the following equation:

[0064] [Math.6] T(t+dt)[i, j] - qij = n(t, i, j) / n__total(t, i)

[0065] with T(t) the matrix of transitions at time t, T(t)[i, j] which represents the transition coefficient from state i to state j at time t. n(t, i, j) the number of occurrences of the transition from state i to state j observed between time t and time t+dt, n_total(t, i) the total number of occurrences of all transitions starting from state i between time t and time t+dt, that is to say the sum of n(t, i, j) for all states j.

[0066] Taking the step into account, the expression can be expressed as follows:

[0067] [Math.7] T(m+l)[i, j] = qij = n(m, i, j) / n_total(m, i)

[0068] This means that the probability of going from state i to state j between time t and time t+dt is equal to the number of occurrences of this observed transition divided by the total number of occurrences of all transitions starting from state i in the period between time t and time t+dt.

[0069] Preferably, the general equation for updating row i of the transition matrix is ​​as follows:

[0070] [Math. 8] T[i, j] = PG|i) = Qÿ

[0071] with T[i, j] the element of the transition matrix located in row i and column j.

[0072] Since this also represents the transition probability from state i to state j, P(jli) is the conditional probability of moving from state i to state j. Contextual information

[0073] The method according to the invention advantageously uses information from the driving context and vehicle cab, including information about the user, the predetermined objective, the environment, previous actions, or geographical or road information.

[0074] Context information advantageously comes from time series of data from a plurality of sensors included in the vehicle, said time series being transformed into probability vectors and grouped into data groups according to predetermined training parameters.

[0075] The numerous contextual information or "Car Data" reflects the journey context or "Driving Task", as well as the vehicle cab context or "In Vehicle Task".

[0076] This information evolves over time. It can be optimally collected through digital representations automatically learned from raw data, enabling machine learning models to better understand and utilize the information contained within that data. These representations, also called "AI embeddings" or simply "embeddings," of extracted features, embeddings, or compact semantic vectors, play a crucial role in many areas of artificial intelligence by facilitating the processing and analysis of complex data.

[0077] In the field of intelligent automobiles, these representations are useful in analyzing a vehicle's driving information. When a vehicle is moving, it generates a large amount of data from integrated sensors, such as GPS sensors, accelerometers, speed sensors, cameras, ADAS (Advanced Driver Assistance Systems), etc. This data is critical for understanding and improving vehicle behavior and performance, safety, maintenance, and ADAS systems.

[0078] These time series of data correspond to the numerical descriptions of the state of the driver / vehicle control system at a given moment, which include all the information relevant to understanding the instantaneous state of the journey, such as such as position, speed, acceleration, altitude, etc. By creating time series of contextual data for each moment of the journey, a detailed view of what happens throughout the journey is obtained.

[0079] These time series of data are derived from different sensors, each emitting its own data stream providing different information, for example: - Driver Monitoring System (DMS) camera: all DMS reports, such as alertness, nervousness, and fatigue, - Visual attention / road / landscape / cabin: measurement of the field of vision, - Pedal pressure: pressure and its frequency on the pedals indicating the level of attention to pedal actions. - Steering wheel grip and contact: a tight or loose grip on the steering wheel, with one or two hands, indicates the driver's attention or lack of focus. - HMI and radio manipulation, air conditioning...: the HMI system and its operating states at the time of measurement, - Blinking: camera measurements allow the periodicity or aperiodicity of the driver's blinks to be captured. - Driving time: the duration of driving. - Risk-taking: a driver in a hurry may adopt a driving style that tends to disregard the highway code and the speed limit.

[0080] Data related to the road context can also be recorded, such as, for example: - Functional Road Class #0...8: road type identification, - Weather conditions: the weather situation, - Winding / Straight line: the evolution of the road's shape, - Slope / Ascent: the evolution of the road's shape, - Urban context (city, density), Peri-urban, Countryside, Mountain, - Anticipating road events: navigation is a critical indicator for anticipating the cognitive load required for upcoming road situations in the near future. - Smooth traffic flow, Night / Day, Works, Stop, Hour, Day.

[0081] Data related to the vehicle context can also be recorded, such as, for example: - Speed, Acceleration (deceleration): long, lat, - Exceeding the speed limit, - Front-rear safety index / sensors: AD AS alarms, steering angle, angular velocity, - AD / AS controllers: the status of AD / AS controllers is directly related to measurement data, such as: • AEBS: Automatic Emergency Braking System • ACC: Adaptive Cruse Control • FCW: Forward Collision Warning • EAPM: Emergency Assist for Pedal Misapplication • ESP: Electronic Stability Program • DW: Distance Warning • OSP: Over Speed Prévention • ISA: Intelligent Speed Assistance • LDW: Lane Departure Waming / Emergency • LKA: Lane Keeping Assist • LDP: Lane Departure Prévention • LSS: Advanced Latéral Support • UTA: Unstable Trajectory Alert • DAA: Driver Attention Alert • BSW: Blind Spot Warning • BSI: Blind Spot Intervention • CTA: Cross Traffic Alert • BCI: Back up Collision Intervention • R-CTA: Rear Cross Traffic Alert • CASP : Contextual Adaptive Speed Prévention.

[0082] Raw data from various sensors can be complex and difficult to process directly. Machine-learned numerical representations allow this raw data to be converted into more understandable numerical representations, while preserving essential information. In this context, time series of data are advantageously transformed into probability vectors. These are designed to compactly represent the key characteristics of the vehicle's various driving lists.

[0083] Using these digital representations, several useful analyses can be carried out, such as, among others, anomaly detection, machine learning, and driving optimization.

[0084] As is known, "AI embeddings" in the context of vehicle driving data use the concept of an encoder / decoder based on neural networks to transform raw data into meaningful representations. As shown in [Fig. 3], the encoder takes driving data as input, such as speed, acceleration, steering angles, etc., and converts it into compact numerical representations that capture the key characteristics of the driving lists. The decoder, in turn, uses these numerical representations as initial context to generate useful information, such as anomaly detection and prediction. driving events or driving optimization. "AI embeddings" therefore allow us to focus on essential and relevant data within the very large set of vehicle data.

[0085] AI embeddings are, as is known, grouped into data sets or "clusters," as shown in [Fig. 4]. Clustering numerical representations from vehicle driving data is a powerful method for grouping similar driving lists into clusters. This offers several significant advantages in the analysis of vehicle driving data, such as:

[0086] - Detection of driving patterns: using clustering algorithms on With "embeddings," similar driving patterns can be automatically identified across different recorded lists. This can reveal characteristic driving styles, recurring behaviors, or even specific scenarios such as parking maneuvers, sudden acceleration, etc. This information can be invaluable for understanding driver behavior and improving road safety.

[0087] - Segmentation of the process: clustering the "embeddings" allows the division of Driving data broken down into segments or specific journeys. This can be useful for analyzing each journey individually, identifying high-risk driving areas, frequent routes, or particular traffic conditions.

[0088] - Prediction of driving events: the clusters obtained from the " Embeddings can be used as input for machine learning models. By training a model on clustered driving data, it becomes possible to predict certain events, such as dangerous situations, sudden changes in driver behavior, or even future traffic conditions.

[0089] A non-exhaustive list of road context information may be: - Accident on the main road, anticipated question: "Is there any information about the accident and any suggestions to avoid further delays?" - Road closed due to construction work, anticipated question: "Is there a recommended detour route to avoid the work zone?" - Heavy traffic jam, anticipated question: "Is there a faster alternative to avoid this traffic jam?" - Bad weather conditions (heavy rain), anticipated question: "Are there safer roads to drive on in the rain?" - Major roadworks, anticipated question: "Is there an alternative route due to the works?" - Temporary closure of the main road, anticipated question: "What is the best way to bypass the road closure?" - Traffic jam during the morning rush hour, anticipated question: "Are there any less congested secondary roads at this time?" - Prolonged traffic jam, anticipated question: "What is the estimated delay for my delivery due to the current traffic jam?" - Road closed due to a sporting event, anticipated question: "Is there an alternative route to bypass the road closure due to the sporting event?" - Level crossing blocked by a train, anticipated question: "Is there an alternative route to avoid the level crossing blocked by the train?" - Congestion due to a demonstration, anticipated question: "Can you direct me to a less congested route due to the ongoing demonstration?" - Roadworks with diversion, anticipated question: "What is the best way to reach the destination using the current roadworks diversion?" - Winter weather conditions (snow), anticipated question: "Are there clear and safe roads for driving during periods of snow?" - Strong gusts of wind, anticipated question: "Are there roads less exposed to strong winds to ensure safer driving?" - Construction work on a motorway, anticipated question: "How can I avoid delays due to construction work on the motorway?" - Traffic lights out of service, anticipated question: "Are there any special instructions for safely crossing intersections without traffic lights?" - Traffic jam caused by an accident, anticipated question: "Can you provide me with information on the duration of the traffic jam and the alternative routes available?" - Temporary closure of a lane due to maintenance work, anticipated question: "Are there any detours in place to bypass the lane closure due to maintenance work?" - Traffic slowdown caused by road repair work, anticipated question: "What is the expected impact on travel time due to the ongoing road repair work?"

[0090] A non-exhaustive list of vehicle and driver context information may be : - Excessive fatigue, anticipated question: "Are there any recommended rest areas nearby where I can take a break and rest?" - Low fuel level, anticipated question: "Where can I find the nearest gas station to refuel?" - Need to take a break for food, anticipated question: "What restaurants or fast food establishments are available on my route?" - Need to find a toilet, anticipated question: "Are there any accessible rest areas or public toilets near my current location?" - Need to find secure parking, anticipated question: "Can you tell me about available parking lots nearby where I can safely park my delivery vehicle?" - Need to find an electric charging point, anticipated question: "Where can I locate the nearest electric vehicle charging stations to charge my delivery vehicle?" - Need mechanical assistance, anticipated question: "Can you recommend a reliable garage or car breakdown service nearby?" - The delivery driver wants to know if the vehicle data indicates a problem with the brakes, thus ensuring safety during stops and slowdowns; anticipated question: "Is there a problem with the brakes?" - This question helps to check if the tires are correctly inflated, as inadequate pressure can affect handling and fuel consumption. Anticipated question: "What is the tire pressure level?" - The delivery driver wants to know the average fuel consumption of his vehicle, which will help him to assess costs and optimize his routes for better energy efficiency. Anticipated question: "What is the average fuel consumption?"

[0091] A subset of probable questions can be associated with each data group in order to improve natural voice comprehension processing.

[0092] The frequencies of the N questions can be counted for each of the M data groups. At the nth occurrence n_i of group i, question k, issued among the N questions, can be counted. The probability of question k in data group i advantageously corresponds to the number of times this question was issued in the number of occurrences of the context of data group i, expressed by the following equation, for all possible questions k:1..N:

[0093] [Math.9] P(Qk Clusteri) = nb(Qk) / nb(n_i)

[0094] Each data group is advantageously associated with a probability vector of the user's questions about the context.

[0095] The invention may relate to a method for assisting speech recognition, particularly in a motor vehicle, adapted to a predetermined user objective and using:

[0096] - at least one voice sensor disposed in said vehicle, in particular a microphone, connected to a voice assistant that at least implements natural speech understanding processing,

[0097] - a predetermined list of a number N of questions Q i...N previously asked and learned according to the driving and vehicle cab context, each question corresponding to a state, said list being translated into a system of states, and

[0098] - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned,

[0099] the process comprising at least the following steps: - after detection by the voice sensor and recognition of a question posed to the voice assistant, time series of data from a plurality of sensors included in the vehicle are acquired, said time series providing information on the driving context and vehicle cabin, - said time series are transformed into probability vectors and grouped into data sets according to predetermined training parameters, - The following question is investigated by choosing the two most probable states based on the transition matrix, - identify a pair of keywords associated with the first and second states chosen in the previous step from predefined keywords for each question in the predetermined list, - when a new question is detected by the voice sensor and recognized by the voice assistant, search in said question for the presence of the keywords of the pairs identified in step b) and:

[0100] - if no keyword is detected, recognition by the voice assistant continues by natural voice comprehension processing without additional information, or

[0101] - if the pair of keywords associated with the first state chosen in step a) is detected, the Natural voice comprehension processing continues with the indication of the question associated with this first state, or

[0102] - if the pair of keywords associated with the second state chosen in step a) is detected, The natural voice comprehension process continues with the indication of the question associated with this second state.

[0103] - continue the procedure with the selected question in order to facilitate future voice recognition.

[0104] The probability vector is preferably updated for the group concerned.

[0105] In one embodiment, questions related to the predetermined objective and Questions related to the driving environment and vehicle cab are used. The final choice is thus optimized thanks to the suggestions from these two aspects.

[0106] The selection of the two most probable states is advantageously carried out using the probabilities of the transition matrices of the two state systems. The process then continues with the identification of the keyword pair. The transition matrix is ​​updated with the selected question, as well as the probability vector. computer program product

[0107] The invention also relates, in another of its aspects, to a computer program product comprising a medium and, recorded on this medium, instructions readable by a processor so that, when executed, they enable the implementation of the speech recognition assistance method according to the invention.

[0108] The characteristics stated in relation to the process apply to the computer program product and vice versa. Device

[0109] According to another aspect of the invention, the invention relates to a speech recognition aid device, particularly in a motor vehicle, adapted to a predetermined user objective and comprising at least one voice sensor disposed in said vehicle, in particular a microphone, connected to a voice assistant implementing at least natural speech understanding processing and comprising or being connected to at least one module, said module having access to: - a predetermined list of a number N of questions Q^ia previously asked and learned according to the predetermined objective and / or the driving and vehicle cab context, each question corresponding to a state, said list being translated into a system of states, and

[0110] - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned,

[0111] said module being configured to perform at least the following steps: a. After the voice sensor detects and recognizes a question posed to the voice assistant, search for the next question by choosing the two most probable states based on the transition matrix. b. identify a pair of keywords associated with the first and second states chosen in step a) from among predefined keywords for each question in the predetermined list, c. When a new question is detected by the voice sensor and recognized by the voice assistant, search for the presence in said question of the keywords of the pairs identified in step b) and:

[0112] - if no keyword is detected, recognition by the voice assistant continues by natural voice comprehension processing without additional information, Or

[0113] - if the pair of keywords associated with the first state chosen in step a) is detected, the Natural voice comprehension processing continues with the indication of the question associated with this first state, or

[0114] - if the pair of keywords associated with the second state chosen in step a) is detected, The natural voice comprehension process continues with the indication of the question associated with this second state.

[0115] d) continue the process with the selected question in order to facilitate future speech recognitions.

[0116] The characteristics stated in relation to the method apply to the device and vice versa. Motor vehicle

[0117] According to another aspect of it, the invention relates to a motor vehicle comprising a powertrain and at least one device according to the invention.

[0118] The characteristics stated in relation to the method apply to the vehicle and vice versa. Brief description of the drawings

[0119] The invention will be better understood upon reading the detailed description that follows, a non-limiting example of its implementation, and upon examination of the accompanying drawing, in which

[0120] [Fig-1] [Fig.1], already described, illustrates an example of the use of recognition voice command in a motor vehicle to assist a delivery person

[0121] [Fig.2] The [Fig.2], already described, illustrates different steps implemented during a delivery route

[0122] [Fig.3] The [Fig.3], already described, illustrates the mechanism of "AI embeddings" in the context of vehicle driving data,

[0123] [Fig.4] Fig.4, already described, illustrates the grouping of data in the form of data groups,

[0124] [Fig.5] [Fig.5] is a logic diagram illustrating the steps of an example of implementation in the implementation of the process according to the invention,

[0125] [Fig.6] [Fig.6] is a logic diagram illustrating the steps of another example of implementing the process according to the invention, and

[0126] [Fig.7] [Fig.7] is a logic diagram illustrating the steps of another example implementation of the process according to the invention. Detailed description

[0127] Figure 5 illustrates an example of an implementation of the speech recognition assistance method according to the invention. In the illustrated example, the method, adapted to a predetermined objective of a delivery person, takes place in a motor vehicle and uses:

[0128] - the voice sensor located in said vehicle and connected to a voice assistant putting at less so in implementing a natural voice comprehension process,

[0129] - a predetermined list of a number N of questions Q.-iy previously asked and learned according to the predetermined objective, each question corresponding to a state, said list being translated into a system of states, and

[0130] - a transition matrix representing the possible transitions between each state said system of states, the transitions having been previously learned.

[0131] In the case of a delivery route, the questions that the driver asks the voice assistant are already predefined in a specific area. They refer to route statuses such as: route summary, next stop, remaining tasks, time remaining, urgent delivery request, cancellation, call / text to fleet manager / boss, customer call, and others.

[0132] Here is a non-exhaustive list of questions: 1. What is my next step? 2. Where do I need to go? 3. And after this delivery? 4. Where should I deliver the next package? 5. What is the address of the next delivery? 6. Where will I drop off my next package? 7. How many packages do I have left? 8. What remains to come? 9. How many are left? 10. How many addresses are left? 11. Are there any urgent deliveries? 12. Am I late somewhere? 13. Urgent packages? 14. Do we have any priority mail today? 15. Do you know what time we have to leave today? 16. How long will it take? 17. What is the expected time of our final delivery? 18. Can I cancel this job? 19. Can you repeat what you were saying? 20. Sorry, could you repeat that?

[0133] Each question is posed in a very specific list that perfectly reflects the sequence of actions carried out during the delivery route. This pre-established list is the result of meticulous planning aimed at optimizing the efficiency of the delivery route, taking into account unforeseen events.

[0134] For example, when the delivery driver asks question #3, it means he is at a certain stage of his route, which is defined as a state in the process according to the invention. From this state, corresponding to a specific associated stage, he will not ask just any subsequent question, but the one most likely to be asked in relation to the next stage of his route.

[0135] After steps 0 and 1 of detection by the vehicle's voice sensor and recognition of a question asked to the voice assistant, the next question is sought by choosing the two most probable states according to the transition matrix, in steps 2 and 3.

[0136] In step 4, a pair of keywords associated with the first and second states chosen in the previous step is identified from among predefined keywords for each question in the predetermined list. For example, for the question "What is the address of the next delivery?", the word pair could be "address" and "delivery". For the question "Do we have priority mail today?", the word pair could be "mail" and "priority".

[0137] At a step 5, when a new question is detected by the voice sensor and recognized by the voice assistant, a search is conducted in a step 6 for the presence in said question of the keywords of the pairs identified in step 4.

[0138] At step 7, if no keyword is detected, the voice assistant continues with natural speech understanding processing without additional information. Conversely, if the keyword pair associated with the first state chosen in step 3 is detected, natural speech understanding processing continues with the indication of the question associated with this first state. Finally, if the keyword pair associated with the second state chosen in step 3 is detected, the processing of Understanding of the natural voice continues with the indication of the question associated with this second state.

[0139] The process is then continued in step 8 with the selected question to facilitate future speech recognition. The transition matrix is ​​updated with the selected question in step 9.

[0140] In the variant of [Fig. 2], after steps 0 and 1 of detection by the vehicle's voice sensor and recognition of a question posed to the voice assistant, driving and vehicle cabin context information is acquired during a step 2, including information about the user, the predetermined objective, the environment, previous actions, or geographical information, as defined previously. This context information comes from time series data from a plurality of sensors included in the vehicle.

[0141] These time series are transformed into probability vectors and grouped into data groups according to predetermined training parameters, in a step 4. A subset of probable questions is associated with each data group in order to improve natural voice comprehension processing, in a step 5.

[0142] In a step 6, after the voice sensor detects and recognizes a question asked to the voice assistant, the next question is sought by choosing the two most probable states according to the transition matrix, in a step 7.

[0143] In step 8, a pair of keywords associated with the first and second states chosen in the previous step is identified from among predefined keywords for each question in the predetermined list. When a new question is detected by the voice sensor and recognized by the voice assistant in step 9, step 10 searches the question for the presence of the keyword pairs identified in step 8.

[0144] In step 11, if no keyword is detected, the voice assistant continues its recognition process by processing natural speech without additional information. Conversely, if the keyword pair associated with the first state selected in step a) is detected, the natural speech understanding process continues with the prompting of the question associated with that first state. Finally, if the keyword pair associated with the second state selected in step a) is detected, the natural speech understanding process continues with the prompting of the question associated with that second state.

[0145] The process continues in step 12 with the selected question to facilitate future speech recognition. The probability vector is updated for the relevant group.

[0146] In the variant of [Fig.7], at step 2, the two treatments of figures 5 and 6 are performed. Questions related to the predetermined objective and questions related to the driving and vehicle cab context are used in step 3. The selection of the two most probable states is carried out based on the probabilities of the transition matrices of the two state systems. The process then continues with the identification of the keyword pair. The transition matrix is ​​updated with the selected question, as well as the probability vector.

[0147] The invention is not limited to the examples just described.

[0148] The invention can be used by a voice assistant in a passenger vehicle identifying target areas and using the principle of vehicle cab and driving context information.

[0149] The invention can be implemented in any field using speech recognition specific to a given field of activity, such as aviation for example.

Claims

Demands

1. A method for assisting with speech recognition, particularly in a motor vehicle, adapted to a predetermined user objective and using: - at least one voice sensor installed in said vehicle, including a microphone, connected to a voice assistant implementing at least natural voice understanding processing, - a predetermined list of N Qi-r.x questions previously asked and learned according to the predetermined objective and / or the driving and vehicle cab context, each question corresponding to a state, said list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of said system of states, the transitions having been previously learned, the process comprising at least the following steps: a. After the voice sensor detects and recognizes a question posed to the voice assistant, search for the next question by choosing the two most probable states based on the transition matrix. b. identify a pair of keywords associated with the first and second states chosen in step a) from among predefined keywords for each question in the predetermined list, c. When a new question is detected by the voice sensor and recognized by the voice assistant, search for the presence in said question of the keywords of the pairs identified in step b) and: - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or - if the pair of keywords associated with the first state chosen in step a) is detected, the natural voice understanding processing continues with the indication of the question associated with this first state, or - if the pair of keywords associated with the second state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this second state, a. continue the process with the selected question in order to facilitate future speech recognitions.

2. A method according to claim 1, wherein, at step d), the transition matrix is ​​updated with the selected question.

3. A method according to claim 1 or 2, using vehicle cab and driving context information, including information about the user, the predetermined objective, the environment, previous actions, or geographical information.

4. A method according to the preceding claim, wherein the context information is derived from time series of data from a plurality of sensors included in the vehicle, said time series being transformed into probability vectors and grouped into data groups according to predetermined training parameters.

5. A method according to the preceding claim, wherein a subset of probable questions is associated with each data group in order to improve natural voice understanding processing.

6. A method according to any one of claims 4 or 5, wherein the frequencies of the N questions are counted for each of the M data groups, at the nth occurrence n_i of group i, question k is counted among the N questions, the probability of question k in data group i corresponding to the number of times that this question was asked in the number of occurrences of the context of data group i, expressed by the following equation, for all possible questions k:1..N: [Math.9] P(Qk | Cluster i) = nb(Qk) / nb(ni)

7. A method according to any one of claims 4 and 5, wherein the probability vector is updated for the group concerned.

8. A method according to any one of the preceding claims, wherein the transition matrix is ​​a square matrix of size N x N, where each element of the matrix represents the probability of transition from one state to another.

9. A method according to any one of the preceding claims, wherein maximum likelihood is used to candidate the next two most probable states, represented by the two highest probabilities.

10. A speech recognition assistance device, particularly in a motor vehicle, adapted to a predetermined user objective and comprising at least one voice sensor disposed in said vehicle, including a microphone, connected to a voice assistant implementing at least natural speech understanding processing and comprising or being connected to at least one module, said module having access to: - a predetermined list of a number N of questions Q previously asked and learned according to the predetermined objective and / or the driving and vehicle cabin context, each question corresponding to a state, said list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of said system of states, the transitions having been previously learned, said module being configured to perform at least the following steps: a.after the voice sensor detects and recognizes a question asked to the voice assistant, search for the next question by choosing the two most probable states based on the transition matrix, b. identify a pair of keywords associated with the first and second states chosen in step a) from keywords predefined beforehand for each question in the predetermined list, c. when a new question is detected by the voice sensor and recognized by the voice assistant, search in. said question the presence of the keywords of the pairs identified in step b) and: - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or - if the pair of keywords associated with the first state chosen in step a) is detected, the natural voice understanding processing continues with the indication of the question associated with this first state, or - if the pair of keywords associated with the second state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this second state, d) continue the process with the selected question in order to facilitate future speech recognitions.

11. Product computer program comprising a medium and, stored on this medium, instructions readable by a processor so that, when executed, they enable the implementation of the speech recognition assistance method according to any one of claims 1 to 9.

12. Motor vehicle comprising a powertrain and at least one identification device according to claim 10.

Citation Information

Patent Citations

  • Logistics management method and system based on speech recognition

    CN108537480A

  • Field area vehicle loading method based on voice guidance system

    CN110059996A

  • Automobile accessory logistics management method and system based on voice intelligent man-machine interaction

    CN115330290A

  • A voice assistant system for a vehicle cockpit system

    US20210358496A1

  • Using predictive user models for language modeling on a personal device

    US20070239637A1