Method for voice recognition, in particular in a motor vehicle
The method enhances speech recognition in noisy vehicle environments by using a voice sensor and transition matrix to identify probable states and keywords, addressing the limitations of existing systems and improving accuracy and user experience.
Patent Information
- Application Number
- PCT/EP2025/069107
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-08
AI Technical Summary
Existing speech recognition systems in motor vehicles are not adequately adapted to noisy environments and specific professional contexts, leading to decreased accuracy and user experience.
A method utilizing a voice sensor with a microphone connected to a voice assistant, employing a predetermined list of questions and a transition matrix to enhance speech recognition by identifying probable states and keywords, and updating based on contextual data from vehicle sensors.
Improves speech recognition accuracy and user experience in noisy vehicle environments by focusing on specific user objectives and adapting to professional needs, ensuring robustness and efficiency.
Smart Images

Figure EP2025069107_08012026_PF_FP_ABST
Abstract
Description
[0001] Description
[0002] Title: Voice recognition method, particularly in a motor vehicle
[0003] technical field
[0004] The present invention relates to a method of voice recognition, particularly in a motor vehicle.
[0005] Previous technique
[0006] Automotive voice recognition offers numerous practical and safety benefits to drivers. First, it allows drivers to control various vehicle functions, such as the climate control, audio system, and navigation, without taking their eyes off the road. This can help reduce distractions, limit accidents, and improve road safety. Furthermore, voice recognition can also facilitate communication between drivers and passengers or people outside the vehicle, for example, during hands-free phone calls or when dictating text messages. Finally, in a professional context, voice recognition could be useful for delivery drivers, who can read delivery instructions and interact with the delivery interface without having to take their eyes off the road.
[0007] It is known, particularly in the vehicles of the Filing Company, to use voice recognition. This technology allows drivers to control various functions of their car, such as the audio system, with voice commands. This allows drivers to interact with their vehicle using voice commands to perform tasks such as selecting music and making hands-free phone calls. The use of a voice-activated personal assistant system provides a smoother and more intuitive experience for drivers.
[0008] This advanced voice recognition system is integrated into select vehicles. It allows the driver to control various vehicle functions, such as navigation, music, and climate control, using simple and natural voice commands. It utilizes speech recognition technology to understand the driver's instructions and requests, providing clear and precise feedback. It is also compatible with connected vehicle controls, enabling drivers to control their favorite apps with voice commands, such as playing music and replying to messages. The voice personal assistant system is integrated with the vehicle's infotainment system (IVI), offering an intuitive and user-friendly interface for the driver.
[0009] Parcel delivery has become an increasingly important activity, particularly in the context of online sales or e-commerce. Commercial vehicles, such as vans, are used to transport and deliver parcels. Shippers face complex tasks in organizing and executing their own deliveries efficiently. In this context, voice recognition has become a valuable tool for delivery drivers, as it allows them to control their vehicles and manage their work more effectively. As shown in Figure 1, connected commerce vehicles are now equipped with voice interfaces, enabling shippers to interact with their vehicles and manage functions related to delivery, navigation, mailing lists, and, where applicable, communicate with customers using simple voice commands.This makes the delivery driver's job easier and faster, thus contributing to improved efficiency and productivity. As a result, delivery companies and carriers using commercial vehicles equipped with voice recognition can provide faster and more reliable service to their customers.
[0010] By using voice commands, drivers can interact with their vehicle without taking their eyes off the road or using their work tablet, reducing distractions and improving road safety. Furthermore, voice recognition can also help drivers navigate unfamiliar or complex environments by providing clear and precise voice guidance that keeps them focused on the road. In short, voice recognition can significantly improve road safety and make driving a commercial vehicle less accident-prone by allowing drivers to concentrate on driving without being distracted by other tasks.
[0011] As shown in Figure 2, in step 1, the delivery driver speaks the keyword and then asks their delivery-related question (Q) to the system. In step 2, the keyword is detected by the "WuW" ("Wake up Word") processing. In step 3, the Automatic Speech Recognition (ASR) system reports the detected words agnostically. Natural Voice Understanding (NLU) reports the speaker's intended meaning. Step 4 involves B2B speech recognition processing for the Q-based voice command. The system queries the customer database and receives the associated response. In step 5, the system formats the response in a bimodal way: via audio to communicate the answer, using Text-to-Speech, and via the screen for navigation if necessary.
[0012] Application CN 115330290 discloses a distribution logistics management system based on intelligent voice-activated human-machine interaction. The system comprises several steps, including the configuration of basic information and the linking of a driver and a vehicle, the generation of packaging information, the generation of delivery order information, the generation of dispatch order information, the loading of goods by the driver, the transport of goods to the target delivery point, and the confirmation of delivery. The voice interaction module identifies and analyzes the voice via the mobile application, generates corresponding text, and performs semantic analysis to generate a voice instruction.
[0013] CN application 110059996 describes a vehicle loading system using a voice guidance system. The system includes several constraints and steps to optimize the loading of goods into a given cargo space. The system uses a voice guidance system as the basis for goods management. It can provide voice instructions to the operator regarding the order in which goods are loaded and their target positions within the cargo space.
[0014] Application CN 108537480 describes a logistics management process based on speech recognition. A user ID corresponding to the goods order number is assigned, and input information for the user ID is retrieved. If the input information is speech data, it is analyzed to obtain an identification text corresponding to the speech data. A semantic analysis is then performed on the identification text, and if the semantic analysis matches cargo information, the cargo information is recorded in the cargo logistics information. Speech recognition enables the identification of users' speech data and its conversion into text for semantic analysis. Thus, instructions, requests, and reports issued verbally can be processed by the logistics management system.
[0015] US patent application 2021 / 0358496 describes a system for a motor vehicle that uses speech recognition. The system includes a multimedia device that recognizes the user's voice, converts speech to text, and analyzes the text to understand the requested action. Depending on the requested action, the system uses specific capabilities to perform that action, either locally in the vehicle or via a connection to a cloud service provider. The system then generates a voice response that is played through the vehicle's speaker.
[0016] It should be noted that existing speech recognition systems are not perfectly suited to professional needs. Indeed, these systems are currently designed to be very general, which may not meet the specific requirements of a particular application domain. In fact, general speech recognition is largely context-agnostic and is primarily designed to understand numerous commands and different queries in various contexts, which can be useful for more general applications. Furthermore, these systems are mostly network-based, whereas having embedded speech recognition would be more advantageous.
[0017] Furthermore, noisy environments and regional accents can make communication difficult, which can negatively impact driver productivity, as accuracy is crucial, and customer service quality. As shown in Figure 2, in step 3 of F ASR, the noisy environment of the commercial vehicle cab can significantly impact ASR and NLU, leading to decreased overall system performance and consequently a poor user experience.
[0018] In the automotive industry, noisy environments can include ambient noise such as engine noise, cabin noise, traffic noise, and conversations of people present, as well as potentially any other source of unwanted noise. Background noise can mask spoken words, making word recognition difficult. Speech recognition techniques use sophisticated algorithms to identify a voice and distinguish it from background noise. This requires a precise analysis of various speech characteristics, such as frequency, amplitude, and duration, to determine the phonemes, syllables, and spoken words.
[0019] Another problem is understanding speech in noisy environments. Ambient noise can alter the meaning of a sentence, making it difficult to understand the delivery person's request or instructions.
[0020] There is therefore a need to further improve speech recognition processes, particularly in motor vehicles, in order to offer speech recognition that is robust to the noise of the motor vehicle cabin and adapted to the context considered.
[0021] Summary of the invention
[0022] The present invention addresses this need through, in one aspect, a method for assisting with speech recognition, particularly in a motor vehicle, adapted to a predetermined user objective and using:
[0023] - at least one voice sensor installed in said vehicle, including a microphone, connected to a voice assistant implementing at least natural voice understanding processing,
[0024] - a predetermined list of a number N of questions Qi=i...N previously asked and learned according to the predetermined objective and / or the driving and vehicle cab context, each question corresponding to a state, said list being translated into a system of states, and
[0025] - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned, the process comprising at least the following steps: a) after detection by the voice sensor and recognition of a question posed to the voice assistant, search for the next question by choosing the two most probable states according to the transition matrix, b) identify a pair of keywords associated with the first and second states chosen in step a) from among keywords predefined beforehand for each question in the predetermined list, c) when a new question is detected by the voice sensor and recognized by the voice assistant, search in said question for the presence of the keywords of the pairs identified in step b) and:
[0026] - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or - if the keyword pair associated with the first state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this first state, or
[0027] - if the pair of keywords associated with the second state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this second state, d) continue the process with the selected question in order to facilitate future speech recognitions.
[0028] The invention improves the performance of voice recognition in the vehicle cabin. It offers a professional-focused solution, meaning it addresses more specific and precise voice recognition needs. The invention is tailored to these specific requirements to ensure maximum accuracy and an optimal user experience.
[0029] Furthermore, the invention is placed in an embedded model, which avoids performance losses due to lack of connectivity when voice recognition is on the network.
[0030] The method according to the invention makes it possible to fully exploit the processing, interpretation, and speech recognition capabilities of the voice assistant. Sequencing logic is taken into account; for example, if the user is a delivery driver, questions about the first delivery will be asked before those concerning the last.
[0031] The invention implements a probabilistic model that focuses more specifically on intention domains related to the user's goal, aiming to optimize the efficiency and accuracy of the voice assistant. This means the voice assistant can be adapted to respond more accurately and quickly to the user's questions. For example, if the user is a delivery driver, the voice assistant can efficiently answer questions about their delivery route, thus improving the overall efficiency of the delivery process.
[0032] The invention assists natural speech understanding processing by providing it with prior information about what it is expected to interpret. Thus, natural speech understanding processing is no longer solely responsible for the correct or incorrect interpretation of the user's intent. In a preferred embodiment, contextual information from the driving environment and vehicle cabin is used, including information about the user, the predetermined objective, the environment, previous actions, or geographical information. This allows for the estimation of a prediction of the user's questions based on the context.
[0033] State system and transition matrix
[0034] A transition matrix is a well-known concept used in graph theory and probability to represent the change of state in a system.
[0035] In the invention, the transition matrix represents the possible transitions between each state of said state system, learned during past occurrences, according to the predetermined objective, and updated after each occurrence, for example after each delivery in the case where the user is a delivery person.
[0036] The transition matrix is advantageously a square matrix of size N x N, where each element of the matrix represents the probability of transition from one state to another, or from one question to another.
[0037] In a preferred embodiment, the transition matrix T contains an element in the i-th row and j-th column, denoted T[i, j] or qj, which represents the probability of moving from question Qi to question Qj. The elements of the transition matrix T advantageously satisfy the known properties of a transition matrix, namely that each element of the matrix must be between 0 and 1, since it represents a probability. The sum of the elements in each row of the matrix must be equal to 1 or 100%.
[0038] [Math 1]
[0039] This matrix effectively models the sequencing of questions asked during the different stages of the predetermined objective, particularly during a delivery route. For the mth stage, the matrix can be presented as follows:
[0040] [Math 2]
[0041] For the m+l-th step, the matrix can be represented as follows:
[0042] [Math 3]
[0043] Furthermore, when no questions have yet been asked by the user (equivalent to step m = 0, or m = 1 depending on the chosen indexing method), different strategies for initializing the transition matrix can be used. For example, it is possible to initialize the matrix with identical values for each question to represent the case where each question has the same probability of being asked. It is also possible to initialize the matrix with the probability that a question is the first one.
[0044] Maximum likelihood can be used to candidate the next two most probable states, represented by the two highest probabilities.
[0045] The maximum likelihood principle, as is well known, involves estimating the parameters of a probabilistic model by finding the value that makes the observed data most probable. This allows for statistical decisions to be made based on real data and for obtaining accurate estimates of unknown parameters.
[0046] At step m, for the question Qk (m) for k :1..N, the following two most probable questions can be chosen:
[0047] [Math 4]
[0048] Qki(
[0049] Qkz( (qkj , j 5 j)
[0050] When a new question is detected by the voice sensor and recognized by
[0051] 1. Voice assistant, searches in said question for the presence of keywords from the pairs identified in step b). This will allow the selection of the most probable question from the two questions retained.
[0052] For the question Qi for i: 1..N:
[0053] WuW(i,l): wake-up word #1
[0054] WuW(i,2): wake word #2.
[0055] To achieve this, a 'wake-up word' algorithm is used. This type of well-known algorithm is more robust at detecting words without necessarily recognizing them, as it is not a natural speech comprehension process. This principle is used in speech recognition systems to activate the system based on a specific short spoken keyword, before triggering automatic speech recognition and natural speech comprehension processing. This mechanism allows devices to listen continuously but only react when the wake-up word is detected.
[0056] Popular voice assistants use a wake word. These assistants are in standby mode and don't actively listen to conversations. However, as soon as the specific wake word is spoken, the assistant wakes up and is ready to respond to commands.
[0057] This principle relies on signal feature recognition technology (mel-frequency cepstral coefficients, MFCCs, or neural networks). When the device is in listening mode, it continuously records ambient audio, but without actively processing it. However, it performs a rapid and continuous analysis of the audio signal to detect if the wake word is spoken. Wake word detection is typically achieved using machine learning models, such as neural networks, which have been trained on large amounts of speech data to specifically recognize the wake word.
[0058] The alarm word detection must be precise enough to prevent accidental activations. These systems are therefore qualified for their robustness and are continuously improved to reduce false triggers and only react to genuine alarm words.
[0059] The keyword pairs are preferably distinct in pairs so that no overlap and therefore no double detection is possible.
[0060] This "wake-up word" detection principle allows the most probable question to be identified based on signal characteristics. The wake-up words present in the candidate question indicate its likelihood, which can then be communicated to the natural speech understanding processing. Based on their detection, the processing task will be facilitated.
[0061] When the new question is finally determined, it is advantageously passed on to a processing module associated with the predetermined objective.
[0062] During step d), the transition matrix can be updated with the selected question.
[0063] The updating of the coefficients of the transition matrix is carried out by the method based on the data related to the occurrences of the questions during an occurrence according to the predetermined objective.
[0064] The coefficients can be updated according to the following equation:
[0065] [Math 5]
[0066] Tp, j] = qij = n(i, j) / n Jotal(l) with n(i, j) the number of occurrences of the transition from state i to state j observed in the empirical data. The total number of occurrences of all transitions starting from state i is given by n_total(i).
[0067] This means that the probability of going from state i to state j is equal to the number of occurrences of this transition divided by the total number of occurrences of all transitions starting from state i.
[0068] The updating of the coefficients of the transition matrix between time t (question asked at t equivalent to step m) and time t+dt (question that will be asked at t+dt equivalent to step m+1) is preferably based on empirical observations and questions of transitions between the states of the system.
[0069] The update of the coefficients of the transition matrix between time t and time t+1 can be done according to the following equation:
[0070] [Math 6]
[0071] T(t+dt)[i, j] = qij = n(t, i, j) / n_total(t, i) with T(t) the transition matrix at time t, T(t)[i, j] which represents the transition coefficient from state i to state j at time t. n(t, i, j) the number of occurrences of the transition from state i to state j observed between time t and time t+dt, n_total(t, i) the total number of occurrences of all transitions starting from state i between time t and time t+dt, that is to say the sum of n(t, i, j) for all states j.
[0072] Taking the step into account, the expression can be expressed as follows:
[0073] [Math 7]
[0074] T(m+l)[i, j] = qij = n(m, i, j) / n otal(m, i)
[0075] This means that the probability of going from state i to state j between time t and time t+dt is equal to the number of occurrences of this observed transition divided by the total number of occurrences of all transitions starting from state i in the period between time t and time t+dt.
[0076] Preferably, the general equation for updating row i of the transition matrix is as follows:
[0077] [Math 8]
[0078] T[i, j] = P(j|i) = qij with T[i, j] the element of the transition matrix located in row i and column j.
[0079] Since this also represents the probability of transition from state i to state j, P(j |i) is the conditional probability of moving from state i to state j.
[0080] Contextual information
[0081] The method according to the invention advantageously uses information from the driving context and vehicle cab, including information about the user, the predetermined objective, the environment, previous actions, or geographical or road information.
[0082] Contextual information is advantageously derived from time series of data from a plurality of sensors included in the vehicle, said time series being transformed into probability vectors and grouped into data groups according to predetermined training parameters.
[0083] The numerous contextual information or "Car Data" reflects the journey context or "Driving Task", as well as the vehicle cab context or "In Vehicle Task".
[0084] This information evolves over time. It can be optimally collected through digital representations automatically learned from raw data, allowing machine learning models to better understand and utilize the information contained within that data. These representations, also called "AI embeddings" or simply "embeddings," of extracted features, embeddings, or compact semantic vectors, play a crucial role in many areas of artificial intelligence by facilitating the processing and analysis of complex data.
[0085] In the field of smart cars, these visualizations are useful for analyzing a vehicle's driving data. When a vehicle is in motion, it generates a wealth of data from integrated sensors, such as GPS sensors, accelerometers, speed sensors, cameras, ADAS (Advanced Driver Assistance Systems), and more. This data is critical for understanding and improving vehicle behavior and performance, safety, maintenance, and ADAS systems.
[0086] These time series of data correspond to numerical descriptions of the state of the driver / vehicle control system at a given moment, which include all the information relevant to understanding the instantaneous state of the journey, such as position, speed, acceleration, altitude, etc. By creating contextual data time series for each instant of the journey, a detailed view of what is happening throughout the journey is obtained.
[0087] These time series of data come from different sensors, each emitting its own data stream providing different information, for example:
[0088] Driver Monitoring System (DMS) camera: all DMS reports, such as alertness, nervousness, and fatigue.
[0089] Visual attention / road / landscape / cabin: measurement of the field of vision,
[0090] Pedal pressure: pressure and its frequency on the pedals indicating the level of attention to pedal actions.
[0091] Steering wheel grip and contact: a tight or loose grip on the steering wheel, with one or two hands, indicates the driver's attention or lack of focus.
[0092] HMI and radio manipulation, air conditioning...: the HMI system and its operating states at the time of measurement,
[0093] Blinking: camera measurements allow the periodicity or aperiodicity of the driver's blinks to be captured.
[0094] Driving time: the duration of driving. Risk-taking: a driver in a hurry may adopt a driving style that tends to disregard the highway code and the speed limit.
[0095] Data related to the road context can also be recorded, such as, for example:
[0096] Functional Road Class #0...8: road type identification,
[0097] Weather conditions: the weather situation,
[0098] Winding / Straight line: the evolution of the road's shape,
[0099] Slope / Ascent: the evolution of the road's shape,
[0100] Urban context (city, density), Peri-urban, Countryside, Mountain
[0101] Anticipating road events: navigation is a critical indicator for anticipating the cognitive load required for upcoming road situations in the near future.
[0102] Smooth traffic flow, Night / Day, Works, Stop, Hour, Day.
[0103] Data related to the vehicle context can also be recorded, such as:
[0104] Speed, Acceleration (deceleration): long, lat.,
[0105] Speeding
[0106] Front-rear safety index / sensors: ADAS alarms, steering angle, angular velocity,
[0107] Contrôleurs ADAS: l’état des contrôleurs ADAS sont directement des données de mesures, comme par exemple : o AEBS: Automatic Emergency Braking System o ACC: Adaptive Cruse Control o FCW: Forward Collision Warning o EAPM: Emergency Assist for Pedal Misapplication o ESP: Electronic Stability Program o DW: Distance Warning o OSP: Over Speed Prevention o ISA: Intelligent Speed Assistance o LDW: Lane Departure Warning / Emergency o LKA: Lane Keeping Assist o LDP: Lane Departure Prevention o LS S: Advanced Lateral Support o UTA: Unstable Trajectory Alert o DAA: Driver Attention Alert o B SW: Blind Spot Warning o BSI: Blind Spot Intervention o CTA: Cross Traffic Alert o BCI: Back up Collision Intervention o R-CTA: Rear Cross Traffic Alert o CASP : Contextual Adaptive Speed Prevention.
[0108] Raw data from various sensors can be complex and difficult to process directly. Machine-learned numerical representations allow this raw data to be converted into more understandable numerical representations, while preserving essential information. In this context, time series data are advantageously transformed into probability vectors. These are designed to compactly represent the key characteristics of the vehicle's various driving modes.
[0109] Using these digital representations, several useful analyses can be performed, such as, among others, anomaly detection, machine learning, and driving optimization.
[0110] As is well known, AI embeddings in the context of vehicle driving data use the concept of an encoder / decoder based on neural networks to transform raw data into meaningful representations. As shown in Figure 3, the encoder takes driving data as input, such as speed, acceleration, steering angles, etc., and converts it into compact numerical representations that capture the key characteristics of driving patterns. The decoder, in turn, uses these numerical representations as initial context to generate useful information, such as anomaly detection, driving event prediction, or driving optimization. AI embeddings thus allow users to focus on the essential and relevant data within the very large dataset of vehicle data.
[0111] AI embeddings are, as is well known, grouped into data sets or clusters, as shown in Figure 4. Clustering numerical representations from a vehicle's driving data is a powerful method for grouping similar driving patterns. This offers several significant advantages in analyzing vehicle driving data, such as
[0112] - Driving pattern detection: By using clustering algorithms on embeddings, similar driving patterns can be automatically identified across different recorded lists. This can reveal characteristic driving styles, recurring behaviors, or even specific scenarios such as parking maneuvers, sudden acceleration, etc. This information can be invaluable for understanding driver behavior and improving road safety.
[0113] - Driving segmentation: Clustering embeddings allows driving data to be divided into segments or specific journeys. This can be useful for analyzing each journey individually, identifying high-risk driving areas, frequent routes, or particular traffic conditions.
[0114] - Prediction of driving events: the clusters obtained from the "embeddings" can be used as input for machine learning models. By training a model on the clustered driving data, it becomes possible to predict certain events, such as dangerous situations, sudden changes in driver behavior, or even future traffic conditions.
[0115] A non-exhaustive list of road context information may include:
[0116] Accident on the main road, anticipated question: "Is there any information about the accident and any suggestions to avoid further delays?"
[0117] Road closed due to construction work, anticipated question: "Is there a recommended detour to avoid the work zone?" Heavy traffic jam, anticipated question: "Is there a faster alternative to avoid this traffic jam?"
[0118] Bad weather conditions (heavy rain), anticipated question: "Are there safer roads to drive on in the rain?"
[0119] Major roadworks, anticipated question: "Is there an alternative route due to the works?"
[0120] Temporary closure of the main road, anticipated question: "What is the best way to get around the road closure?" Traffic jam during the morning rush hour, anticipated question: "Are there any less congested secondary roads at this time?"
[0121] Prolonged traffic jam, anticipated question: "What is the estimated delay for my delivery due to the current traffic jam?"
[0122] Road closed due to a sporting event, anticipated question: "Is there an alternative route to bypass the road closure due to the sporting event?"
[0123] Level crossing blocked by a train, anticipated question: "Is there an alternative route to avoid the level crossing blocked by the train?"
[0124] Traffic congestion due to a demonstration, anticipated question: "Can you direct me to a less congested route due to the ongoing demonstration?"
[0125] Roadworks with detour, anticipated question: "What is the best way to reach the destination using the current roadworks detour?" Winter weather conditions (snow), anticipated question: "Are there any clear and safe roads to drive on during periods of snow?"
[0126] Strong gusts of wind, anticipated question: "Are there roads less exposed to strong winds to ensure safer driving?"
[0127] Construction work on a highway, anticipated question: "How can I avoid delays due to construction work on the highway?"
[0128] Traffic lights out of service, anticipated question: "Are there any special instructions for safely crossing intersections without traffic lights?" Traffic jam caused by an accident, anticipated question: "Can you provide me with information on the duration of the traffic jam and available alternative routes?"
[0129] Temporary closure of a lane due to maintenance work, anticipated question: "Are there any detours in place to bypass the lane closure due to maintenance work?"
[0130] Traffic slowdown caused by road repair work, anticipated question: "What is the expected impact on travel time due to the ongoing road repair work?"
[0131] A non-exhaustive list of vehicle and driver contextual information may include: Excessive fatigue, anticipated question: "Are there any recommended rest areas nearby where I can take a break and rest?"
[0132] Low fuel level, anticipated question: "Where can I find the nearest gas station to refuel?"
[0133] Need to take a break for food? Anticipate the question: "What restaurants or fast food outlets are available along my route?" Need to find a restroom? Anticipate the question: "Are there any accessible rest areas or public restrooms near my current location?"
[0134] Need to find a secure parking lot, anticipating the question: "Can you tell me about available parking lots nearby where I can safely park my delivery vehicle?"
[0135] Need to find an electric charging point, anticipated question: "Where can I locate the nearest electric vehicle charging stations to charge my delivery vehicle?"
[0136] Need mechanical assistance? Anticipated question: "Can you recommend a reliable garage or roadside assistance service nearby?"
[0137] The delivery driver wants to know if the vehicle data indicates a problem with the brakes, thus ensuring safety during stops and slowdowns; the anticipated question is: "Is there a problem with the brakes?"
[0138] This question helps to check if the tires are correctly inflated, as inadequate pressure can affect handling and fuel consumption. Anticipated question: "What is the tire pressure level?"
[0139] The delivery driver wants to know the average fuel consumption of his vehicle, which will help him to assess costs and optimize his routes for better energy efficiency. Anticipated question: "What is the average fuel consumption?"
[0140] A subset of likely questions can be associated with each data group to improve natural voice understanding processing.
[0141] The frequencies of the N questions can be counted for each of the M data groups. At the nth occurrence n_i of group i, question k, issued among the N questions, can be counted. The probability of question k in data group i corresponds advantageously to the number of times this question was issued in the number of occurrences of the context of data group i, expressed by the following equation, for all possible questions k:1..N:
[0142] [Math 9]
[0143] P(Qk | Cluster J) = nb(Qk) / nb(n)
[0144] Each data group is advantageously associated with a probability vector of the user's questions about the context.
[0145] The invention may relate to a method for assisting speech recognition, particularly in a motor vehicle, adapted to a predetermined user objective and using:
[0146] - at least one voice sensor installed in said vehicle, including a microphone, connected to a voice assistant implementing at least natural voice understanding processing,
[0147] - a predetermined list of a number N of questions Qi=i...N previously asked and learned according to the driving and vehicle cab context, each question corresponding to a state, said list being translated into a system of states, and
[0148] - a transition matrix representing the possible transitions between each state of said system of states, the transitions having been previously learned, the process comprising at least the following steps: after detection by the voice sensor and recognition of a question posed to the voice assistant, time series of data from a plurality of sensors included in the vehicle are acquired, said time series providing information on driving context and vehicle cabin, said time series are transformed into probability vectors and grouped into data groups according to predetermined training parameters, the next question is sought by choosing the two most probable states according to the transition matrix,identify a pair of keywords associated with the first and second states chosen in the previous step from predefined keywords for each question in the predetermined list; when a new question is detected by the voice sensor and recognized by the voice assistant, search in said question for the presence of the keywords from the pairs identified in step b) and :,
[0149] - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or
[0150] - if the pair of keywords associated with the first state chosen in step a) is detected, the natural voice understanding processing continues with the indication of the question associated with this first state, or
[0151] - if the pair of keywords associated with the second state chosen in step a) is detected, the natural voice comprehension processing continues with the indication of the question associated with this second state,
[0152] - continue the process with the selected question in order to facilitate future speech recognition.
[0153] The probability vector is preferably updated for the group concerned.
[0154] In one embodiment, questions related to the predetermined objective and questions related to the driving and vehicle cab context are used. The final choice is thus optimized through the suggestions from these two aspects.
[0155] The selection of the two most probable states is advantageously carried out using the probabilities of the transition matrices of the two state systems. The process then continues with the identification of the keyword pair. The transition matrix is updated with the selected question, as well as the probability vector.
[0156] computer program product
[0157] The invention also relates, in another of its aspects, to a computer program product comprising a medium and, recorded on this medium, instructions readable by a processor so that, when executed, they enable the implementation of the speech recognition assistance method according to the invention.
[0158] The characteristics stated in relation to the process apply to the computer program product and vice versa.
[0159] Device
[0160] According to another aspect, the invention relates to a speech recognition aid device, particularly in a motor vehicle, adapted to a predetermined user objective and comprising at least one voice sensor disposed in said vehicle, in particular a microphone, connected to a voice assistant implementing at least natural speech understanding processing and comprising or being connected to at least one module, said module having access to: a predetermined list of a number N of questions Qi=i...N previously asked and learned according to the predetermined objective and / or the driving and vehicle cabin context, each question corresponding to a state, said list being translated into a system of states, and
[0161] - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned, said module being configured to perform at least the following steps: a) after detection by the voice sensor and recognition of a question posed to the voice assistant, search for the next question by choosing the two most probable states according to the transition matrix, b) identify a pair of keywords associated with the first and second states chosen in step a) from among keywords predefined beforehand for each question in the predetermined list, c) when a new question is detected by the voice sensor and recognized by the voice assistant, search in said question for the presence of the keywords of the pairs identified in step b) and:
[0162] - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or
[0163] - if the pair of keywords associated with the first state chosen in step a) is detected, the natural voice understanding processing continues with the indication of the question associated with this first state, or
[0164] - If the keyword pair associated with the second state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this second state; d) continue the process with the selected question to facilitate future speech recognition. The characteristics stated in relation to the process apply to the device and vice versa.
[0165] Motor vehicle
[0166] According to another aspect, the invention relates to a motor vehicle comprising a powertrain and at least one device according to the invention.
[0167] The characteristics stated in relation to the process apply to the vehicle and vice versa.
[0168] Brief description of the drawings
[0169] The invention will be better understood upon reading the detailed description that follows, a non-limiting example of its implementation, and upon examination of the accompanying drawing, in which
[0170] [Fig 1] Figure 1, already described, illustrates an example of the use of voice recognition in a motor vehicle to assist a delivery driver,
[0171] [Fig 2] Figure 2, already described, illustrates different steps implemented during a delivery route,
[0172] [Fig 3] Figure 3, already described, illustrates the mechanism of "AI embeddings" in the context of vehicle driving data,
[0173] [Fig 4] Figure 4, already described, illustrates the grouping of data in the form of data groups,
[0174] [Fig 5] Figure 5 is a logic diagram illustrating the steps of an example implementation of the process according to the invention,
[0175] [Fig 6] Figure 6 is a logic diagram illustrating the steps of another example of implementing the method according to the invention, and
[0176] [Fig 7] Figure 7 is a logic diagram illustrating the steps of another example of implementation of the process according to the invention.
[0177] Detailed description
[0178] Figure 5 illustrates an example of the implementation of the speech recognition assistance method according to the invention. In the illustrated example, the method, adapted to a predetermined objective of a delivery person, takes place in a motor vehicle and uses: - the voice sensor located in said vehicle and connected to a voice assistant implementing at least natural speech understanding processing,
[0179] - a predetermined list of a number N of questions Qi=i...N previously asked and learned according to the predetermined objective, each question corresponding to a state, said list being translated into a system of states, and
[0180] - a transition matrix representing the possible transitions between each state of said system of states, the transitions having been previously learned.
[0181] In the case of a delivery route, the questions the driver asks the voice assistant are predefined within a specific area. They refer to route statuses such as: route summary, next stop, remaining tasks, time remaining, urgent delivery request, cancellation, call / text to fleet manager / boss, customer call, and others.
[0182] Here is a non-exhaustive list of questions:
[0183] 1. What is my next step?
[0184] 2. Where do I need to go?
[0185] 3. And after this delivery?
[0186] 4. Where should I deliver the next package?
[0187] 5. What is the address of the next delivery?
[0188] 6. Where will I drop off my next package?
[0189] 7. How many packages do I have left?
[0190] 8. What remains to come?
[0191] 9. How many are left?
[0192] 10. How many addresses are left?
[0193] 11. Are there any urgent deliveries?
[0194] 12. Am I late somewhere?
[0195] 13. Urgent packages?
[0196] 14. Do we have any priority mail today?
[0197] 15. Do you know what time we have to leave today?
[0198] 16. How long will it take?
[0199] 17. What is the expected time of our final delivery?
[0200] 18. Can I cancel this job?
[0201] 19. Can you repeat what you said? 20. Sorry, can you repeat that?
[0202] Each question is posed within a specific list that perfectly reflects the sequence of actions performed during the delivery route. This pre-established list is the result of meticulous planning aimed at optimizing the efficiency of the delivery route, taking into account unforeseen circumstances.
[0203] For example, when the delivery driver asks question #3, it means he is at a certain stage of his route, which is defined as a state in the process according to the invention. From this state, corresponding to a specific associated stage, he will not ask just any subsequent question, but the one most likely to be asked in relation to the next stage of his route.
[0204] After steps 0 and 1 of detection by the vehicle's voice sensor and recognition of a question asked to the voice assistant, the next question is sought by choosing the two most probable states according to the transition matrix, in steps 2 and 3.
[0205] In step 4, a pair of keywords associated with the first and second states chosen in the previous step is identified from predefined keywords for each question in the predetermined list. For example, for the question "What is the address of the next delivery?", the keyword pair might be "address" and "delivery". For the question "Do we have priority mail today?", the keyword pair might be "mail" and "priority".
[0206] At step 5, when a new question is detected by the voice sensor and recognized by the voice assistant, a search is conducted in step 6 for the presence in said question of the keywords of the pairs identified in step 4.
[0207] At step 7, if no keyword is detected, the voice assistant continues its natural speech recognition process without further input. Conversely, if the keyword pair associated with the first state selected in step 3 is detected, the natural speech recognition process continues, including the prompt for the question associated with that first state. Finally, if the keyword pair associated with the second state selected in step 3 is detected, the natural speech recognition process continues, again including the prompt for the question associated with that second state.
[0208] The process is then continued in step 8 with the selected question to facilitate future speech recognition. The transition matrix is updated with the selected question in step 9. In the variant shown in Figure 6, after steps 0 and 1 of detection by the vehicle's voice sensor and recognition of a question posed to the voice assistant, driving and vehicle cabin context information is acquired in step 2, including information about the user, the predetermined objective, the environment, previous actions, and geographical information, as defined previously. This contextual information comes from time series data from multiple sensors included in the vehicle.
[0209] These time series are transformed into probability vectors and grouped into data groups according to predetermined training parameters, in step 4. A subset of probable questions is associated with each data group in order to improve natural voice comprehension processing, in step 5.
[0210] In step 6, after the voice sensor detects and recognizes a question asked to the voice assistant, the next question is sought by choosing the two most probable states based on the transition matrix, in step 7.
[0211] In step 8, a pair of keywords associated with the first and second states chosen in the previous step is identified from a predefined list of keywords for each question. When a new question is detected by the voice sensor and recognized by the voice assistant in step 9, step 10 searches the question for the presence of the keyword pairs identified in step 8.
[0212] In step 11, if no keyword is detected, the voice assistant continues its natural speech recognition process without further input. Conversely, if the keyword pair associated with the first state selected in step a) is detected, the natural speech recognition process continues, including the prompt for the question associated with that first state. Finally, if the keyword pair associated with the second state selected in step a) is detected, the natural speech recognition process continues, including the prompt for the question associated with that second state.
[0213] The process continues in step 12 with the selected question to facilitate future speech recognition. The probability vector is updated for the relevant group. In the variant shown in Figure 7, in step 2, both processes from Figures 5 and 6 are performed. Questions related to the predetermined objective and questions related to the driving and vehicle cab context are used in step 3. The selection of the two most probable states is based on the probabilities of the transition matrices of the two state systems. The process then continues with the identification of the keyword pair. The transition matrix is updated with the selected question, as well as the probability vector.
[0214] The invention is not limited to the examples just described.
[0215] The invention can be used by a voice assistant in a passenger vehicle by identifying target areas and using the principle of driving and vehicle cabin context information.
[0216] The invention can be implemented in any field using speech recognition specific to a given field of activity, such as aviation for example.
Claims
Demands 1. A method for assisting with speech recognition, particularly in a motor vehicle, adapted to a predetermined user objective and using: - at least one voice sensor installed in said vehicle, including a microphone, connected to a voice assistant implementing at least natural voice understanding processing, - a predetermined list of a number N of questions Qi=i...N previously asked and learned according to the predetermined objective and / or the driving and vehicle cab context, each question corresponding to a state, said list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned, the process comprising at least the following steps: a) after detection by the voice sensor and recognition of a question posed to the voice assistant, search for the next question by choosing the two most probable states according to the transition matrix, b) identify a pair of keywords associated with the first and second states chosen in step a) from among keywords predefined beforehand for each question in the predetermined list, c) when a new question is detected by the voice sensor and recognized by the voice assistant, search in said question for the presence of the keywords of the pairs identified in step b) and: - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or - if the pair of keywords associated with the first state chosen in step a) is detected, the natural voice understanding processing continues with the indication of the question associated with this first state, or - if the pair of keywords associated with the second state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this second state, d) continue the process with the selected question in order to facilitate future speech recognitions.
2. A method according to claim 1, wherein, at step d), the transition matrix is updated with the selected question.
3. Method according to claim 1 or 2, using vehicle cab and driving context information, including information on the user, on the predetermined objective, on the environment, on previous actions, or geographical information.
4. A method according to the preceding claim, wherein the context information comes from time series of data from a plurality of sensors included in the vehicle, said time series being transformed into probability vectors and grouped into data groups according to predetermined training parameters.
5. A method according to the preceding claim, wherein a subset of probable questions is associated with each data group in order to improve natural voice understanding processing.
6. A method according to any one of claims 4 or 5, wherein the frequencies of the N questions are counted for each of the M data groups, at the nth occurrence n_i of group i, question k is counted among the N questions, the probability of question k in data group i corresponding to the number of times that this question was asked in the number of occurrences of the context of data group i, expressed by the following equation, for all possible questions k:1..N: [Math 9] P(Qk | Cluster ) = nb(Qk) / nb(n_i) 7. A method according to any one of claims 4 and 5, wherein the probability vector is updated for the group concerned.
8. A method according to any one of the preceding claims, wherein the transition matrix is a square matrix of size N x N, where each element of the matrix represents the probability of transition from one state to another.
9. A method according to any one of the preceding claims, wherein maximum likelihood is used to candidate the next two most probable states, represented by the two highest probabilities.
10. A speech recognition assistance device, particularly in a motor vehicle, adapted to a predetermined user objective and comprising at least one voice sensor located in said vehicle, including a microphone, connected to a voice assistant implementing at least natural speech understanding processing and comprising or being connected to at least one module, said module having access to: a predetermined list of a number N of questions Qi=i...N previously asked and learned according to the predetermined objective and / or the driving and vehicle cabin context, each question corresponding to a state, said list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of said state system, the transitions having been previously learned, said module being configured to perform at least the following steps: d) after detection by the voice sensor and recognition of a question posed to the voice assistant, search for the next question by choosing the two most probable states according to the transition matrix, e) identify a pair of keywords associated with the first and second states chosen in step a) from among keywords predefined beforehand for each question in the predetermined list, f) when a new question is detected by the voice sensor and recognized by the voice assistant, search in said question for the presence of the keywords of the pairs identified in step b) and: - if no keyword is detected, the voice assistant continues recognition by processing natural speech understanding without additional information, or - if the pair of keywords associated with the first state chosen in step a) is detected, the natural voice understanding processing continues with the indication of the question associated with this first state, or - if the pair of keywords associated with the second state chosen in step a) is detected, the natural speech understanding processing continues with the indication of the question associated with this second state, d) continue the process with the selected question in order to facilitate future speech recognitions.
11. Product computer program comprising a medium and, recorded on this medium, instructions readable by a processor so that, when executed, they enable the implementation of the speech recognition assistance method according to any one of claims 1 to 9.
12. Motor vehicle comprising a powertrain and at least one identification device according to claim 10.
Citation Information
Patent Citations
Logistics management method and system based on speech recognition
CN108537480A
Field area vehicle loading method based on voice guidance system
CN110059996A
Automobile accessory logistics management method and system based on voice intelligent man-machine interaction
CN115330290A
A voice assistant system for a vehicle cockpit system
US20210358496A1
Using predictive user models for language modeling on a personal device
US20070239637A1