Method for a motor vehicle for processing voice inputs of several occupants of the motor vehicle, computer program and / or computer-readable medium, data processing device and motor vehicle

The method improves voice command processing in motor vehicles by identifying and prioritizing individual speech inputs based on context and assignment, addressing errors from overlapping speech, ensuring accurate function execution.

DE102024101092A1Pending Publication Date: 2025-07-17BAYERISCHE MOTOREN WERKE AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE102024101092
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing voice command systems in motor vehicles struggle to accurately process multiple overlapping speech inputs from different occupants, leading to errors and misinterpretations, particularly when multiple people speak simultaneously.

Method used

A method for processing voice inputs in motor vehicles that involves detecting and identifying individual speech inputs, determining semantic content, assigning each input to a specific occupant, and applying a location-independent prioritization based on speech content and context to execute functions accurately.

Benefits of technology

Enhances the robustness of voice command processing by accurately identifying and prioritizing individual speech inputs, allowing for efficient execution of functions while minimizing errors, even with overlapping or alternating speech from multiple occupants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for a motor vehicle for processing speech inputs from multiple occupants of the motor vehicle; the method comprising: detecting the speech inputs, wherein at least two of the speech inputs are spoken by different occupants; determining semantic speech content, wherein each of the speech contents corresponds to one of the speech inputs; determining assignment information, wherein the assignment information comprises an assignment between each of the speech contents and one of the occupants; determining a location-independent prioritization of the speech inputs taking into account the speech content and the assignment information; and outputting an execution signal for executing a function based on the speech inputs taking into account the prioritization.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present disclosure relates to a method for a motor vehicle for processing voice inputs from multiple occupants of the motor vehicle. The disclosure also relates to a computer program and / or computer-readable medium, a data processing device, and a motor vehicle.

[0002] Voice input by occupants of motor vehicles or vehicles is a widespread technology. Special systems, such as third-party personal assistants based on cloud services, are usually used to enable voice input. Content spoken by a occupant can be converted by the assistant into a command to apply a function, particularly that of the assistant. In-house developments of such assistants by motor vehicle manufacturers are also known. These systems enable more comprehensive access to motor vehicle functions, for example, to control the vehicle's navigation system via voice command, to make calls via the vehicle, and / or to play music in the vehicle.

[0003] Even though speech quality and recognition have improved significantly in recent years, speech systems and assistants can still be prone to errors, particularly when several people are in the vehicle speaking at the same time, or when their voice inputs at least partially overlap and / or alternate. This is because the systems typically attempt to record all voices and interpret their content as a command to perform a function. When several people are speaking at the same time, the content of a command cannot usually be fully determined. For example, if the driver says the sentence "Drive to XY-Allee 4 in Regensburg" and is interrupted in the middle of the sentence by another passenger, for example with "Dad, can you turn on the radio play XYZ?", problems with speech processing can arise.

[0004] To avoid such problems, it is known to localize the spatial origin of the voice command in order to differentiate speakers from speech inputs. US 2006 / 0212291 A1 discloses a speech recognition system, a speech recognition method, and a storage medium capable of recognizing the voice command of a single speaker, even in a case where multiple speakers input overlapping voice commands, and of making a single application program usable in execution among the speakers. The received voice commands are separated and recognized according to the respective speakers as needed.Optionally, a priority level indicating the priority for selecting a speech recognition result for each speaker is stored, or a priority level is specified in the utterance order, and a selection means selects a speech recognition result of a speech uttered by a speaker with a prestored highest priority level. Ultimately, prioritization is thus based on spatial allocation.

[0005] DE 10 2014 114 604 A1 discloses a method for processing a plurality of audio streams in an on-board computing system of a vehicle. The method comprises receiving the plurality of audio streams from a plurality of locations in a vehicle; prioritizing each of the plurality of audio streams to generate a prioritization result; and executing an application associated with each of the plurality of audio streams depending on the prioritization result.

[0006] In other words, DE 10 2014 114 604 A1 describes a method that receives multiple audio streams in a vehicle's on-board computer system, prioritizes them, and executes a corresponding application based on the prioritization result. An application is identified to be assigned to each audio stream, and a priority is set for each application.

[0007] According to the current state of the art, the processing of voice commands is subject to a relatively rigid set of rules. Consideration of context and / or correction of a voice command from one occupant by a voice command from another occupant does not appear to be provided.

[0008] Against the background of this prior art, one object of the present disclosure is to provide a method suitable for enriching the prior art and improving at least the above-mentioned aspects of the prior art. In particular, the object of the disclosure is to provide improved processing of voice commands from multiple occupants of a motor vehicle.

[0009] The problem is solved by the features of the independent claims. The subclaims contain further developments of the disclosure.

[0010] According to one aspect of the disclosure, the object is achieved by a method for a motor vehicle for processing speech inputs from a plurality of occupants of the motor vehicle; the method comprising: detecting the speech inputs, wherein at least two of the speech inputs are spoken by different occupants; determining semantic speech content, wherein each of the speech contents corresponds to one of the speech inputs; determining assignment information, wherein the assignment information comprises an assignment between each of the speech contents and one of the occupants; determining a location-independent prioritization of the speech inputs taking into account the speech content and the assignment information; and outputting an execution signal for executing a function based on the speech inputs taking into account the prioritization.

[0011] The method can enable more robust speech input in the motor vehicle, particularly when multiple occupants overlap and / or alternate speaking, which can be detected as speech input from the respective occupant(s) overlapping and / or alternating with another speech input. The core idea is to recognize the speech inputs or voices of all occupants or people and to process the spoken speech content or texts together using defined rules in order to be able to perform exactly one function. These rules are assigned to the people and can in turn have different meanings depending on the action to be performed. For this purpose, the speech content of a speech input from one of the occupants can be linked to the occupant using the assignment information. The assignment information therefore includes details about which occupant produced which speech content resulting from one of the speech inputs.The rules for joint processing are covered by the prioritization.

[0012] Prioritization is location-independent and can therefore depend solely on the occupant or their identity. This makes it possible, for example, for a passenger to be assigned different seating positions within the vehicle on different journeys.

[0013] Prioritization is based on the speech content and the mapping information. The prioritization thus reflects the occupants' context-related competence, reflected by the consideration of the speech content, to be able to define and / or execute a function in a context that can be assigned to the speech content. In other words, prioritization can be used to establish contextual knowledge about the individual people in order to use this knowledge to process the speech content of the various occupants in the context.

[0014] The method is particularly useful for motor vehicles, as it allows for comparatively efficient and reliable identification of individuals within the vehicle, for example, based on voice input, user devices, tags, and / or visual features. Furthermore, the method is limited to a relatively small number of individuals, which can simplify implementation. Spatial allocation of the occupants may be unnecessary, but is possible.

[0015] Optionally, the prioritization can be defined statically and / or context-related. This can ensure that the voice input of, for example, a child and / or another predetermined occupant is statically disregarded when determining and outputting the function, i.e., regardless of the voice content. In addition, the voice input of, for example, an occupant identified as the driver can be given a comparatively high priority and thus be considered to be worthy of consideration. Context-related prioritization, i.e., prioritization dependent on the voice content, can, for example, allow a child to execute a function for playing audio books, but not for navigation. In other words, contextual knowledge about the occupants can be determined ad hoc, e.g., based on age, characteristic (child yes / no), and / or retrievably stored information about the occupants.This ensures that prioritization is not subject to a rigid set of rules.

[0016] Optionally, the method includes classifying the speech content, and prioritizing it depends on the classification. This allows for efficient prioritization. For example, the speech content can be classified according to whether it is relevant to driving the motor vehicle or not, but rather relevant for entertainment.

[0017] Optionally, the prioritization can be defined by a voice input from one of the occupants. This allows the prioritization to be specified via voice input, for example, by a driver and / or owner of the vehicle.

[0018] Optionally, the procedure includes checking the prioritization and / or function. This allows a consistency check to be performed. Before the function is executed, the prioritization and / or function can be checked, for example, using the voice input of the driver and / or vehicle owner, to ensure that no functions are executed against the driver's will and thus potentially distract the driver.

[0019] Optionally, the function is based on voice input from several of the occupants. It was recognized that the voice input from multiple occupants can lead to the execution of exactly one function. To determine and output the execution signal, the voice content is evaluated jointly according to the prioritization, and a joint request to execute the function is generated or synthesized.

[0020] Optionally, the speech content of a voice input from a first of the occupants at least partially replaces the speech content of a voice input from a second of the occupants. It was recognized that the method makes it possible, for example, to supplement, replace, and / or correct part of the speech content of the first occupant with speech content from the second occupant. This can, for example, allow an adult passenger in the back seat to refine and / or correct a voice input from the driver for navigation. For example, the driver says "Set the destination Kaiserstrasse 85," the rear passenger corrects "Kaiserstrasse 85a," and the function to be executed can then be "Set the destination Kaiserstrasse 85a."

[0021] According to one aspect of the disclosure, a computer program and / or a computer-readable medium is provided. The computer program and / or the computer-readable medium comprise instructions which, when the program or instructions are executed by a data processing device, cause the device to perform the method according to the disclosure and / or steps thereof. Optionally, the computer program and / or the computer-readable medium comprise instructions which, when the program or instructions are executed by a data processing device, cause the device to perform the method steps described as advantageous or optional in order to achieve an associated technical effect.

[0022] According to one aspect of the disclosure, a data processing device for a motor vehicle is provided. The data processing device is configured to perform the method described above. Optionally, the data processing device is configured to perform a method step described as advantageous or optional and / or to implement a method feature in order to achieve an associated technical effect.

[0023] According to one aspect of the disclosure, a motor vehicle comprising the data processing device described above is provided. Optionally, the data processing device of the motor vehicle and / or the motor vehicle is configured to perform a method step described as advantageous or optional and / or to implement a method feature in order to achieve an associated technical effect.

[0024] In the following, one embodiment is described with reference to the figures. Fig. 1 schematically shows a motor vehicle according to one aspect of the disclosure; Fig. 2 schematically shows a flow diagram of a method according to one aspect of the disclosure; and Fig. 3 shows a schematic representation of a computer program and / or computer-readable medium according to one aspect of the disclosure.

[0025] Fig. 1 schematically shows a motor vehicle 50 according to one aspect of the disclosure. The motor vehicle 50 is a land vehicle. The motor vehicle 50 is a passenger car.

[0026] A plurality of occupants 10, 10a, 10b are arranged in the motor vehicle 50, i.e., in an interior (not indicated) or a cabin of the motor vehicle 50. Their arrangement and number are shown only schematically and / or by way of example. One of the occupants 10, 10a is, for example, a driver of the motor vehicle 50, i.e., a person who drives the motor vehicle 50 and / or is responsible for driving the motor vehicle 50. The other occupants 10, 10b are passengers.

[0027] The motor vehicle 50 has a data processing device 51. The motor vehicle 50 or the data processing device 51 is designed to process the data relating to Fig. 2 described method 100. For this purpose, the motor vehicle 50 has Fig. 1 has an input device 55 and a functional device 56.

[0028] The input device 55 is configured to capture voice inputs 11, wherein at least two of the voice inputs 11 are spoken by different occupants 10, 10a, 10b. The input device 55 comprises one or more microphones for this purpose. The input device 55 and the data processing device 51 are communicatively coupled to one another so that the input device 55 can transmit the voice inputs 11 to the data processing device 51, whereby the data processing device 51 can capture the voice inputs 11.

[0029] The voice inputs 11 can overlap in time, meaning that the occupants 10, 10a, 10b speak at least partially simultaneously. The voice inputs 11 can also be spoken in a complementary manner, for example, in the form of a dialogue, whereby the voice inputs 11 of the occupants 10, 10a, 10b can alternate: for example, a first occupant 10a makes a voice input 11, followed by a second occupant 10b.

[0030] The data processing device 51 is configured to determine semantic speech content 60 of the speech inputs 11, wherein each of the speech contents 60 corresponds to one of the speech inputs 11. In other words, speech content 60 is extracted from each speech input 11 in order to be able to perform a function 80 based on the speech content 60. The speech inputs 11 are evaluated for content, for example, by a computer-assisted method, in particular supported by machine learning and / or artificial intelligence.

[0031] The data processing device 51 is configured to determine assignment information 65, wherein the assignment information 65 comprises an assignment between each of the speech contents 60 and one of the occupants 10, 10a, 10b. Thus, each of the speech contents 60 can be precisely and reliably assigned to one of the occupants 10, 10a, 10b, for example, by a computer-assisted method, in particular by machine learning and / or artificial intelligence. For example, the voices of the occupants 10, 10a, 10b can be recognized. Optionally, further, optionally non-acoustic information can be processed to determine the assignment information 65.

[0032] In other words, the data processing device 51 is configured to identify the occupants 10, 10a, 10b based on their voices and to separate the spoken texts or speech contents 60 by AI using voice recognition.

[0033] The data processing device 51 is configured to determine a location-independent prioritization 70 of the voice inputs 11, taking into account the voice content 60 and the assignment information 65. The prioritization 70 weights the voice content 60 according to the assignment information 65 in order to be able to execute a function 80 based on the voice content 60. The prioritization 70 is static, i.e., permanently defined for one or more of the occupants 10, 10a, 10b, and / or context-related, i.e., defined for one or more of the occupants 10, 10a, 10b depending on the voice content 60. The prioritization 70 can be defined by a voice input 11 from one of the occupants 10, 10a, 10b, in particular by the driver. The prioritization 70 and / or features for determining the prioritization 70 can be stored in the motor vehicle 50 and / or in a cloud in a retrievable manner. In other words, the prioritization 70 comprises a comparison of the texts with defined rules using a language model.

[0034] The data processing device 51 is configured to classify the speech content 60. For this purpose, the speech content 60 is semantically analyzed and divided into classes or groups. The prioritization 70 depends on the classification 126 or the classes. This allows the respective conversation content to be classified, for example, into one of the classes "Navigation," "Audio," etc., particularly using a language model (large language model, Ilm).

[0035] The functional device 56 is configured to perform a function 80 or a vehicle function. The function 80 is therefore a function of the motor vehicle 50. The function 80 can comprise data input and / or actuation of a component of the motor vehicle 50. The functional device 56 is, for example, a user interface of the motor vehicle 50. The data processing device 51 and the functional device 56 are communicatively coupled to one another so that the data processing device 51 can transmit an execution signal 75 to the functional device 56. The execution signal 75 thus serves to execute the voice commands or voice inputs 11.

[0036] The data processing device 51 is configured to output the execution signal 75 for executing the function 80 based on the voice inputs 11, taking into account the prioritization 70, and to transmit it to the functional device 56. For this purpose, the data processing device 51 is configured to determine the execution signal 75, taking into account the prioritization 70 and the voice content 60. The function 80 is based on voice inputs 11 from several of the occupants 10, 10a, 10b, as schematically illustrated by arrows. The voice content 60 of a voice input 11 from a first of the occupants 10, 10a at least partially replaces the voice content 60 of a voice input 11 from a second of the occupants 10, 10b.

[0037] The data processing device 51 is configured to check the prioritization 70 and / or the function 80.

[0038] An example dialogue with occupants 10, 10a, 10b named P1, P2, and P3 could be: P1 says "Drive to Maxdiamantstrasse 143," P3 says: "Diamant what?", P2 says: "183," P1 says "Ah, ok, 183." P3's voice input 11 is ignored and has a low priority. The voice inputs 11 from P1 and P2 are given high priority and synthesized into the execution signal 75 to set "Drive to Maxdiamantstrasse 183" as function 80. In an example dialogue with the same occupants 10, 10a, 10b: P3 says "Drive to Paul at Steinstrasse 17," a function 80 may be omitted because P3 lacks the authority to influence the navigation, meaning his voice command 11 is given a low priority. In an example dialogue with the same occupants 10, 10a, 10b: P3 says “Go to Paul at Steinstrasse 17”.If P1 says "OK, let's go to Paul's," the voice inputs 11 from P1 and P3 are synthesized to set "Drive to Steinstrasse 17" as function 80, since P1 has increased the prioritization 70 of P3 through his voice input 11 by confirming P3's voice input 11. For person P1, it is specified that this occupant 10, 10a, 10b always has the final say on navigation topics. It therefore does not matter where this occupant 10, 10a, 10b is sitting and / or whether he or she is driving or not. If the occupant 10, 10a, 10b is in the motor vehicle 50, the voice input 11 of the occupant 10, 10a, 10b is always processed with a high weighting in a dialogue with the other occupants 10, 10a, 10b. P1, P2, and / or P3 could also be a child. If no context is specified for occupants 10, 10a, and 10b, a simple classification is applied: "Child" means "no or low priority 70" for navigation topics compared to adults.

[0039] Fig. 2 schematically shows a flow diagram of a method 100 according to one aspect of the disclosure. The method 100 according to Fig. 2 is a method 100 for a motor vehicle 50 for processing voice inputs 11 of several occupants 10, 10a, 10b of the motor vehicle 50. Such a motor vehicle 50 is described with reference to Fig. 1 described. Fig. 2 is with reference to Fig. 1 described.

[0040] The method 100 comprises: capturing 110 the voice inputs 11, wherein at least two of the voice inputs 11 are spoken by different occupants 10, 10a, 10b.

[0041] The method 100 comprises: determining 120 semantic speech contents 60, wherein each of the speech contents 60 corresponds to one of the speech inputs 11.

[0042] The method 100 comprises: determining 125 association information 65, wherein the association information 65 comprises an association between each of the speech contents 60 and one of the occupants 10, 10a, 10b.

[0043] The method 100 comprises: classifying 126 the speech contents 60.

[0044] The method 100 includes determining 130 a location-independent prioritization 70 of the voice inputs 11, taking into account the voice content 60 and the assignment information 65. The prioritization 70 is defined statically and / or context-related. The prioritization 70 can be defined by a voice input 11 from one of the occupants 10, 10a, 10b. The prioritization 70 depends on the classification 126.

[0045] The method 100 comprises: checking 135 the prioritization 70.

[0046] The method 100 comprises: outputting 140 an execution signal 75 for executing a function 80 based on the voice inputs 11, taking into account the prioritization 70. The function 80 is based on voice inputs 11 of several of the occupants 10, 10a, 10b. The voice content 60 of a voice input 11 of a first of the occupants 10, 10a at least partially replaces the voice content 60 of a voice input 11 of a second of the occupants 10, 10b.

[0047] Alternatively or in addition to checking 135 the prioritization 70, the method 100 comprises: checking 135 the function 80.

[0048] The person skilled in the art will recognize that the method 100 according to Fig. 2 can also be performed in a different order than that shown. In particular, it is possible for steps of method 100 to be interchanged, shifted, and / or performed simultaneously.

[0049] Fig. 3 shows a schematic representation of a computer program and / or computer-readable medium 200 according to one aspect of the disclosure. The computer program and / or computer-readable medium 200 includes instructions (not shown) which, when the program or instructions are executed by a data processing device 51, cause the data processing device 51 to execute the method 100 and / or the steps of the method 100 according to Fig. 2 to be carried out.

[0050] The instructions can be present as program code in any code or language, in particular in code suitable for controlling and / or monitoring motor vehicles 50. The computer program and / or computer-readable medium 200 can be or include any digital data storage device, such as a USB stick, a hard drive, a CD-ROM, an SD card, or an SSD card. The computer program does not necessarily have to be stored on such a computer-readable storage medium, but can also be accessible via the Internet or otherwise. Reference symbol (part of the description) 10 inmates 10a first occupant, occupant 10b second occupant, occupant 11 Voice input 50 motor vehicles 51 Data processing device 55 Input device 56 Functional device 60 language content 65 Assignment information 70 Prioritization 75 execution signals 80 Function 100 procedures 110 Capture 120 Determining language content 125 Determining mapping information 126 Classify 130 Determining a Prioritization 135 Check 140 Issues 200 Computer program and / or computer-readable medium QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature

[0000] US 2006 / 0212291 A1

[0004] DE 10 2014 114 604 A1 [0005, 0006]

Claims

[1] Method (100) for a motor vehicle (50) for processing voice inputs (11) of a plurality of occupants (10, 10a, 10b) of the motor vehicle (50); the method (100) comprising: - detecting (110) the voice inputs (11), wherein at least two of the voice inputs (11) are spoken by different occupants (10, 10a, 10b); - determining (120) semantic speech contents (60), wherein each of the speech contents (60) corresponds to one of the speech inputs (11); - determining (125) association information (65), wherein the association information (65) comprises an association between each of the speech contents (60) and one of the occupants (10, 10a, 10b); - determining (130) a location-independent prioritization (70) of the speech inputs (11) taking into account the speech content (60) and the assignment information (65); and - Outputting (140) an execution signal (75) for executing a function (80) based on the speech inputs (11) taking into account the prioritization (70). [2] Method (100) according to claim 1, wherein the prioritization (70) is defined statically and / or contextually. [3] The method (100) of claim 1 or 2, wherein the method (100) comprises: - Classifying (126) the language content (60); and - the prioritization (70) depends on the classification (126). [4] Method (100) according to one of the preceding claims, wherein the prioritization (70) can be defined by a voice input (11) of one of the occupants (10, 10a, 10b). [5] Method (100) according to one of the preceding claims, wherein the method (100) comprises: - Checking (135) the prioritization (70) and / or the function (80). [6] Method (100) according to one of the preceding claims, wherein the function (80) is based on voice inputs (11) of several of the occupants (10, 10a, 10b). [7] Method (100) according to claim 6, wherein the speech content (60) of a speech input (11) of a first of the occupants (10, 10a) at least partially replaces the speech content (60) of a speech input (11) of a second of the occupants (10, 10b). [8] Computer program and / or computer-readable medium (200) comprising instructions which, when the program or instructions are executed by a data processing device (51), cause the device (51) to carry out the method (100) and / or the steps of the method (100) according to one of claims 1 to 7. [9] Data processing device (51) for a motor vehicle (50), wherein the data processing device (51) is configured to carry out the method (100) according to one of claims 1 to 7. [10] Motor vehicle (50) comprising the data processing device (51) according to claim 9.

Citation Information

Patent Citations

  • Method and device for processing multiple audio streams in an on-board computer system of a vehicle

    DE102014114604A1

  • Speech recognition system, speech recognition method and storage medium

    US20060212291A1

  • Systems and Methods for Speech Command Processing

    US20130018659A1

  • Method for processing voice signals of multiple speakers, and electronic device according thereto

    US20230040938A1