Method for operating a speech dialogue system and speech dialogue system
The speech dialogue system addresses the challenge of handling unfamiliar users and limited attention by combining non-target and target-guided analyses, ensuring intuitive and comprehensive device interaction through context-aware and personalized responses.
Patent Information
- Application Number
- DE102019217751
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2019-11-18
- Publication Date
- 2025-08-21
- Estimated Expiration
- 2039-11-18
AI Technical Summary
Existing speech dialogue systems face challenges in handling natural language inputs, particularly when users are unfamiliar with voice commands or available options, and struggle to balance conversational and task-oriented interactions, especially in contexts like vehicle operation where user attention is limited.
A speech dialogue system that combines non-target-guided and target-guided dialogue analyses, using machine learning and rule-based approaches to generate responses based on relevance probabilities, dynamically adjusting to context and user feedback, and integrating environmental and personal data for personalized interactions.
Enables intuitive and comprehensive access to device functions, supporting both conversational and task-oriented tasks, enhancing user experience and safety by adapting to user familiarity and context, and providing personalized support.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to a method for operating a speech dialogue system and a speech dialogue system, in particular in a vehicle.
[0002] Speech dialogue systems can be used in a variety of contexts to enable particularly simple operation of electronic devices. The user does not have to operate any physical controls, but can activate or set functions, enter inputs, or perform communication tasks using spoken utterances. However, inputs in natural language often pose problems for existing systems, for example, when the user is unfamiliar with the correct voice command or is unaware of the available control and input options. Furthermore, not all speech recognition and processing approaches are equally suitable for all tasks, such as conducting a conversation with the user and operating specific electronic devices.
[0003] US 2016 / 0071518 A1 discloses a speech recognition system that determines an intention for a user's utterances and selects a suitable recognition engine based on that intention. A search is performed, and the results are presented to the user.
[0004] DE 102017 115 936 A1 describes a method in which a voice assistant is activated when the context of the speech indicates that audible voice assistance is appropriate. Several words are used as input arguments to determine and output additional information.
[0005] In the speech dialogue system described in US 2018 / 0090132 A1, a multitude of different dialogue scenarios are stored, and a dialogue text is generated to respond to a user's utterance. This check checks whether the user's response to an initial utterance from the system corresponds to an expected response, and if necessary, a suitable further utterance from the system is output as a response. If the user's utterance does not correspond to an expected response, a new dialogue scenario is selected that corresponds to the content of the utterance.
[0006] EP 1 346 556 B1 relates to a dialogue system for human-machine interaction, through which users can communicate with various dialogue experts.
[0007] DE 10 2013 222 757 A1 generally relates to speech systems, and more specifically to methods and systems for adapting components of the speech systems based on data determined from user interactions and / or from one or more systems, for example a vehicle.
[0008] DE 10 2013 219 649 A1 relates to a method for creating or supplementing a user-specific language model in a local data storage device connectable to a terminal device, wherein the language model is configured to assign control commands for controlling the terminal device to natural language utterances of a user. The invention further relates to a system with which the method can be implemented.
[0009] The present invention is based on the object of providing a speech dialogue system and a method for its operation, whereby the user can carry out a speech operation in the simplest and most intuitive way possible.
[0010] According to the invention, this object is achieved by a method having the features of claim 1 and a speech dialogue system having the features of claim 9. Advantageous embodiments and further developments emerge from the dependent claims.
[0011] In the method according to the invention for operating a speech dialogue system, a speech input is recorded. A first response output is generated based on the speech input using a non-target-guided dialogue analysis, and a second response output is generated based on the speech input using a target-guided dialogue analysis. A first relevance probability for the first response output and a second relevance probability for the second response output are determined. A speech output is generated based on the response output with the highest relevance probability.In a further training, the first relevance probability for the first answer output and the second relevance probability for the second answer output are compared with a relevance threshold, whereby only answer outputs with a relevance probability above the relevance threshold are taken into account for the generation of the speech output and whereby the relevance threshold is generated dynamically and is determined to be higher the higher a criticality determined for a context is.
[0012] This advantageously allows for selecting between response outputs generated using different analysis methods and providing the optimal speech output depending on the respective context. The speech dialogue system can thus implement both a natural language, non-goal-guided conversation with the user and a goal-guided dialogue for controlling a device.
[0013] A user, such as the driver of a vehicle, is often not familiar with all of the operable functions and applications, such as the functions of a vehicle or other devices such as smartphones or an external computer system. This is due, for example, to the wealth of available functions. This is particularly important when using rental vehicles with which the user is unfamiliar. In addition, only limited attention is available for operation, especially while driving the vehicle, which makes it difficult to access and use various functions. The method now makes it possible to provide the user with a small talk function that supports and informs them and is also perceived as useful in terms of content, while at the same time allowing comprehensive use of existing functions, such as those of a vehicle or an infotainment system, and other functions.
[0014] In particular, the speech dialog system is not specifically invoked to operate a specific application; rather, the context is automatically recognized and a suitable system for dialog analysis is selected. This means that the user does not need to be familiar with the desired application and its available functions from the outset. Furthermore, the system can be used flexibly in a wide variety of situations.
[0015] The recorded speech input is, in particular, natural language, meaning it is not limited to predefined commands or keywords, but a user can enter freely formulated inputs. Recording occurs in a conventional manner, in particular using a microphone. The speech input is converted into text or another data form that can be automatically processed by a computer system. Known methods for converting speech into text (speech to text, SST) can be used for this purpose. Conversely, the speech output is generated in an acoustically perceptible manner, using known methods for converting text into speech (text to speech, TTS).
[0016] In a “dialogue analysis” within the meaning of the invention, at least one response output or a series of candidates for a response output is generated based on at least one speech input and a dialogue history.
[0017] In particular, non-goal-guided dialogue analysis does not assume a dialogue that follows a predefined sequence of user inputs and clearly assigned system responses, for example, to configure a functionality as the goal of the interaction. Instead, a small-talk functionality is executed, for example, in which the speech output is generated as a response from the system in such a way that an ongoing conversation with the user is continued. Such a non-goal-guided analysis can be performed, for example, using a data-driven system in which the response output is generated based on training data from previous dialogues.
[0018] In an embodiment according to the invention, the first response output in the non-target-guided dialogue analysis is generated using a machine learning system. In particular, the machine learning system comprises a deep neural network (DNN). This advantageously allows the system to be trained particularly comprehensively and using a large number of previously known dialogue histories and continuously newly acquired data.
[0019] In a further development, the machine learning system accesses a personalized preferences database to generate the initial response output. This database is updated based on a dialogue history. This advantageously allows the machine learning system to be adapted particularly flexibly to a user, specific contexts, and requirements.
[0020] The preferences database is created, for example, based on recorded data about the user's acceptance of speech output. For example, the user can indicate through an input that a speech output from the system is not relevant to their voice input or that they would like to steer the dialogue to a different topic. From this, it can then be determined that the output is not relevant to the user, and negative feedback can be stored in the preferences database. Conversely, positive feedback can be stored if the user accepts or confirms the speech output.
[0021] The dialogue history considered in the dialogue analysis is, in particular, formed in such a way that it encompasses the course of a current conversation. The history can, for example, comprise a series of consecutively recorded voice inputs and response outputs, or it can comprise the voice inputs recorded within a specific period of time and the response outputs generated for them. The dialogue history is, in particular, stored in a database, which can be accessed, for example, to generate the first and / or second response output.
[0022] The dialogue history may further include additional information about a conversation context. For example, this may be information about the operating status of a vehicle, the current traffic situation, the position, and / or the geographical surroundings of the vehicle. The additional information may also include current information or be related to the past and future, such as a user's appointments or data on the use of a communication device, or information about other people in the vicinity of a user, for example, in the same vehicle.
[0023] Goal-guided or rule-based dialogue analysis is based on a "scripted," deterministically predefined dialogue flow. This means that a speech input is assigned to a dialogue state stored in a database, such as a step in operating a device. A response output is then assigned to this dialogue state. This type of dialogue analysis is therefore suitable, for example, for guiding a user through a specific operation, capturing an input, or executing another predefined dialogue flow.
[0024] In a training course, goal-guided dialogue analysis uses a rule-based expert system to generate operating instructions for a function controlled by the speech dialogue system. The dialogue can thus be advantageously used to capture a specific action instruction in the system.
[0025] In this case, an operating instruction corresponds to the goal of goal-guided dialogue analysis, especially a scripted dialogue, for example, to operate a specific device or to capture a specific input. A predefined knowledge base is used to determine a relevant response output.
[0026] The expert system can be implemented locally in the method, for example, with an integrated storage unit for the speech dialogue system to store a corresponding database. The speech dialogue system can also access an external unit with the expert system, for example, via a computer network such as the Internet or through a connection to an external unit, such as a mobile user device. In this way, different expert systems can be integrated modularly.
[0027] The instruction specifically relates to a unit that is data-linked to the speech dialog system, to which the operating instruction is subsequently transmitted. For example, the speech dialog system is integrated into a vehicle, and the instruction is generated for another device in the vehicle. Such devices can, for example, relate to settings for the vehicle's driving characteristics, an infotainment system, or a telecommunications device.
[0028] In a design for generating the second response output, the speech input is assigned to a dialog scenario, where the dialog scenario includes an input intention and an output response predefined for the input intention. This advantageously allows a scripted dialog to be executed particularly efficiently.
[0029] When defining a dialog scenario, the context of the input is determined, taking into account the dialog history and other data about the state of the speech dialog system and the respective environment. The dialog scenario describes a potential dialog flow with system responses associated with specific inputs and potentially expected further user inputs. The input intention is understood as the operation of a specific unit, the invocation of a function, or the execution of a task to provide specific information. This means that the input intention specifies the goal of the goal-guided dialog analysis.
[0030] In particular, the process continuously checks the relevance of the last determined input intention. If it is determined that a different intention appears more relevant, the dialog scenario is redefined accordingly. This allows for a reaction if the user's input intention changes during the dialog or if the user indicates that an incorrect input intention was detected.
[0031] In a further embodiment, environmental data of a user are recorded, and the first and / or second response output is further generated based on the recorded environmental data of the user. This advantageously allows a particularly relevant response output to be generated.
[0032] The environmental data can include, for example, the operating status of a vehicle or other device, a current, past, or predicted position and characteristics associated with that position, as well as personal data of a user. They can also relate to user-related information, which can be used to recognize a user's state and / or emotions. The first and / or second response output can therefore also be generated based on the recorded information about a user's state and / or emotions. They can also relate to information determined and / or stored in a spatial environment of the user. The spatial environment can be viewed for a current, past, or future point in time.
[0033] Based on the environmental data, a context for the speech input is determined and used in either non-goal-guided or goal-guided dialogue analysis. For example, the set of potentially considered response outputs can be limited based on the environmental data, for example, because the linguistic relationship between the speech input can be clarified based on the environmental data.
[0034] The first and second relevance probabilities are determined in a conventional manner. The probabilities are determined, in particular, through a statistical analysis of the generated first and second response outputs and output together with the respective response output. The relevance probability indicates the probability that a response output represents a relevant reaction for the user to the previously recorded speech input. Different methods can be used, for example, to analyze relationships between the response outputs and the speech input and to determine a context and input intention based on the speech input and the dialogue history. Environmental data can also be taken into account when determining the relevance probabilities.
[0035] The relevance threshold can be fixed. It can also vary depending on the method used to generate the response outputs. For example, different relevance thresholds can be provided if response outputs are generated using a neural network or an expert system. The relevance threshold can also be generated dynamically and, for example, be set higher the higher the criticality determined for the context. For example, in the context of a vehicle, it can be determined that a higher criticality exists during high traffic volumes than during low traffic volumes, and the threshold can be set higher in this case to avoid unnecessary distraction from less relevant responses. For example, the environmental data can be taken into account to determine the relevance probability.
[0036] During training, an entity encompassed by the speech input is identified, and the first and second relevance probabilities are determined based on the identified entity. The entities can be identified using conventional methods, such as a Named Entity Resolver or a Named Entity Recognizer (NER). Based on these entities, particularly relevant response outputs can be advantageously generated, tailored to the respective context.
[0037] In the context of the invention, "entities" are understood to mean, in particular, linguistic objects that contain collected information. For example, attributes and predicates can be assigned to the entities to further determine the content of the speech input. In this way, they are used to generate response outputs and provide information about both the context of the speech input and the content or input intention of the user.
[0038] The specific entities are therefore central to the decision as to which of the generated response outputs is most relevant to the user. In particular, the entities are used to distinguish whether a response output from the non-goal-guided dialogue analysis or a response output from the goal-guided dialogue analysis should be generated.
[0039] In particular, entities related to functionalities that can be operated by the voice dialogue system, such as vehicle and infotainment functionalities in a vehicle, are trained and learned by a machine learning system and stored in a database as a so-called knowledge base. The use of operable functionalities is particularly personalized and linked to a user. Entities used in a current dialogue are generated, for example, by a machine learning system by evaluating voice dialogues and environmental data, such as data recorded by a vehicle about the driving situation, the interior or surroundings of the vehicle, a traffic situation, and recorded information about the driver's condition. In particular, the driver's personal interests and / or emotions specific to the user are also evaluated in order to be able to react flexibly to the driver's current mood.
[0040] The information about the user's state, which may be included in the environmental data, can itself comprise various data and is recorded and evaluated, in particular, during user emotion and user state recognition. This can include evaluating movement sequences, facial expressions, and gestures of the driver or another user. Furthermore, speech parameters can be analyzed, such as voice parameters, speaking rate, volume, phrases used, or conversation dynamics. Furthermore, physiological parameters or vital signs of the user of the speech dialogue system can be taken into account. Driver emotion and driver state recognition are particularly provided for this purpose in a vehicle. Furthermore, a smartphone application can be used to determine and classify information about the user's state, such as movement sequences, movement patterns, gestures, and facial expressions.Furthermore, the user's physiological or vital signs can be recorded via sensors integrated into clothing or wearable measuring devices, for example, and located on or near the user's body. The data thus recorded can then be read out, for example, via a smartphone application, and evaluated and used for the voice dialogue system.
[0041] The speech dialogue system according to the invention comprises a detection unit which is configured to detect a speech input, a first dialogue analysis unit which is configured to generate a first response output based on the speech input by means of a non-target-guided dialogue analysis, and a second dialogue analysis unit which is configured to generate a second response output based on the speech input by means of a target-guided dialogue analysis.It further comprises a control unit configured to determine a first relevance probability for the first response output and a second relevance probability for the second response output, and an output unit configured to generate a voice output based on the response output with the highest relevance probability, the control unit configured to compare the first relevance probability for the first response output and the second relevance probability for the second response output with a relevance threshold, wherein only response outputs with a relevance probability above the relevance threshold are taken into account for generating the voice output, and wherein the relevance threshold is generated dynamically and is determined to be higher the higher a criticality determined for a context is.The speech dialogue system according to the invention is particularly designed to implement the above-described method according to the invention. The speech dialogue system thus has the same advantages as the method according to the invention.
[0042] In one embodiment of the speech dialog system according to the invention, the second dialog analysis unit is configured to generate, based on a rule-based expert system, an operating instruction for a functionality controllable by the speech dialog system. This advantageously generates the response output based on an existing knowledge base. Furthermore, expert systems can be integrated modularly into the speech dialog system or provided by an external unit via a data connection.
[0043] A voice interaction can be used to call up functionalities for which the user already knows predefined terms, expressions and words. This is done in particular via a goal-guided dialogue analysis, especially a scripted dialogue. Furthermore, a non-goal-guided dialogue analysis can be used to call up, activate and use functionalities and applications that are unknown to them or that they must first be made aware of. For this purpose, for example, a small talk application is used in which recommended functionalities are determined based on data about a context or situation. Using the small talk application and, if necessary, situation recognition, it can be determined how the user would like to be supported at the moment or what relevant support can be offered, whereby less relevant functionalities are not offered.A Small Talk application allows for more targeted and faster access to functions and applications, without requiring the user to memorize the voice commands. Functions and applications commonly used by users in a specific context can be offered in a targeted manner.
[0044] This provides extensive support for the use of operable functions, such as those in a vehicle and for operating an infotainment system, and intuitive access to a comprehensive range of functions and applications is provided via the voice dialog system. For example, vehicle maintenance or the use of infrastructure, such as parking spaces, can also be intuitively supported.
[0045] Furthermore, personalized learning can be provided for different users regarding which specific functionalities should be used in specific situations. Of all the functionalities and applications available, for example, in a vehicle, the most relevant ones are determined and used based on the context and any data collected about the environment.
[0046] Furthermore, the speech dialogue system and non-goal-guided dialogue analysis can be used to implement "free speech" or "small talk" without scripted dialogue states. However, such states can also be used to perform targeted tasks for a user, so that the speech dialogue system also provides for goal-guided dialogue analysis, particularly in parallel. The decision as to which dialogue analysis will be used to generate the response output is made based on relevance probabilities, which are determined primarily by entities. This means that the various dialogue analyses of the speech dialogue system are linked using entities, which are used to determine the respective context and the current dialogue content.
[0047] The process makes it possible to combine a small talk application with a rule-based analysis of the dialogue in such a way that different functionalities, such as those of a vehicle and external devices, can be operated in a uniform operating concept.
[0048] Goal-guided dialogue analyses are used to perform specific tasks, for example with speech inputs such as "What's the weather like?" or "Buy me two tickets!". Scripted dialogue states are used in particular, where the speech input is assigned to a dialogue scenario for which specific response tasks are predefined. In contrast, non-goal-guided dialogue analyses are used to enable free conversation with the speech dialogue system, also known as "small talk." Scripted dialogue states are dispensed with here, and relevant response outputs are generated using data-driven machine learning systems, such as a deep neural network (DNN). Here, the content of an individual speech input is typically not determined; instead, the response output is calculated using statistical methods.The procedure combines the two approaches of dialogue analysis and switches between them depending on the application.
[0049] The two approaches are linked, in particular, by entities that are learned, for example, by a deep neural network in a training phase. The available entities are stored in a database with associated information. When a response output is generated by the DNN end, a dialogue manager can determine the dialogue content and, in particular, the user's input intention based on the entities used and generated in the speech input. If a task is identified that goes beyond small talk, a scripted dialogue can be conducted using goal-guided dialogue analysis until the task is completed.
[0050] The method can also personalize the speech dialogue system. Using a reinforcement approach, personalized user information is generated and used to expand or adapt the system for the user. For example, if a speech output is generated based on a response that is irrelevant to the user, the system can detect this through a change of topic or other feedback from the user. Negative feedback is used through reinforcement learning to avoid such irrelevant output in future dialogues of a similar nature.
[0051] In addition, other suitable entities can be recognized and assigned for which a small talk function or an entity model of an empathic assistant already exists. An entity model assigns matching entities to one another. For example, by merging data from various environmental sensors or other sensors, e.g. for situation recognition in a vehicle interior, objects and situations are recognized that can be assigned to a software object and / or a term, i.e. a possible entity, in different systems. In particular, parameters set in vehicle and infotainment systems as well as features and properties of current media and app usage in the vehicle are also stored. Attributes for individual objects (entities) can be assigned and stored in the vehicle and infotainment system.
[0052] Examples of object or situation recognition include the recognition of a building or geographical surroundings, the classification of buildings or other classifiable facilities, such as a theater, opera house, cinema, hotel, school, town hall, swimming pool, hospital, doctor's office and therapy facility, restaurant, bus stop, train station, parking lot or similar. Information about points of interest (POI) of a navigation system can also be taken into account. Furthermore, a traffic situation, route, area or immediate surroundings can be recorded and taken into account. In addition, features and properties of the current use of media or apps, such as volume, a music or media title, a radio station, an app or functionality used, a short description of a content title can be used.A navigation system can provide POls arranged in a geographical area as well as POls of individual interest to the user.
[0053] These objects and / or their features and attributes can be stored and continuously updated in a Small Talk database (entity data store), for example, in an external backend or in the vehicle. The Small Talk database for storing attributes of data units can be used additionally or connected. During Small Talk, recognized entities from the conversation are related to entities in the (vehicle) Small Talk database, and the Small Talk can thus be steered in a direction that informs or supports the driver.
[0054] By analyzing the dialogue spoken by the driver and / or other vehicle occupants, entities previously stored in the Small Talk system can be found. Entities can be defined via a database, for example, in data modeling using an entity-relationship model, for knowledge-based systems (knowledge base) and related to each other using semantic entity models, so that further appropriate questions, offers, suggestions, information, and / or answers that further specify the topic can be generated and output by the Small Talk system as response outputs.
[0055] The interrelated entities are stored in entity models and can be used to further refine thematic content for small talk conversations. Additional entities stored in entity models are associated, particularly based on entities already recognized in the speech dialogue. The data model, in which the entities and their relationships are defined, can incorporate predictions of experience-based linguistic relationships, which were previously collected statically from data analyses. For example, certain terms frequently occur in combination with others. The entity model can therefore be expanded and personalized for a specific person through self-learning.
[0056] In this context, offers relating to already existing applications and services offered, for example, by the vehicle system can also be used, such as organizational applications, the possibility of reserving cinema or theater tickets, ordering options, for example, for food, music, news, information, advice, and so on.
[0057] Furthermore, this type of small talk tailored to the driver can be used to increase driving safety, recognize the traffic situation or surroundings, or introduce the driver to new vehicle functions and respond to their current emotions. For example, if the vehicle system detects driver emotions such as annoyance, rage, sadness, or impatience, small talk expanded to reflect the driver's current interests can be used to ask the driver specific questions from the vehicle's applications and services in order to positively influence the driver's state, such as desired orders and calling up the infotainment system with music, news, information, or specific advice. This allows emotion recognition to be incorporated into small talk with an empathetic assistant.
[0058] The invention will now be explained using embodiments with reference to the drawings. Fig. 1 shows a vehicle with an embodiment of the speech dialogue system according to the invention and Fig. 2 shows a detailed view of the embodiment of the speech dialogue system according to the invention.
[0059] With reference to Fig. 1, a vehicle with an embodiment of the speech dialogue system according to the invention is explained.
[0060] The vehicle 1 comprises a control unit 3, to which a detection unit 2 and an output unit 6 are coupled. The control unit 3 comprises a first dialogue analysis unit 4 and a second dialogue analysis unit 5.
[0061] The recording unit 2 is designed in a manner known per se and includes, in particular, a microphone. The recording of vocal utterances of a user, in particular a driver of the vehicle 1, occurs continuously in the exemplary embodiment. In the example, recorded utterances are stored in a ring buffer such that only utterances within a past time interval of a predetermined length are stored, and older utterances are deleted. Only when it is detected that the user wishes to engage in a dialogue with the speech dialogue system are utterances stored over a longer period of time.In further embodiments, the recording and storage takes place in response to an input from the user, for example triggered by actuating an input element or by a signal from a device that can be operated by voice control, which prompts the user to enter voice commands and activates the voice dialogue system to record the voice commands.
[0062] The output unit 5 is also designed in a manner known per se and, in particular, comprises a loudspeaker. In the exemplary embodiment, it is integrated into an infotainment system of the vehicle 1, which implements various functions for media playback, the operation of communication devices, and the output of messages from driver assistance systems, as well as the associated recording of user inputs.
[0063] With reference to Fig. 2, an embodiment of the method according to the invention is explained. In this case, the method described above with reference to Fig.1 explained embodiment of the speech dialogue system according to the invention, which is further specified by the description of the method.
[0064] In the exemplary embodiment of the method, a user's voice input is first captured by the capture unit 2 and converted into machine-readable text in a first step S1. Known conversion methods (speech-to-text, STT) are used. In a second step S2, the text thus generated is transmitted to a natural language understanding unit (NLU) and processed there. In particular, the processing is carried out by a named entity recognizer (NER), which establishes a data connection to a first database DB1 for this purpose.
[0065] The first database DB1 stores general and domain-specific knowledge (knowledge base), which specifically represents the semantic relationships between different entities. Such relationships can be represented, for example, as follows: (Bill Gates; born in; Seattle), (Seattle; located in; USA). Entities identified in step S2 are then used to capture associated information and recognize the context of a conversation between the user and the speech dialog system.
[0066] Subsequently, response outputs matching the speech input are determined. This is done by the first 4 and second dialogue analysis unit 5. The first dialogue analysis unit 4 processes the speech input and the entities determined therein in a step S4 using a deep neural network (DNN), which uses statistical methods to determine a first response output or a series of potential first response outputs. For this purpose, the DNN also accesses another database DB3, which stores a dialogue history between the user and the system. A context for the current speech input can therefore be determined from the data stored there. A non-goal-guided dialogue analysis is carried out, i.e., a goal towards which the dialogue is directed is not determined, for example, obtaining specific information as user input.Rather, the dialogue analysis and the generation of the first response output serve to continue the dialogue until it is recognized that a specific operating action should be carried out and a targeted dialogue analysis is more relevant.
[0067] The DNN also accesses another database, DB2, in which personalized information about the user's behavior, interests, and preferences is stored. This means that the exemplary embodiment assumes that the user's identity is known. This identity is determined using conventional methods, in particular by means of an input, a personal vehicle key, or a mobile user device.
[0068] The second dialogue analysis unit 5 comprises a script unit, by means of which a second response output is generated in a step S5. Here, a goal-guided dialogue analysis is carried out using a rule-based expert system, in which the voice input is assigned to a predetermined current dialogue state, for which in turn a specific second response output is defined. For this purpose, the script unit also accesses the dialogue history stored in the database DB3. In particular, an input intention is first determined, for example a specific operable function of a device of the vehicle 1, for which an operating action is to be recorded with a specific probability in the current context. Such an operating action then represents the goal of the operation by means of the second dialogue analysis unit 5, wherein in this case the dialogue is conducted in such a way that the necessary inputs from the user for the operating action are recorded.
[0069] In a further embodiment, the script unit also accesses the DB2 database with personalized information and uses the data about the user stored there to more precisely determine, for example, the input intention or the current dialog state.
[0070] In a further exemplary embodiment, environmental data is also recorded. This includes, for example, information about a state of the vehicle 1, its driving operation, or a situation in its interior. Furthermore, information about the position of the vehicle can be recorded, both regarding the current time and previous times or positions on a planned route. Based on the position information, further data can be recorded, for example about points of interest in a surrounding area, buildings and facilities, restaurants and shops, or the like. Furthermore, stored data and information, such as furnishings in a user's smart home, can also be taken into account. Furthermore, further information about the surroundings of the vehicle 1 or the user can be determined and stored for further determination of the situation.The acquisition of environmental data or other information can be performed not only by sensors of vehicle 1, but also by other information sources, such as a computer network, user inputs, and / or available storage media. The response outputs are then also generated based on this environmental data and / or other information, with a context being determined, for example, based on the available data.
[0071] The generated first and second response outputs are transmitted to a dialogue manager, which decides in step S6 whether the first or second response output should be output. This means that the dialogue manager decides between responses to the speech input generated using different dialogue analyses. For this purpose, relevance probabilities are used, which in the exemplary embodiment were determined by the first dialogue analysis unit 4 and the second dialogue analysis unit 5 when generating the response outputs.
[0072] The relevance probabilities are determined in a manner known per se. For example, the response output or a plurality of potential response outputs is determined by the DNN of the first dialogue analysis unit 4 using statistical methods, and in the process, a probability is also determined with which the first response output is relevant as a response to the voice input. This probability is passed along with the first response output to the dialogue manager. In a similar manner, the script unit of the second dialogue analysis unit 5 determines a relevance probability for the output of the second response output, taking into account, for example, a confidence factor when determining the input intention, and passes this probability to the dialogue manager.
[0073] To decide between the first and second response output, the context of the dialogue being conducted is also taken into account, in particular by accessing the database DB3, which stores the dialogue history, as well as the database DB2 containing personalized information. Based on the data collected there, the context is further determined, and the relevance probabilities for the first and second response output can be specified. Conversely, in step S6, data is also provided to update the databases DB2 and DB3, for example, by supplementing the dialogue history with the final response output and storing a personal preference determined by the user.
[0074] In a further step S7, the response output with the highest relevance probability is converted into spoken language, whereby the text is converted into spoken language (text-to-speech, TTS) in a conventional manner. The output is finally output via output unit 6 in vehicle 1.
[0075] A dialogue between the user and the speech dialogue system known as a “bot” can, for example, proceed as follows: User: "I saw the trailer for the new "Lion King" movie yesterday. What do you think of the film?" Bot: “I like the photorealistic depiction of the animals in the film.” User: “So that means there are no real animals in the film?” Bot: “The entire film was created on the computer.” User: "That sounds impressive. Can you tell me if the film is showing anywhere nearby?"
[0076] Up to this point, the bot's outputs are generated using the neural network of the first dialogue analysis unit 4, which implements a small-talk bot. The responses are generated data-driven, meaning the system generates the most probable responses based on previously learned, previous dialogues.
[0077] At this point in the dialogue, the system recognizes from the user's question that an instruction has been given to the system, namely to search for a nearby cinema showing the film. This function can be performed by a function of the system in vehicle 1. To do this, by identifying the entities in the user's request and in the dialogue history, the entity "the film" is resolved as a reference to the film title "The Lion King" and passed on to the speech dialogue system together with the command to search for a corresponding venue, generally therefore for a POI search. Responses are then output by the second dialogue analysis unit 5, whereby a script-based dialogue is carried out to process a user's request for the POI search and, if necessary, further information is requested from the user.For example, an operating action can now also be recorded from the dialogue by which the navigation system of vehicle 1 is set for a trip to a corresponding cinema and / or an appointment for a visit to the cinema is planned in the user's calendar. List of reference symbols 1 vehicle 2 recording unit 3 Control unit 4 First dialogue analysis unit 5 Second dialogue analysis unit 6 Output unit S1 to S7 step DB1, DB2, DB3 database DM Dialog Manager NER Entity Recognition Unit
Claims
[1] Method for operating a speech dialogue system in which a voice input is recorded; a first response output is generated based on the speech input by means of a non-targeted dialogue analysis; and a second response output is generated based on the speech input by means of a targeted dialogue analysis; a first relevance probability for the first response output and a second relevance probability for the second response output are determined; and a speech output is generated based on the answer output with the highest relevance probability, characterized by , that the first relevance probability for the first answer output and the second relevance probability for the second answer output are compared with a relevance threshold, where only response outputs with a relevance probability above the relevance threshold are taken into account for the generation of the speech output and where the relevance threshold is generated dynamically and is determined higher the higher the criticality determined for a context is. [2] Method according to claim 1, characterized by that the first response output in non-targeted dialogue analysis is generated using a machine learning system. [3] Method according to claim 2, characterized by that the machine learning system accesses a personalized preferences database to generate the first response output, which is updated depending on a dialogue history. [4] Method according to one of the preceding claims, characterized bythat in the targeted dialogue analysis, an operating instruction for a functionality that can be controlled by means of the speech dialogue system is generated using a rule-based expert system. [5] Method according to claim 4, characterized by , that to generate the second response output, the speech input is assigned to a dialogue scenario; where the dialogue scenario includes an input intention and an output response specified for the input intention. [6] Method according to one of the preceding claims, characterized by , that A user’s environmental data is collected and the first and / or second response output is further generated based on the recorded environmental data of the user. [7] Method according to one of the preceding claims, characterized by , that an entity encompassed by the speech input is determined; and the first relevance probability and the second relevance probability are determined depending on the specific entity. [8] Speech dialogue system, comprehensive a detection unit (2) configured to detect a voice input; a first dialogue analysis unit (4) which is configured to generate a first response output based on the speech input by means of a non-target-guided dialogue analysis; a second dialogue analysis unit (5) which is configured to generate a second response output based on the speech input by means of a targeted dialogue analysis; a control unit (3) configured to determine a first relevance probability for the first response output and a second relevance probability for the second response output; and an output unit (6) which is designed to generate a voice output based on the response output with the highest relevance probability, characterized by , that the control unit (3) is configured to compare the first relevance probability for the first response output and the second relevance probability for the second response output with a relevance threshold value, where only response outputs with a relevance probability above the relevance threshold are taken into account for the generation of the speech output and where the relevance threshold is generated dynamically and is determined higher the higher the criticality determined for a context is. [9] Speech dialogue system according to claim 8, characterized bythat the second dialogue analysis unit (5) is designed to generate, on the basis of a rule-based expert system, an operating instruction for a functionality controllable by means of the speech dialogue system.
Citation Information
Patent Citations
Method and system for creating or supplementing a user-specific language model in a local data storage device connectable to an end device.
DE102013219649A1
Adaptation methods and systems for language systems
DE102013222757A1
Dialog system for man-machine interaction, comprising co-operating dialog devices
EP1346556B1