Conversation system and method with improved man-machine conversation concept

By designing input interface, preprocessor and multiple information extraction processors in the dialogue system, and selecting processors related to the current dialogue state, the high memory consumption and computing complexity problems of dialogue systems in the prior art during multi-domain and multi-intention processing are solved, and a more efficient and robust information extraction effect is achieved.

CN120019380APending Publication Date: 2025-05-16FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069545.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing dialogue systems have problems with excessive memory consumption and computational complexity when dealing with multi-domain and multi-intention user input, especially when evaluating multiple classifiers simultaneously.

Method used

A dialogue system is designed, which includes an input interface, a preprocessor and a plurality of information extraction processors. The resulting information is generated from the preprocessing information by selecting the information extraction processor related to the current dialogue state and utilizing information extraction rules specific to each processor.

Benefits of technology

By optimizing the selection and use of information extraction processors, the system reduces memory consumption and computing complexity, and improves the efficiency and robustness of the dialogue system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019380A_ABST
    Figure CN120019380A_ABST
Patent Text Reader

Abstract

A dialog system according to an embodiment is provided. The dialog system comprises an input interface (105) for obtaining an input representation of an input by receiving the input and deriving an input representation from the input, or by receiving the input representation, the input representation being an audio signal representation or a speech representation or a textual representation, where the input representation comprises a plurality of input representation elements. Furthermore, the dialog system comprises a pre-processor (110) for pre-processing the input representation to generate pre-processed information such that the pre-processed information comprises a plurality of pre-processed information elements, and such that each of two or more of the plurality of pre-processed information elements depends on at least two of the plurality of information representation elements. Further, the dialog system comprises two or more information extraction processors (120, 123), where each of the two or more information extraction processors (120, 123) is adapted to extract the dialog information in accordance with the information extraction processor-specific and associated with the two or more information extraction processors (120, 123). The resulting information is generated from the pre-processed information by an information extraction rule that is different from the information extraction rule of any other one of the information extraction rules. Furthermore, the dialog system comprises an output interface (135) for generating outputs from the resulting information from one or more of the two or more information extraction processors (120, 123), the outputs being audio outputs and / or textual outputs and / or visual outputs and / or signals for manipulating the machine.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] manual

[0002] The present invention relates to a dialog system and method with an improved human-computer dialog concept, and in particular to effective information extraction in a dialog system.

[0003] Human-machine interaction (HMI) technology, especially human-computer interaction (HCI) focuses on the design and use of computer technology to provide new interaction interfaces and ways between humans and machines such as computers.

[0004] In this field, when humans communicate with machines, and vice versa, when machines communicate with humans, interfaces for interaction between humans and machines, such as voice interfaces, can be used. In order to realize meaningful dialog systems for human-machine interaction, complex natural language processing concepts can be used.

[0005] Natural language processing (NLP) is a subfield of computer science and artificial intelligence that deals with the interaction between machines (especially computers and humans), and in particular with processing and analyzing large amounts of natural language data. Artificial intelligence can be used to enable computers to understand content, including the contextual nuances of language therein. Analyzing spoken content involves accurately extracting information from speech representations and classifying the information.

[0006] Conversational dialogue systems or short dialogue systems (e.g., voice assistant systems) play an important role in the digitization of industrial processes, home automation, or entertainment applications. Users can interact with dialogue systems using a voice interface. Another example is a chatbot, where users can interact by typing text into the chatbot interface. In goal-oriented dialogue systems (see [4]), users are typically guided by the dialogue system in order to complete use case-specific tasks, such as booking a ticket, starting a phone call, gathering information, or managing a calendar. A common approach to defining the possible interactions of a user with a dialogue system is to specify the states and transitions of the dialogue path, e.g., using a graph-based representation in a tree structure or representing them as blocks in a flow chart.

[0007] Figure 2 A dialog diagram depicting an example of a user interacting with a dialog system is shown, depicting states and transitions of a dialog path. Depending on the state within the dialog, i.e., the interaction with the user, the dialog system needs to perform different behaviors, e.g., provide an appropriate response to the user. In some cases, the dialog system needs to recognize the actual intent of the user based on the user input. For example, this user input can be a voice or voice representation of a user's utterance in a voice assistant system. Or, for example, the user input can be text that has been typed into a chatbot, or a text transcription thereof.

[0008] For example, in a conversational system for a banking application, the system may need to identify whether the user's intention is to check an account balance or to initiate a transfer to another account. The task of mapping user input to a set of predefined intent categories is often called intent classification (see [2]).

[0009] In the following, a distinction is made between so-called local intents and so-called global intents. Local intents are represented as intents that are only relevant in a specific dialog state, i.e. a specific state of the user's interaction with the dialog system. The corresponding intent classifiers are called local intent classifiers. On the other hand, global intents can be triggered by the user at any point during the interaction with the dialog system. Their relevance does not depend on the specific state of the dialog. Examples of global intents include a "stop" command for stopping the dialog system, or a "play music" intent for playing music on a smart speaker at any point during a session with the dialog system. A dialog system may need to recognize both types of intents simultaneously, where the specific local intent to be recognized depends on the actual state of the dialog.

[0010] In some applications, the dialogue system is configured to serve different tasks or domains, which is often referred to as a multi-domain dialogue system (see [5]). For example, the dialogue configurations for these different domains can be represented by a set of parallel disconnected dialogue graphs. Figure 3 Such a dialog configuration is shown. In some cases, subgraphs associated with different domains may be connected to each other, but different entries or starting points of these subgraphs are associated with different domains. Such a multi-domain dialog system uses a so-called domain classifier to decide which available domain the user input involves.

[0011] In some use cases, dialogue systems aim not only to extract information related to the user intent or domain related to a specific user input, but also perform additional classification tasks. Examples include dialogue act classification, i.e., classifying user input according to the type of utterance: for example, whether it represents a question, a confirmation, a rejection, or whether the user is providing information to the system. In some cases, dialogue systems need to extract information included in the input, which can be considered as variables or so-called entities. For example, if the user input is "What is the weather in Berlin?", the dialogue should recognize that the corresponding dialogue act is a question, that the user's intent is to obtain information about the weather, and further extract the term Berlin as an entity related to location information. The extraction of information about entities or variables is often called entity recognition (see [6]).

[0012] More generally, dialogue systems need to solve various information extraction tasks on the received input related to natural language understanding, some examples of which are provided above. These tasks can be summarized as being solved by an information extraction processor.

[0013] In modern dialogue systems, for example, the classification of intent, domain or dialogue act can be performed based on deep neural networks.

[0014] In prior art dialog systems, each of the different intent classifiers, domain classifiers or other classifiers is typically implemented by a separate dedicated deep neural network (DNN) that has been trained with the corresponding training data (including example sentences for each category to be recognized by the classifier). In some cases, part of the neural network parameters are initialized with pre-trained parameters of a separately trained neural network (see [1]), i.e., so-called transfer learning methods are applied (see [7]). The entire neural network is then trained using labeled training data to adapt it to the actual classification task. Since the size of DNNs typically used for intent or domain classification is usually very large, for example, they may include millions of parameters, this approach implies huge demands on memory consumption and computational complexity if several classifiers are evaluated simultaneously for the same user input.

[0015] The object of the invention is to provide an improved concept for a dialog system concept. The object of the invention is achieved by a dialog system according to claim 1, a dialog system according to claim 20, a method according to claim 24, a method according to claim 25 and a computer program according to claim 26.

[0016] A dialogue system according to an embodiment is provided. The dialogue system includes an input interface for obtaining an input representation of the input by receiving the input and deriving an input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation includes a plurality of input representation elements. In addition, the dialogue system includes a preprocessor for preprocessing the input representation to generate preprocessing information, so that the preprocessing information includes a plurality of preprocessing information elements, and so that each of two or more of the plurality of preprocessing information elements depends on at least two of the plurality of information representation elements. In addition, the dialogue system includes two or more information extraction processors, wherein each of the two or more information extraction processors is adapted to generate the obtained information from the preprocessing information according to an information extraction rule that is specific to the information extraction processor and is different from the information extraction rule of any other one of the two or more information extraction processors. In addition, the dialogue system includes an output interface for generating an output based on the obtained information from one or more of the two or more information extraction processors, wherein the output is an audio output and / or a text output and / or a visual output and / or a signal for manipulating a machine.

[0017] According to an embodiment, the dialog system may, for example, be configured to select at least one of the two or more information extraction processors such that only those of the two or more information extraction processors that have been selected and their information extraction rules generate the resulting information.

[0018] In an embodiment, at least two of the two or more information extraction processors may generate derived information from the pre-processed information, for example, according to their information extraction rules.

[0019] In addition, a dialogue system according to another embodiment is provided. The dialogue system includes an input interface for obtaining an input representation of the input by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation includes a plurality of input representation elements. In addition, the dialogue system includes two or more information extraction processors, wherein each of the two or more information extraction processors is adapted to generate the resulting information according to the input representation, according to an information extraction rule specific to the information extraction processor and different from the information extraction rule of any other one of the two or more information extraction processors. In addition, the dialogue system includes an output interface for generating an output according to the resulting information from one or more of the two or more information extraction processors, wherein the output is an audio output and / or a text output and / or a visual output and / or a signal for manipulating a machine. At least two of the two or more information extraction processors are dialogue state dependent. The dialogue system is configured to select one or more of the at least two information extraction processors whose dialogue states are dependent, according to the current state of the dialogue, so that only those information extraction processors associated with the current state of the dialogue among the at least two information extraction processors are selected. The selected one or more information extraction processors are configured to generate the obtained information according to their information extraction rules.

[0020] In addition, a method according to an embodiment is also provided. The method includes:

[0021] - obtaining an input representation of the input via an input interface of the dialog system, wherein the input interface obtains the input representation by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements.

[0022] - preprocessing the input representation by a preprocessor of the dialog system to generate preprocessed information, such that the preprocessed information comprises a plurality of preprocessed information elements, and such that each of two or more of the plurality of preprocessed information elements is dependent on at least two of the plurality of information representation elements, wherein each of the two or more information extraction processors of the dialog system is adapted to generate resulting information from the preprocessed information according to an information extraction rule that is specific to the information extraction processor and different from the information extraction rule of any other one of the two or more information extraction processors. And:

[0023] - generating an output via an output interface of the dialog system in dependence on the obtained information from one or more of the two or more information extraction processors, the output being an audio output and / or a textual and / or a visual output and / or a signal for operating the machine.

[0024] According to an embodiment, for example, selection of at least one of two or more information extraction processors through a dialogue system may be performed so that only those selected information extraction processors of the two or more information extraction processors generate the obtained information according to their information extraction rules.

[0025] and / or:

[0026] In an embodiment, for example, generating the derived information from the pre-processed information by at least two of the two or more information extraction processors according to their information extraction rules may be performed.

[0027] Furthermore, another method according to an embodiment is provided, comprising:

[0028] - obtaining an input representation of the input through an input interface of the dialog system, wherein the input interface obtains the input representation by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements; wherein each of the two or more information extraction processors of the dialog system is adapted to generate resulting information according to an information extraction rule specific to the information extraction processor and different from the information extraction rule of any other one of the two or more information extraction processors based on the input representation. And:

[0029] - generating an output through an output interface of the dialog system based on the obtained information from the two or more information extraction processors, the output being an audio output and / or a textual and / or a visual output and / or a signal for operating the machine.

[0030] At least two of the two or more information extraction processors are dialog state dependent. The method includes selecting, by the dialog system, one or more of the at least two dialog state dependent information extraction processors based on the current state of the dialog, so that only those of the at least two information extraction processors that are associated with the current state of the dialog are selected. In addition, the method includes generating the obtained information based on the information extraction rules of the one or more information extraction processors that have been selected.

[0031] Furthermore, a computer program is provided for implementing one of the above methods when executed on a computer or a signal processor.

[0032] In the following, further examples are provided.

[0033] According to an embodiment, input (e.g., input text or a representation of input text) may be received, for example, in a dialog system, an information extraction processor may be selected, which may be related to a current state of the dialog, for example, and the input may be processed, for example, by the selected information extraction processor, to extract information from the input.

[0034] In an embodiment, an input (e.g., input text or a representation of input text) may be received, for example, in a dialog system. The input may be processed, for example, by a feature extractor to obtain an input feature vector from the input, and the obtained input feature vector may be processed, for example, by at least two different information extraction processors to extract information from the input.

[0035] According to an embodiment, an input (e.g., input text or a representation of input text) may be received, for example, in a dialogue system, the input may be processed, for example, by a feature extractor to obtain an input feature vector from the input, an information extraction processor may be selected, for example, which may be related to a current state of a dialogue, and the obtained input feature vector may be processed, for example, by the selected information extraction processor, to extract information from the input.

[0036] In an embodiment, input (e.g., input text or a representation of input text) may be received, for example, in a dialog system, the input may be processed, for example, by a feature extractor to obtain an input feature vector from the input, a classifier may be selected, for example, that may be related to a current state of the dialog, and classification may be performed, for example, based on a set of different category representation vectors representing a set of categories supported by the classifier block, and the classification may be performed using a distance metric between the category representation vector and the input feature vector of the input.

[0037] According to an embodiment, an input (e.g., input text or a representation of the input text) may be received, for example, in a dialog system, the input may be processed, for example, by a feature extractor to obtain an input feature vector from the input, a classifier may be selected, for example, which may be related to the current state of the dialog, the input feature vector may be processed, for example, by the selected classification block to obtain a classifier-specific feature vector for each classifier block, and classification in each classifier block may be performed, for example, based on the corresponding classifier-specific feature vector. Optionally, for example, classification may be performed based on a set of different class representation vectors representing a set of classes supported by the classifier block, and classification may be performed using a distance metric between the class representation vector and the input classifier-specific feature vector. For example, an information extraction processor may be utilized in such embodiments.

[0038] In an embodiment, the input to be classified (eg, input text or a representation of input text) may be represented, for example, by a corresponding vector x (eg, a numeric vector x).

[0039] According to an embodiment, the input may be, for example, input text, or may be, for example, a representation of input text.

[0040] In an embodiment, a speech recognition system may generate input text, for example, from user speech recorded by one or more microphones.

[0041] In an embodiment, the input may, for example, be speech, for example, be a speech signal, or may, for example, be a representation of speech.

[0042] According to an embodiment, the input may, for example, comprise a plurality of speech posterior probability maps, or may, for example, comprise representations of a plurality of speech posterior probability maps.

[0043] In the following, embodiments of the present invention are described in more detail with reference to the accompanying drawings, in which:

[0044] Fig. 1a shows a dialog system according to an embodiment.

[0045] FIG. 1 b shows a dialog system according to another embodiment.

[0046] Figure 2 A dialog diagram depicting an example of a user interacting with a dialog system is shown, depicting states and transitions of a dialog path.

[0047] Figure 3 The dialog configurations for different domains are shown as a parallel set of disconnected dialog graphs.

[0048] Figure 4 The general form of a deep neural network is shown for classifying intent, domain, or dialog act based on an input vector.

[0049] Figure 5 A neural network of a dialog system according to an embodiment is shown, wherein a first part of the neural network is configured to determine a feature representation, and a second part of the neural network is configured to perform classification based on the feature representation.

[0050] Figure 6 An embodiment is shown in which different classifiers are applied simultaneously to provide classification results for different classification tasks.

[0051] Figure 7 An embodiment is shown in which different classifier-specific feature vector representations are utilized for different classifiers.

[0052] Figure 8 States and transitions of a dialog system according to an embodiment are shown.

[0053] Fig. 9 A neural network structure according to an embodiment is shown.

[0054] According to the embodiment shown in Figure 1a, the dialogue system includes an input interface 105, which is used to obtain an input representation of the input by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation includes multiple input representation elements.

[0055] Furthermore, the dialog system comprises a preprocessor 110 for preprocessing the input representation to generate preprocessed information, such that the preprocessed information comprises a plurality of preprocessed information elements, and such that each of two or more of the plurality of preprocessed information elements depends on at least two of the plurality of information representation elements.

[0056] Furthermore, the dialog system comprises two or more information extraction processors 120, 123, wherein each of the two or more information extraction processors 120, 123 is adapted to generate resulting information from the pre-processed information according to information extraction rules that are specific to the information extraction processor and are different from the information extraction rules of any other one of the two or more information extraction processors 120, 123.

[0057] Furthermore, the dialog system comprises an output interface 135 for generating an output depending on the obtained information from one or more of the two or more information extraction processors 120, 123, the output being an audio output and / or a text output and / or a visual output and / or a signal for operating the machine.

[0058] According to an embodiment, the dialog system may be configured, for example, to select at least one of the two or more information extraction processors 120, 123 such that only those of the two or more information extraction processors 120, 123 that have been selected generate resultant information according to their information extraction rules.

[0059] In an embodiment, at least two of the two or more information extraction processors 120, 123 may generate the resulting information from the pre-processed information, for example, according to their information extraction rules. For example, the at least two of the two or more information extraction processors 120, 123 may be configured to generate the resulting information from the pre-processed information in parallel.

[0060] According to an embodiment, the dialogue system can be configured, for example, to select at least one of the two or more information extraction processors 120, 123 based on the current state of the dialogue, so that the selected ones of the two or more information extraction processors 120, 123 generate the obtained information based on their information extraction rules.

[0061] In an embodiment, at least two of the two or more information extraction processors 120, 123 may be, for example, dialog state dependent. The dialog system may be configured, for example, to select one or more of the at least two information extraction processors 120, 123 that are dialog state dependent, based on the current state of the dialog, so that only those of the at least two information extraction processors 120, 123 that are associated with the current state of the dialog are selected. Each of the selected one or more information extraction processors 120, 123 may be configured, for example, to generate the resulting information based on its information extraction rule.

[0062] According to an embodiment, the dialog system includes three or more information extraction processors 120, 123 as two or more information extraction processors 120, 123. At least one information extraction processor 120, 123 of the three or more information extraction processors 120, 123 may, for example, be dialog state independent. The dialog state independent at least one information extraction processor 120, 123 may, for example, be configured to always generate the resulting information according to its information extraction rule independently of the current state.

[0063] According to an embodiment, each of at least two of the two or more information extraction processors 120, 123 may be adapted to generate specific information specific to the information extraction processor according to a modification rule, wherein the information extraction processor is adapted to generate the resulting information from the specific information for the information extraction processor according to the information extraction rule specific to the information extraction processor. The information extraction processor may be adapted to generate specific information for the information extraction processor according to the modification rule, such that the specific information for the information extraction processor is different from any specific information of any other information extraction processor of the at least two information extraction processors 120, 123.

[0064] According to an embodiment, each of at least one of the at least two information extraction processors 120, 123 may, for example, be configured to generate specific information for the information extraction processor using the obtained information of another one of the at least two information extraction processors 120, 123.

[0065] In an embodiment, the dialogue system can be configured, for example, to select at least one of the at least two information extraction processors 120, 123 based on the current state of the dialogue, so that each of the at least one information extraction processor can, for example, generate the resulting information from specific information used for the one of the at least one information extraction processor.

[0066] According to an embodiment, each of the two or more information extraction processors 120, 123 may, for example, be a classification unit. Each of the two or more classification units may, for example, be adapted to generate derived information from the preprocessing information such that the derived information indicates whether the input representation may, for example, be associated with a category, or indicates a probability that the input representation may, for example, be associated with a category.

[0067] In an embodiment, the pre-processing information may, for example, comprise a numerical feature vector.The plurality of pre-processing information elements may, for example, comprise a plurality of numerical vector components of the feature vector.

[0068] According to an embodiment, the input interface 105 may be configured to obtain an original input text, i.e., a word sequence, as an input representation. For example, the preprocessor 110 may be configured to tokenize the original input text using a tokenization method to obtain a plurality of tags. In addition, the preprocessor 110 may be configured to generate a multidimensional numerical vector for each of the plurality of tags to obtain a plurality of multidimensional numerical vectors. In addition, the preprocessor 110 may be configured to generate a numerical feature vector of preprocessing information by combining a plurality of multidimensional numerical vectors for a plurality of tags.

[0069] According to an embodiment, for each of the at least one information extraction processors 120, 123 that have been selected, the information extraction processor can, for example, be configured to generate specific information, such that the specific information includes a numerical feature vector that depends on the numerical feature vector of the preprocessing information. In addition, the information extraction processor can, for example, be configured to generate the resulting information for the information extraction processor by determining a distance metric between the specific information of the information extraction processor and the numerical category representation vector associated with the information extraction processor.

[0070] In an embodiment, each of the two or more information extraction processors 120, 123 may, for example, include a neural network, wherein the neural network may, for example, include at least one of an attention layer, a pooling layer, and a fully connected layer. The neural network may, for example, be configured to receive preprocessed information as input, and may, for example, be configured to output the resulting information; or the neural network may, for example, be configured to receive specific information for the information extraction processor as input, and may, for example, be configured to output the resulting information.

[0071] According to an embodiment, the pre-processor may, for example, be configured to generate the pre-processed information such that each of the plurality of information elements is dependent on each of the plurality of information representation elements.

[0072] In an embodiment, the preprocessor may, for example, include a neural network, which may, for example, be configured to receive a plurality of input representation elements as input, and may, for example, be configured to output a plurality of preprocessed information elements as output. For example, the neural network may include at least two of an attention layer, a pooling layer, and a fully connected layer.

[0073] According to an embodiment, the input representation may, for example, include a numerical multidimensional sentence representation vector, or wherein the preprocessor 110 may, for example, be configured to generate a numerical multidimensional sentence representation vector from the input representation, wherein the numerical multidimensional sentence representation vector (for example) identifies a sentence or part of a sentence based on a user utterance, wherein the multidimensional sentence representation vector may, for example, include three or more numerical vector elements, wherein each of the three or more numerical vector elements may, for example, be associated with one of a plurality of dimensions.

[0074] In an embodiment, for each two pairs of the plurality of numerical multidimensional sentence representation vectors for the plurality of sentences represented by the input, the two numerical multidimensional sentence representation vectors of the first pair of numerical multidimensional sentence representation vectors that identify two first sentences with semantically related meanings in the two pairs of numerical multidimensional sentence representation vectors have a smaller spatial distance in the multidimensional space in which the plurality of numerical multidimensional sentence representation vectors are defined than the two numerical multidimensional sentence representation vectors of the second pair of numerical multidimensional sentence representation vectors that identify two second sentences with semantically unrelated meanings in the two pairs of numerical multidimensional sentence representation vectors. Alternatively, the preprocessor 110 may, for example, be configured to generate the plurality of numerical multidimensional sentence representation vectors so that for each two pairs of the plurality of numerical multidimensional sentence representation vectors, the two numerical multidimensional sentence representation vectors of the first pair of numerical multidimensional sentence representation vectors that identify two first sentences with semantically related meanings in the two pairs of numerical multidimensional sentence representation vectors have a smaller spatial distance in the multidimensional space in which the plurality of numerical multidimensional sentence representation vectors are defined than the two numerical multidimensional sentence representation vectors of the second pair of numerical multidimensional sentence representation vectors that identify two second sentences with semantically unrelated meanings in the two pairs of numerical multidimensional sentence representation vectors. For example, the spatial distance may be the Euclidean distance or the sine similarity or the cosine similarity.

[0075] FIG. 1 b shows a dialog system according to another embodiment.

[0076] The dialogue system of FIG. 1ab includes an input interface 105 for obtaining an input representation of the input by receiving the input and deriving an input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation includes a plurality of input representation elements.

[0077] In addition, the dialogue system includes two or more information extraction processors 120, 123, wherein each of the two or more information extraction processors 120, 123 is adapted to generate resulting information based on an input representation according to information extraction rules that are specific to the information extraction processor and different from the information extraction rules of any other one of the two or more information extraction processors 120, 123.

[0078] Furthermore, the dialog system comprises an output interface 135 for generating an output depending on the obtained information from one or more of the two or more information extraction processors 120, 123, the output being an audio output and / or a text output and / or a visual output and / or a signal for operating the machine.

[0079] At least two of the two or more information extraction processors 120, 123 are dialog state dependent. The dialog system is configured to select one or more of the at least two dialog state dependent information extraction processors 120, 123 according to the current state of the dialog, so that only those of the at least two information extraction processors 120, 123 that are associated with the current state of the dialog are selected. The selected one or more information extraction processors 120, 123 are configured to generate the obtained information according to their information extraction rules.

[0080] According to an embodiment, the dialog system includes three or more information extraction processors 120, 123 as two or more information extraction processors 120, 123. At least one information extraction processor 120, 123 of the three or more information extraction processors 120, 123 may, for example, be dialog state independent. The dialog state independent at least one information extraction processor 120, 123 may, for example, be configured to always generate the resulting information according to its information extraction rule independently of the current state.

[0081] In an embodiment, each of at least two of the two or more information extraction processors 120, 123 may be adapted to generate specific information specific to the information extraction processor according to a modification rule, wherein the information extraction processor is adapted to generate the resulting information from the specific information for the information extraction processor according to the information extraction rule specific to the information extraction processor. The information extraction processor may be adapted to generate specific information for the information extraction processor according to the modification rule, such that the specific information for the information extraction processor is different from any specific information of any other of the at least two information extraction processors 120, 123.

[0082] According to an embodiment, the input interface 105 may be configured to receive an input, the input being a speech signal or an audio signal. The input interface 105 may be configured to apply a speech recognition algorithm to the speech signal or the audio signal to obtain a text representation of the speech signal or the audio signal as an input representation.

[0083] In certain embodiments, the speech signal may, for example, include a command to operate the machine. In response, the machine may, for example, execute the command.

[0084] Hereinafter, specific embodiments are described.

[0085] The input to be classified (eg, input text or a representation of input text) may, for example, be represented by a corresponding vector x, such as a numerical vector x.

[0086] In the following, a preferred embodiment of determining preprocessing information based on text input is described. It follows the concept described in [1].

[0087] First, the dialogue system receives raw input text which is a sequence of words. The input text is first tokenized, for example, using the WordPiece tokenization method, e.g., described by Wu et al.

[11] . Tokens can represent words, subwords, or just parts of words. Each token is then mapped / converted to a high-dimensional numerical vector representation via a matrix of size (N, D) called the token embedding matrix, where N is the size of the vocabulary and D is the size of the vector in the high-dimensional vector space (multidimensional vector space). Each row of the matrix corresponds to a token in the vocabulary. So, for example, if the length of the input text is S=6 (it contains 6 tokens) and D is 768, the corresponding token representations are extracted and the input text is represented as a matrix of size (6, 768) in the vector space.

[0088] If the dialogue system receives a pair of raw input texts instead of a single input text, in order to distinguish the inputs, a "segment embedding" matrix of shape (2, D) can be used, for example. In this matrix, the first row (all values ​​0) is assigned to all tokens belonging to input 1, and the last row (all values ​​1) is assigned to all tokens belonging to input 2. If the input text consists of only one input sequence, then its segment embedding will be just the first row of the segment embedding matrix. Thus, if the total length of the input pair is 10 (4 for the first input and 6 for the second input), then the segment representation of the pair can be, for example, a matrix of size (10, 768), where at most 4 rows are 0 and the other rows are 1. If the input text is a single input of length 6, then the segment embedding can be, for example, a matrix of size (6, 768), where all values ​​are 0. This way, the language model in the dialogue system can identify which token belongs to which input. Note that in our examples, for example, we always have a single input text.

[0089] Then, the position of the token in the input text is represented by a high-dimensional vector representation. To obtain the position vector of each token, a lookup table of size (L, D) called position embedding can be used, for example, where the first row is the vector representation of any token at the first position, the second row is the vector representation of any token at the second position, and so on. Here, L shows the maximum sequence length that the language model of the dialogue system can preprocess, for example 512.

[0090] For example, the token representation, segment representation, and position representation of the input text can be element-wise summed to produce a single representation with shape (S, D), where s is the input text length (number of tokens), and d is the vector size in the high-dimensional vector space. This is, for example, the input representation that can be passed to the first layer of the system's language model.

[0091] The language model may, for example, have several layers, for example, 12 layers, with the same architecture. For example, each layer may include an attention network and a fully connected neural network. For example, both the input and output of a layer may be of shape (S, D). The previously calculated input representation passes through all layers, and the output of the last layer may, for example, be an output representation of the same shape (S, D).

[0092] To compute the pre-processed information input to the information extraction processor, the output representations of size (S, D) from the language model can be aggregated by taking the element-wise average of all S token vectors, for example, which can produce, for example, a single high-dimensional vector representation of length D that can correspond to the input text. This vector can be considered a sentence representation vector or sentence embedding or sentence embedding vector, and it corresponds to the processed information output by the pre-processor 110.

[0093] According to an embodiment, a speech signal may be obtained, for example, from a microphone, and a text may be derived from the recorded microphone, for example, using a speech recognition algorithm. For example, the obtained text may be mapped to an input vector by an algorithm based on a vocabulary.

[0094] Alternatively, in another embodiment, the speech recognition algorithm may be designed to directly map sentences included in the speech signal to indices of the vocabulary.

[0095] According to another embodiment, a speech analysis algorithm may be used, for example, to map a speech signal to an index of a vocabulary of a plurality of speech posterior probability maps to obtain a vector of indices of the vocabulary of speech posterior probability maps. In such a vocabulary of speech posterior probability maps, each speech posterior probability map may be represented, for example, by an index. The mapping algorithm may, for example, map a vector of indices of speech posterior probability maps to an index of a word of a vocabulary of words.

[0096] In another embodiment, the vector of indices of the speech posterior probability map may be directly processed, for example as input to a feature extractor.

[0097] A neural network representing a classifier can usually be represented by a function f(A,x), where the parameter is A and the input is x.

[0098] For example, the output of a classifier could be a vector c where the elements c iIt can be a numerical score, for example, which can be interpreted as the input x belongs to c i The probability of the category represented by . This general classification method is as follows Figure 4 shown.

[0099] Certain embodiments address the complexity issues noted above.

[0100] Some embodiments are based on the observation that a neural network for classification can be interpreted as consisting of two parts:

[0101] The first part of the neural network computes what are called feature representations, or embeddings, from the user input.

[0102] The second part of the neural network performs the actual classification task by processing the embeddings determined by the first part of the neural network. This approach to neural network-based classification systems that interpret, for example, intents, dialogue acts, or domains is described in [1]. Figure 5 shown.

[0103] In an embodiment, user input (e.g., input text or a representation of input text) can be processed by different classifiers, for example, simultaneously (e.g., in parallel) to provide classification results for different classification tasks (e.g., local intent classification, global intent classification, or domain classification). Suggested methods for this scenario are as follows: Figure 6 as shown, and includes multiple processing steps.

[0104] exist Figure 6 , Figure 7 and Fig. 9 In, the preprocessor 110 can be implemented as a feature extractor 610, for example. Figure 6 , Figure 7 and Fig. 9 In the example, the information extraction processors 120, 123 may be implemented as classification units 620, 623, 626, for example.

[0105] First, an input (eg, input text or a representation of input text), such as an input vector, may be processed by a feature extractor 610 to determine a corresponding feature vector or text embedding.

[0106] For example, feature extractor 610 may be implemented by a neural network that has been trained to generate appropriate feature vectors based on the user's input, for example by applying the concepts described in [9].

[0107] According to an embodiment, the concepts described in

[10] , in particular in

[10] , chapter 3.2, may be utilized to derive a feature vector, for example, from a text input or from a representation of a text input.

[0108] For example, in an embodiment, an input, such as an input text, may first be transformed from text into sentence representation vectors or so-called sentence embeddings. For example, these vectors may be digital vectors of dimension L, where the dimension of the sentence representation vector may, for example, include 100 to 1000 dimensions.

[0109] In an embodiment, the text or text representation may, for example, be transformed into sentence representation vectors such that sentences with similar semantics may, for example, be close to each other in space; for example, may have similar vector representations, for example as measured by a distance metric such as Euclidean distance or sine similarity or cosine similarity.

[0110] In an embodiment, a neural network may be used, for example, to generate a feature vector from an input. For example, the neural network may include, for example, an attention layer and / or a pooling layer and / or a fully connected layer. In a particular layer, a configuration of a neural network may be used, for example, including an attention layer followed by a pooling layer. In addition, in a particular embodiment, the neural network may also include, for example, a fully connected layer following an attention layer (see

[10] ). Figure 2 ).

[0111] The feature vectors calculated by the feature extractor 610 can be input to multiple classification units 620, 623, 626, for example. For example, each classification unit 620, 623, 626 can be dedicated to solving a separate classification task and provide information related to the corresponding classification results. Each different classification unit 620, 623, 626 can be trained or configured differently, usually based on different training data.

[0112] For example, the classification units 620, 623, 626 can be implemented using neural networks that provide classification information in the form of an estimate of the probability that an input (e.g., input text or a representation of input text) can be associated with a particular category. Alternatively, other classification methods can also be utilized, such as support vector machines or decision trees, for example.

[0113] Figure 7 Another embodiment is shown, in which the classification unit 620, 623 first determines (e.g., by the first subunit 721, 724) a classifier-specific feature vector representation of the input from the input feature vector to obtain the classifier-specific feature vector representation of the input. Typically, the classifier-specific feature vector representation of the input is different for different classifiers.

[0114] Classification may then be performed, for example, based on a distance measure between the classifier-specific feature vector and a set of different class representation vectors representing a set of classes supported by the respective classification unit 620 , 623 (eg, by the second subunit 722 , 725 ).

[0115] Therefore, different classification units 620, 623 can use different groups of category representation vectors, wherein each category representation vector is associated with a specific category or label. Generally, the closer the feature vector specific to the classifier is to the representation vector of the category, the greater the probability that the input belongs to the category. Classification units 620, 623 (e.g., second subunits 722, 725) can, for example, select the category with the smallest distance measurement between the feature vector specific to the classifier and the representation vector of the category of the input. For example, the Euclidean distance measurement can be utilized. In other embodiments, another distance measurement can be utilized alternatively, for example, the cosine distance, or, for example, the sine distance.

[0116] In some embodiments, the computation of the classifier-specific feature vector representation may be performed, for example, based on a linear projection of the input feature vector (e.g., by the first subunit 721, 724). For example, the linear projection may be represented by a corresponding projection matrix, wherein the projection matrix is ​​typically selected differently for different classifiers, for example based on training data associated with the particular classifier (see [8]).

[0117] In some embodiments, the step of determining a classifier-specific feature vector is omitted. In this case, the category representation vector associated with a particular classifier will be directly compared with the input feature vector of the input text, rather than being compared with the corresponding classifier-specific feature vector.

[0118] In general, the computational complexity of calculating the input feature vector from the input is much higher than the complexity of calculating the classifier-specific feature vector based on the input feature vector. Therefore, the overall computational complexity is much lower for the proposed method. Similarly, the number of parameters of the neural network used to calculate the input feature vector is much higher than the number of parameters required to calculate the classifier-specific feature vector from the input feature vector. This means that the memory requirement using the proposed method is smaller than a set of corresponding separate classifiers.

[0119] According to some embodiments, an additional approach to improve the efficiency and robustness of classification tasks in a dialog system is to evaluate not all classifiers of the dialog system (e.g., for all available intents, domains, or dialog acts), but only those classifiers that are needed to determine the appropriate next behavior of the dialog system at a specific state of the dialog. For example, following the graph-based representation of the dialog, so-called global intents are relevant in every round of conversation between the user and the dialog system, that is, they are required to be independent of the actual state of the dialog or the position of the dialog in the graph-based representation. An example of a global intent is provided by a "stop" command, which should stop the dialog system. On the other hand, so-called local intents are only relevant at a specific state of the dialog or at a specific position in the graph-based representation of the dialog in order to determine the next behavior of the dialog system.

[0120] In an embodiment, classifiers associated with local intent are evaluated only when needed, for example, depending on the specific state of the conversation. According to some embodiments, classifiers associated with global intent are always evaluated.

[0121] Figure 8 An example is shown, in which a dialog system is considered, which supports a user in performing different tasks, like booking a ticket or managing a calendar.

[0122] Design the dialogue so that in state S a Next, the dialogue system requires the user's intention to book a ticket I book Or modify the current reservation? mod This classification task is performed by the corresponding local intent classifier IC a 811 resolved. In another state S b In this case, the dialog system needs information about whether the user's intention is to schedule a new meeting or cancel a calendar entry, which is determined by the corresponding local intent classifier IC b 812 processing. In addition to the local intent classifier, it is assumed that there is also a global intent classifier IC in the dialogue system g 821, used to detect general user intentions, such as stopping the dialog system or restarting a session with the dialog system.

[0123] For the state S a The proposed method is applied to the dialog system as follows: First, the input feature extractor 610 processes the user input (eg, input text or a representation of the input text) and outputs an input feature vector.

[0124] Then, a classifier 811, 812, 821 is selected from a set of available classifiers in the dialog system that corresponds to state S a Related classifiers 811, 821. In the example here, the local intent classifier IC is selected a 811 and Global Intent Classifier IC g 821, and IC b 812 is not considered because it is in state S a The following are irrelevant. Then, the input feature vectors are respectively composed of IC a 811 and IC g 821 corresponding taxonomic units are processed, while IC b The classification unit of 812 is omitted. Similarly, if the dialogue is in state S b , select IC g 821 and IC b The classification unit of 812 further processes the input feature vector, omitting IC a 811 taxa.

[0125] Fig. 9 A neural network structure according to an embodiment is shown. A preprocessor 610 (e.g., a feature extractor) can be implemented, for example, as a first neural network. Two information extraction processors 620, 623 (e.g., two classification units) can be implemented, for example, as two additional neural networks. The output of the last layer of the preprocessor 610 can be, for example, preprocessing information (e.g., a feature vector). It can be, for example, that each preprocessing information element of the preprocessing information is fed to the first information extraction processor 620 and the second information processor 623, which each generate their resulting information (e.g., a classification result, for example, the input fed to the preprocessing processor 610 represents the probability of belonging to a specific category associated with the corresponding classifier 620 or 623).

[0126] Alternatively, the neural network structure may be implemented, for example, by a single neural network. The output of the last layer of the preprocessor 610 may be, for example, preprocessing information (e.g., a feature vector). In such a single neural network structure, for example, there may be no link between the nodes of the information extraction processor 620 and the nodes of the information extraction processor 623. For example, such a single neural network structure may be trained with a large number of training data sets, wherein, for example, for each of the classification units 620 and 623, each training data set has an input representation as an input and a classification result (e.g., 1 = the input belongs to the category, or 0 = the input does not belong to the category).

[0127] Alternatively, the neural network of the feature extractor 610 can be implemented, for example, using a prior art neural network to obtain feature vectors, such as BERT (see [1]) or SBERT (see [9]). For example, the training data of the neural network of each classification unit 620, 623 may include the output of the prior art neural network 610 as input, and the classification result of the classification associated with the corresponding classification unit 620 or 623 as output.

[0128] The above description of the embodiments is provided in particular with respect to classification tasks and intent classification. As described above, for example, there may be additional information extraction tasks that are solved by the information extraction processor of the dialog system. Examples include domain classification, dialog act classification, or entity recognition. In an embodiment, the proposed method described in the context of intent classification may also be applied, for example, similarly to these information extraction tasks. In this case, for example, not only an intent classifier associated with a specific dialog state may be selected, but also (or alternatively) an information extraction processor 120, 123 associated with domain classification, dialog act classification, or entity recognition associated with the dialog state under consideration may be selected from all available information extraction processors 120, 123 of the dialog system. A common input feature vector obtained from the input (e.g., from the input text or from a representation of the input text) is then processed by different information extraction processors 120, 123 associated with the current dialog state in order to extract the desired information.

[0129] Although some aspects have been described in the context of an apparatus or a dialog system, it is clear that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus or dialog system. Some or all of the method steps may be performed by (or using) a hardware device (e.g., a microprocessor, a programmable computer, or an electronic circuit). In some embodiments, one or more of the most important method steps may be performed by such a device.

[0130] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. The implementation may be performed using a digital storage medium such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, the digital storage medium having electronically readable control signals stored thereon, the electronically readable control signals cooperating with (or capable of cooperating with) a programmable computer system to perform the corresponding method. Thus, the digital storage medium may be computer readable.

[0131] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that the system performs one of the methods described herein.

[0132] Generally, the embodiments of the present invention can be implemented as a computer program product with a program code, when the computer program product runs on a computer, the program code is operable to perform one of the above methods.For example, the program code can be stored on a machine-readable carrier.

[0133] Other embodiments comprise the computer program for performing one of the methods described herein, the computer program being stored on a machine readable carrier.

[0134] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0135] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods herein. The data carrier, the digital storage medium, or the recorded medium are typically tangible and / or non-transitory.

[0136] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.For example, the data stream or the sequence of signals may be configured to be transmitted via a data communication connection, for example via the Internet.

[0137] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods herein.

[0138] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0139] A further embodiment according to the invention comprises an apparatus or a system configured to transmit (e.g. electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a storage device, etc. For example, the apparatus or the system may comprise a file server for transmitting the computer program to the receiver.

[0140] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may be coordinated with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.

[0141] The devices described herein may be implemented using hardware devices, or using computers, or using a combination of hardware devices and computers.

[0142] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0143] The above embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to other persons skilled in the art. Therefore, it is intended that the scope of the present invention be limited only by the scope of the pending patent claims and not by the specific details presented herein by way of description and explanation of the embodiments.

[0144] References:

[0145] [1] Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.", In Proceedings of the 2019

[0146] Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.Minneapolis: Association for Computational Linguistics.4171-4186.

[0147] [2]Grosz, Barbara J., and Candace L.Sidner.1986. "Attention, Intentions, and the Structure of Discourse." Computational Linguistics 12: 175--204.

[0148] [3]Louvan,Samuel,and Bernardo Magnini.2020."Recent Neural Methods onSlot Filling and Intent Classification for Task-Oriented Dialogue Systems:ASurvey."Proceedings of the 28th International Conference on ComputationalLinguistics.Barcelona,Spain(Online):International Committee on ComputationalLinguistics.480--496.

[0149] [4]Young,Steve,Milica Blaise Thomson,and Jason D.Williams.2013."POMDP-Based Statistical Spoken Dialog Systems:AReview"Proceedings of the IEEE101(5):1160-1179.

[0150] [5]Qin,Libo,Xiao Xu,Wanxiang Che,Yue Zhang,and Ting Liu.2020."DynamicFusion Network for Multi-Domain End-to-end Task-Oriented Dialog."Proceedingsof the 58th Annual Meeting of the Association for ComputationalLinguistics.Online:Association for Computational Linguistics.6344-6354.

[0152] [6]Yadav,Vikas,and Steven Bethard.2018."ASurvey on Recent Advances inNamed Entity Recognition from Deep Learning models."Proceedings of the 27thInternational Conference on Computational Linguistics.Santa Fe,New Mexico,USA:Association for Computational Linguistics.2145--2158.[7]Pan,Sinno Jialin,and Qiang Yang.2010."ASurvey on Transfer Learning."IEEE Transactions onKnowledge and Data Engineering 1345-1359.

[0153] [8]Weinberger,Kilian Q,and Lawrence K.Saul.2009."Distance MetricLearning for Large Margin Nearest Neighbor Classification."Journal of MachineLearning Research 207-244.

[0154] [9]Nils Reimers and Iryna Gurevych.2019."Sentence-{BERT}:SentenceEmbeddings using{S}iamese

[0155] {BERT}-Networks."Proceedings of the 2019Conference on EmpiricalMethods in Natural Language Processing and the 9th International JointConference on Natural Language Processing(EMNLP-

[0156] IJCNLP).Hong Kong,China:Association for ComputationalLinguistics.3982-3992.

[0157]

[10] H Li,Y Ma,Z Ma,H Zhu,Weibo text sentiment analysis based on bertand deep learning,

[0158] Applied Sciences,2021.

[0159]

[11] Y.Wu et al.,”Google's Neural Machine Translation System:Bridgingthe Gap between Human and Machine Translation,”2016:https: / / arxiv.org / abs / 1609.08144

Claims

1. A dialogue system, comprising: An input interface (105) is used to obtain an input representation of the input by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation includes a plurality of input representation elements, a preprocessor (110; 610) for preprocessing the input representation to generate preprocessed information, such that the preprocessed information includes a plurality of preprocessed information elements, and such that each of two or more of the plurality of preprocessed information elements is dependent on at least two of the plurality of information representation elements, two or more information extraction processors (120, 123; 620, 623, 626), wherein each of the two or more information extraction processors (120, 123; 620, 623, 626) is adapted to generate resultant information from the pre-processed information according to an information extraction rule that is specific to the information extraction processor and that is different from the information extraction rule of any other one of the two or more information extraction processors (120, 123; 620, 623, 626), and An output interface (135) is used to generate an output based on the information obtained from one or more of the two or more information extraction processors (120, 123; 620, 623, 626), the output being an audio output and / or a text output and / or a visual output and / or a signal for operating a machine.

2. The dialogue system according to claim 1, in, The dialog system is configured to select at least one of two or more information extraction processors (120, 123; 620, 623, 626) so that only those selected information extraction processors of the two or more information extraction processors (120, 123; 620, 623, 626) generate the obtained information according to their information extraction rules.

3. The dialogue system according to claim 1 or 2, in, At least two of the two or more information extraction processors (120, 123; 620, 623, 626) generate derived information from the pre-processed information according to their information extraction rules.

4. The dialogue system according to claim 3, in, The at least two (120, 123) of the two or more information extraction processors are configured to generate derived information from the pre-processed information in parallel.

5. A dialog system according to any one of the preceding claims, in, The dialog system is configured to select at least one of two or more information extraction processors (120, 123; 620, 623, 626) based on the current state of the dialog, so that the selected ones of the two or more information extraction processors (120, 123; 620, 623, 626) generate the obtained information based on their information extraction rules.

6. The dialogue system according to claim 5, in, At least two of the two or more information extraction processors (120, 123; 620, 623, 626) are dialog-state dependent, The dialog system is configured to select one or more information extraction processors (120, 123; 620, 623, 626) of at least two information extraction processors (120, 123; 620, 623, 626) on which the dialog state depends, based on the current state of the dialog, so that only those information extraction processors of the at least two information extraction processors (120, 123; 620, 623, 626) associated with the current state of the dialog are selected, and The selected one or more information extraction processors (120, 123; 620, 623, 626) are configured to generate the obtained information according to their information extraction rules.

7. The dialogue system according to claim 6, in, The dialog system includes three or more information extraction processors (120, 123; 620, 623, 626) as two or more information extraction processors (120, 123; 620, 623, 626), wherein at least one of the three or more information extraction processors (120, 123; 620, 623, 626) is independent of the dialog state, Therein, at least one dialog state-independent information extraction processor (120, 123; 620, 623, 626) is configured to always generate the obtained information according to its information extraction rules independently of the current state.

8. A dialog system according to any one of the preceding claims, in, At least two of the two or more information extraction processors (120, 123; 620, 623, 626) are each adapted to generate specific information specific to the information extraction processor according to a modification rule, wherein the information extraction processor is adapted to generate derived information from the specific information for the information extraction processor according to an information extraction rule specific to the information extraction processor, The information extraction processor is adapted to generate specific information for the information extraction processor according to a modification rule, so that the specific information for the information extraction processor is different from any specific information of any other information extraction processor among at least two information extraction processors (120, 123; 620, 623, 626).

9. The dialogue system according to claim 8, in, Each of at least one of the at least two information extraction processors (120, 123; 620, 623, 626) is configured to generate specific information for the information extraction processor using information obtained by another of the at least two information extraction processors (120, 123; 620, 623, 626).

10. The device according to claim 8 or 9, further depending on claim 5, in, The dialog system is configured to select at least one of at least two information extraction processors (120, 123; 620, 623, 626) based on the current state of the dialog, so that each of the at least one information extraction processor generates obtained information from specific information for the one of the at least one information extraction processor.

11. A dialog system according to any one of the preceding claims, in, Each of the two or more information extraction processors (120, 123; 620, 623, 626) is a classification unit, Therein, each of the two or more classification units is adapted to generate derived information from the preprocessed information, such that the derived information indicates whether the input representation is associated with a category, or indicates a probability that the input representation is associated with a category.

12. A dialog system according to any one of the preceding claims, in, The preprocessing information includes numerical feature vectors, The plurality of preprocessing information elements include a plurality of numerical vector components of the feature vector.

13. The dialogue system according to claim 12, in, The input interface (105) is configured to obtain an original input text as an input representation, the original input text being a sequence of words, The preprocessor (110; 610) is configured to tokenize the original input text using a tokenization method to obtain a plurality of tokens. wherein the preprocessor (110; 610) is configured to generate a multidimensional numerical vector for each of the plurality of tags to obtain a plurality of multidimensional numerical vectors, The preprocessor (110; 610) is configured to generate a numerical feature vector of preprocessing information by combining a plurality of multi-dimensional numerical vectors for a plurality of tags.

14. A dialog system according to claim 12 or 13, further depending on claims 10 and 11, in, For each of the at least one information extraction processor (120, 123; 620, 623, 626) that has been selected, The information extraction processor is configured to generate specific information such that the specific information includes a numerical feature vector that depends on a numerical feature vector of the preprocessing information, and The information extraction processor is configured to generate resultant information for the information extraction processor by determining a distance metric between specific information of the information extraction processor and a numerical class representation vector associated with the information extraction processor.

15. A dialog system according to any one of the preceding claims, in, Each of the two or more information extraction processors (120, 123; 620, 623, 626) comprises a neural network, wherein the neural network comprises at least one of an attention layer, a pooling layer, and a fully connected layer, wherein the neural network is configured to receive the preprocessed information as input and is configured to output the resulting information; or Wherein the dialogue system is further according to claim 4, and wherein the neural network is configured to receive specific information for the information extraction processor as input and is configured to output the resulting information.

16. A dialog system according to any one of the preceding claims, in, The preprocessor is configured to generate preprocessed information such that each of the plurality of information elements is dependent on each of the plurality of information representation elements.

17. A dialog system according to any one of the preceding claims, in, The preprocessor includes a neural network configured to receive as input a plurality of input representation elements and configured to output as output a plurality of preprocessed information elements, The neural network includes at least two of an attention layer, a pooling layer, and a fully connected layer.

18. A dialog system according to any one of the preceding claims, wherein the input representation comprises a numerical multidimensional sentence representation vector, or wherein the preprocessor (110; 610) is configured to generate a numerical multidimensional sentence representation vector from the input representation, wherein the multidimensional sentence representation vector comprises three or more numerical vector elements, wherein each of the three or more numerical vector elements is associated with one of a plurality of dimensions.

19. The dialogue system according to claim 18, in, For each two pairs of the multiple numerical multidimensional sentence representation vectors for the multiple sentences represented by the input, the two numerical multidimensional sentence representation vectors of a first pair of numerical multidimensional sentence representation vectors identifying two first sentences having semantically related meanings in the two pairs of numerical multidimensional sentence representation vectors have a smaller spatial distance in the multidimensional space in which the multiple numerical multidimensional sentence representation vectors are defined than the two numerical multidimensional sentence representation vectors of a second pair of numerical multidimensional sentence representation vectors identifying two second sentences having semantically unrelated meanings in the two pairs of numerical multidimensional sentence representation vectors; or, The preprocessor (110; 610) is configured to generate a plurality of numerical multidimensional sentence representation vectors, such that for every two pairs of the plurality of numerical multidimensional sentence representation vectors, the two numerical multidimensional sentence representation vectors of a first pair of numerical multidimensional sentence representation vectors identifying two first sentences having semantically related meanings in the two pairs of numerical multidimensional sentence representation vectors have a smaller spatial distance in the multidimensional space in which the plurality of numerical multidimensional sentence representation vectors are defined than the two numerical multidimensional sentence representation vectors of a second pair of numerical multidimensional sentence representation vectors identifying two second sentences having semantically unrelated meanings in the two pairs of numerical multidimensional sentence representation vectors.

20. A dialogue system, comprising: an input interface (105) for obtaining an input representation of the input by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements, and two or more information extraction processors (120, 123; 620, 623, 626), wherein each of the two or more information extraction processors (120, 123; 620, 623, 626) is adapted to generate resultant information based on an input representation according to an information extraction rule that is specific to the information extraction processor and different from the information extraction rule of any other one of the two or more information extraction processors (120, 123; 620, 623, 626), and an output interface (135) for generating an output based on the information obtained from one or more of the two or more information extraction processors (120, 123; 620, 623, 626), the output being an audio output and / or a text output and / or a visual output and / or a signal for operating a machine, wherein at least two of the two or more information extraction processors (120, 123; 620, 623, 626) are dialog state dependent, The dialog system is configured to select one or more information extraction processors (120, 123; 620, 623, 626) of at least two information extraction processors (120, 123; 620, 623, 626) on which the dialog state depends, based on the current state of the dialog, so that only those information extraction processors of the at least two information extraction processors (120, 123; 620, 623, 626) associated with the current state of the dialog are selected, and The selected one or more information extraction processors (120, 123; 620, 623, 626) are configured to generate the obtained information according to their information extraction rules.

21. The dialogue system according to claim 20, in, The dialog system includes three or more information extraction processors (120, 123; 620, 623, 626) as two or more information extraction processors (120, 123; 620, 623, 626), wherein at least one of the three or more information extraction processors (120, 123; 620, 623, 626) is independent of the dialog state, Therein, at least one dialog state-independent information extraction processor (120, 123; 620, 623, 626) is configured to always generate the obtained information according to its information extraction rules independently of the current state.

22. A dialog system according to any one of claims 20 or 21, in, At least two of the two or more information extraction processors (120, 123; 620, 623, 626) are each adapted to generate specific information specific to the information extraction processor according to a modification rule, wherein the information extraction processor is adapted to generate derived information from the specific information for the information extraction processor according to an information extraction rule specific to the information extraction processor, The information extraction processor is adapted to generate specific information for the information extraction processor according to a modification rule, so that the specific information for the information extraction processor is different from any specific information of any other information extraction processor among at least two information extraction processors (120, 123; 620, 623, 626).

23. A dialog system according to any one of the preceding claims, in, The input interface (105) is configured to receive an input, the input being a speech signal or an audio signal, The input interface (105) is configured to apply a speech recognition algorithm to the speech signal or the audio signal to obtain a text representation of the speech signal or the audio signal as an input representation.

24. A method comprising: obtaining an input representation of the input through an input interface (105) of the dialog system, wherein the input interface (105) obtains the input representation by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements, preprocessing the input representation by a preprocessor (110; 610) of the dialog system to generate preprocessed information, such that the preprocessed information includes a plurality of preprocessed information elements, and such that each of two or more of the plurality of preprocessed information elements is dependent on at least two of the plurality of information representation elements; wherein each of the two or more information extraction processors (120, 123; 620, 623, 626) of the dialog system is adapted to generate resultant information from the pre-processed information according to an information extraction rule that is specific to the information extraction processor and that is different from the information extraction rule of any other one of the two or more information extraction processors (120, 123; 620, 623, 626), and An output is generated through an output interface (135) of the dialog system based on the information obtained from one or more of the two or more information extraction processors (120, 123; 620, 623, 626), the output being an audio output and / or a text and / or a visual output and / or a signal for operating a machine.

25. A method comprising: obtaining an input representation of the input through an input interface (105) of the dialog system, wherein the input interface (105) obtains the input representation by receiving the input and deriving the input representation from the input, or by receiving the input representation, wherein the input representation is an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements; wherein each of the two or more information extraction processors (120, 123; 620, 623, 626) of the dialog system is adapted to generate resultant information based on the input representation according to an information extraction rule that is specific to the information extraction processor and different from the information extraction rule of any other one of the two or more information extraction processors (120, 123; 620, 623, 626), and generating an output, via an output interface (135) of the dialog system, based on the information obtained from one or more of the two or more information extraction processors (120, 123; 620, 623, 626), the output being an audio output and / or a text and / or a visual output and / or a signal for operating a machine, wherein at least two of the two or more information extraction processors (120, 123; 620, 623, 626) are dialog state dependent, The method comprises selecting, by a dialog system, one or more information extraction processors (120, 123; 620, 623, 626) of at least two information extraction processors (120, 123; 620, 623, 626) on which the dialog state depends, according to the current state of the dialog, so that only those information extraction processors of the at least two information extraction processors (120, 123; 620, 623, 626) associated with the current state of the dialog are selected, and The method includes generating the obtained information by one or more selected information extraction processors (120, 123; 620, 623, 626) according to their information extraction rules.

26. A computer program for implementing the method of claim 24 or 25 when executed on a computer or a signal processor.