Spoken language understanding model training method and device and question answering method and device

By constructing a comparison loss training method between syntax tree and semantic tree, the intent detection and slot recognition accuracy of oral comprehension model is improved, and the problem of insufficient utilization of grammatical and semantic information of existing models is solved.

CN120256552APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410009383.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing oral comprehension model has insufficient accuracy in intent detection and slot recognition, and has failed to effectively utilize the grammatical and semantic information of the statement.

Method used

By building a syntax tree and a semantic tree, compute the comparison loss of syntax features and semantic features, and combine joint training losses to train the oral comprehension model to improve the model's understanding of sentences.

Benefits of technology

The accuracy of the oral comprehension model's intention detection and slot recognition of sentences is improved, and the model's ability to analyze grammatical and semantic information is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256552A_ABST
    Figure CN120256552A_ABST
Patent Text Reader

Abstract

The invention provides a spoken language understanding model training method and device and a question answering method and device. The method comprises the following steps: acquiring training sample data comprising a plurality of sample statements and label data corresponding to the sample statements; inputting the sample statement into a spoken language understanding model for intention prediction and slot recognition processing, and calculating joint training loss according to an output intention prediction result, a slot recognition result and label data; constructing a syntax tree of words in the sample statement, and calculating a first comparison loss according to syntax features corresponding to the syntax tree and word features corresponding to the words; constructing a semantic tree corresponding to each sample statement, and calculating a second comparison loss according to the similarity between the sample statements calculated based on the semantic tree; and calculating target loss according to the first comparison loss, the second comparison loss and the joint training loss, and updating parameters of the spoken language understanding model based on the target loss. The method can improve the accuracy of the spoken language understanding model obtained through training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for training a spoken language understanding model, and a method and apparatus for answering questions. Background Art

[0002] With the continuous development of artificial intelligence technologies, task-based human-machine dialogue systems are used in more and more scenarios in people's lives, and the usage frequency is also increasing continuously. The spoken language understanding task is a basic task in the task-based dialogue system, and this task usually consists of two subtasks: an intent detection task and a slot filling task. The intent detection task is essentially a sentence classification task, and the detected sentence intent can be one or more. The slot value filling task is to identify the slot value for the detected intent.

[0003] Currently, the spoken language understanding task can specifically use a spoken language understanding model to implement intent detection and slot value identification. However, the accuracy of the current spoken language understanding model is not high, resulting in insufficient accuracy in intent detection and slot value identification. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method and apparatus for training a spoken language understanding model, and a method and apparatus for answering questions, and this method can improve the accuracy of the spoken language understanding model.

[0005] According to one aspect of the present disclosure, there is provided a method for training a spoken language understanding model, the method including:

[0006] Obtaining training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, the label data including at least one intent label and a slot label corresponding to the intent label;

[0007] Inputting the sample sentences into a spoken language understanding model for intent prediction and slot identification processing, and calculating a joint training loss according to the intent prediction result, slot identification result output by the spoken language understanding model, and the label data;

[0008] Constructing a syntax tree of words in the sample sentences, and calculating a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, the word features being the features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model;

[0009] Constructing a semantic tree corresponding to each of the sample sentences, and calculating a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree;

[0010] Calculate an objective loss based on the first contrast loss, the second contrast loss, and the joint training loss, and update the parameters of the spoken language understanding model based on the objective loss.

[0011] According to one aspect of the present disclosure, there is provided a training apparatus for a spoken language understanding model, the apparatus including:

[0012] A first obtaining unit, configured to obtain training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and slot labels corresponding to the intent label;

[0013] A first calculating unit, configured to input the sample sentences into the spoken language understanding model for intent prediction and slot recognition processing, and calculate a joint training loss according to the intent prediction result, slot recognition result output by the spoken language understanding model, and the label data;

[0014] A second calculating unit, configured to construct a syntax tree of words in the sample sentences, and calculate a first contrast loss according to the syntactic features corresponding to the syntax tree and the word features corresponding to the words, where the word features are features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model;

[0015] A third calculating unit, configured to construct a semantic tree corresponding to each of the sample sentences, and calculate a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree;

[0016] An updating unit, configured to calculate an objective loss according to the first contrast loss, the second contrast loss, and the joint training loss, and update the parameters of the spoken language understanding model based on the objective loss.

[0017] Optionally, in some embodiments, the second calculating unit includes:

[0018] A first constructing subunit, configured to construct a main syntax tree of the sample sentences;

[0019] A generating subunit, configured to generate at least one positive syntax tree and at least one negative syntax tree of words in the sample sentences according to the connection relationship between nodes in the main syntax tree;

[0020] A first calculating subunit, configured to calculate a first contrast loss according to the syntactic features determined by the at least one positive syntax tree and the at least one negative syntax tree, and the word features corresponding to the words.

[0021] Optionally, in some embodiments, the generating subunit includes:

[0022] An extraction module, configured to extract at least one subtree rooted at any target word in the sample statement according to the connection relationships between nodes in the main syntax tree, so as to obtain at least one positive syntax tree of the target word;

[0023] A replacement module, configured to randomly replace nodes in the at least one positive syntax tree to obtain at least one negative syntax tree of the target word.

[0024] Optionally, in some embodiments, the first calculation sub-unit includes:

[0025] A first generation module, configured to generate at least one positive syntax feature based on the at least one positive syntax tree;

[0026] A second generation module, configured to generate at least one negative syntax feature according to the at least one negative syntax tree;

[0027] A determination module, configured to perform text encoding on the sample statement based on a statement encoder in the spoken language understanding model to obtain a sample statement feature, and determine a word feature corresponding to the word in the sample statement feature;

[0028] A first calculation module, configured to calculate a first contrast loss according to the at least one positive syntax feature, the at least one negative syntax feature, and the word feature.

[0029] Optionally, in some embodiments, the first calculation module includes:

[0030] A first calculation sub-module, configured to calculate a first cosine similarity between the at least one positive syntax feature and the word feature, and calculate a second cosine similarity between the at least one negative syntax feature and the word feature;

[0031] A second calculation sub-module, configured to calculate a first contrast loss according to the first cosine similarity and the second cosine similarity.

[0032] Optionally, in some embodiments, the third calculation unit includes:

[0033] A second construction sub-unit, configured to construct a semantic tree corresponding to each sample statement;

[0034] A second calculation sub-unit, configured to calculate the similarity between sample statements according to the similarity relationship between the node combination relationship and the node connection relationship in the semantic tree;

[0035] A third calculation sub-unit, configured to calculate a second contrast loss according to the similarity between sample statements.

[0036] Optionally, in some embodiments, the second construction sub-unit is further configured to:

[0037] Construct a semantic tree for each of the sample sentences according to the label data corresponding to the sample sentences. The semantic tree includes at least one semantic branch, and each branch includes an intent node, a slot node, and a value node;

[0038] The second calculation subunit is further configured to:

[0039] Generate a plurality of node sets corresponding to each sample sentence according to the combination and connection relationship among the intent node, the slot node, and the value node;

[0040] Calculate the similarity of the node sets corresponding to the sample sentences to obtain a plurality of similarities corresponding to the plurality of node sets among the sample sentences;

[0041] The third calculation subunit is further configured to:

[0042] Calculate a second contrastive loss according to the plurality of similarities.

[0043] Optionally, in some embodiments, the third calculation subunit includes:

[0044] An acquisition module, configured to acquire sample sentence features of each sample sentence;

[0045] A second calculation module, configured to calculate a second contrastive loss according to the sample sentence features and the plurality of similarities.

[0046] Optionally, in some embodiments, the first calculation unit includes:

[0047] A processing subunit, configured to input the sample sentence into a spoken language understanding model for intent prediction and slot recognition processing, and obtain an intent prediction result and a slot recognition result output by the spoken language understanding model;

[0048] A fourth calculation subunit, configured to calculate an intent prediction loss according to the intent prediction result and the intent label corresponding to the sample sentence, and calculate a slot recognition loss according to the slot recognition result and the slot label corresponding to the sample sentence;

[0049] A fifth calculation subunit, configured to calculate the sum of the intent prediction loss and the slot recognition loss to obtain a joint training loss.

[0050] Optionally, in some embodiments, the fourth calculation subunit includes:

[0051] A third calculation module, configured to calculate at least one first cross entropy between the intent prediction result and the at least one intent label, and perform a summation calculation on the at least one first cross entropy to obtain an intent prediction loss;

[0052] A fourth calculation module, configured to calculate at least one second cross entropy between the slot recognition result and the at least one slot label, and perform a summation calculation on the at least one second cross entropy to obtain a slot recognition loss.

[0053] Optionally, in some embodiments, the updating unit includes:

[0054] A sixth calculation subunit, configured to calculate the sum of the first contrast loss and the second contrast loss to obtain a contrast loss sum;

[0055] A second obtaining subunit, configured to obtain a first weight corresponding to the contrast loss sum and a second weight corresponding to the joint training loss;

[0056] A seventh calculation subunit, configured to perform a weighted calculation on the contrast loss sum and the joint training loss based on the first weight and the second weight to obtain a target loss;

[0057] An updating subunit, configured to update the parameters of the spoken language understanding model based on the target loss.

[0058] According to an aspect of the present disclosure, there is provided a question answering method, the method including:

[0059] Obtaining a target question;

[0060] Performing intent detection and slot recognition on the target question by using a spoken language understanding model trained based on the training method of the spoken language understanding model provided by the present disclosure to obtain a target intent and a target slot;

[0061] Performing slot filling on the target slot based on the target intent, and generating answer data corresponding to the target question according to the slot filling result.

[0062] According to an aspect of the present disclosure, there is provided a question answering device, the device including:

[0063] A second obtaining unit, configured to obtain a target question;

[0064] A processing unit, configured to perform intent detection and slot recognition on the target question by using a spoken language understanding model trained based on the training method of the spoken language understanding model provided by the present disclosure to obtain a target intent and a target slot;

[0065] A generating unit, configured to perform slot filling on the target slot based on the target intent, and generate answer data corresponding to the target question according to the slot filling result.

[0066] According to one aspect of the present disclosure, there is provided a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the training method or the question answering method of the spoken language understanding model as described above is implemented.

[0067] According to one aspect of the present disclosure, there is provided a storage medium that stores a computer program, and when the computer program is executed by a processor, the training method or the question answering method of the spoken language understanding model as described above is implemented.

[0068] According to one aspect of the present disclosure, there is provided a computer program product that includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the training method or the question answering method of the spoken language understanding model as described above.

[0069] The training method of the spoken language understanding model provided by the embodiments of the present disclosure includes: obtaining training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and a slot label corresponding to the intent label; inputting the sample sentences into the spoken language understanding model for intent prediction and slot recognition processing, and calculating a joint training loss according to the intent prediction result, the slot recognition result, and the label data output by the spoken language understanding model; constructing a syntax tree of words in the sample sentences, and calculating a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, where the word features are the features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model; constructing a semantic tree corresponding to each sample sentence, and calculating a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree; calculating a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and updating the parameters of the spoken language understanding model based on the target loss.

[0070] In the embodiments of the present disclosure, when jointly training the spoken language understanding model, by designing a first contrast loss guided by a syntax tree and a second contrast loss guided by a semantic tree, and training the spoken language understanding model by combining the first contrast loss and the second contrast loss with the joint training loss, the spoken language understanding model can learn the ability to analyze the syntax and semantics of sentences, and a large amount of guiding information of intents and slot values is also hidden in the syntax and semantic information of the sentences. In this way, the method of this case can improve the understanding ability of the spoken language understanding model for sentences, that is, improve the ability to detect intents and recognize slots of sentences, that is, can greatly improve the accuracy of the spoken language understanding model.

[0071] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present disclosure. The objectives and other advantages of the present disclosure may be realized and attained by the structure particularly pointed out in the specification, claims and drawings. Description of the Drawings

[0072] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0073] Figure 1 A system architecture diagram applied to the training method of the spoken language understanding model of the embodiment of the present disclosure;

[0074] Figure 2a A schematic diagram of the embodiment of the present disclosure applied to the AI diagnosis guidance scenario;

[0075] Figure 2b A schematic diagram of the embodiment of the present disclosure applied to the book search scenario;

[0076] Figure 3 A flowchart of the training method of the spoken language understanding model provided by the present disclosure;

[0077] Figure 4 A schematic diagram of the model structure of the spoken language understanding model to be trained in the present disclosure;

[0078] Figure 5 Another flowchart of the training method of the spoken language understanding model provided by the present disclosure;

[0079] Figure 6 A schematic diagram of the main concept of the spoken language understanding model provided by the present disclosure;

[0080] Figure 7 A schematic diagram of the syntax tree of the example sample sentence;

[0081] Figure 8 A schematic diagram of the semantic tree of the example sample sentence;

[0082] Figure 9 A schematic diagram of the subtree corresponding to a word in the syntax tree of the example sample sentence;

[0083] Figure 10 For Figure 9 A schematic diagram of the negative tree corresponding to the shown subtree;

[0084] Figure 11 A flowchart of the question answering method provided by the present disclosure;

[0085] Figure 12Structural schematic diagram of the training device for the spoken language understanding model provided by the embodiments of the present disclosure;

[0086] Figure 13 Structural schematic diagram of the question answering device provided by the embodiments of the present disclosure;

[0087] Figure 14 It is the terminal structure diagram for implementing the various methods according to an embodiment of the present disclosure;

[0088] Figure 15 It is the server structure diagram for implementing the various methods according to an embodiment of the present disclosure. Detailed implementation manners

[0089] In order to make the objectives, technical solutions and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0090] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0091] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that perceives the environment, acquires knowledge and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to the downstream tasks of various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0092] Spoken Language Understanding (SLU): The language understanding technology in a dialogue system generally includes Natural Language Understanding (NLU) or Spoken Language Understanding. The work of NLU is to output a corresponding semantic structured representation given a question. The work of SLU is to output the intent information and slot information in a given statement.

[0093] Intent Detection: Also known as intent classification, it refers to identifying the intent corresponding to the question input by an object in an intelligent question-answering system. For an object question, one or more pre-defined intent category labels can be assigned.

[0094] Contrastive Learning: It is a machine learning method aimed at learning the similarities and differences in data. The basic idea of contrastive learning is to divide data samples into positive samples and negative samples, and learn the representation of data by maximizing the similarity between positive samples and minimizing the similarity between negative samples.

[0095] In related technologies, for the intent detection task and slot recognition task in the spoken language understanding task, two independent models are used for training to obtain two models with the ability to detect the intent of a statement and the ability to recognize slots in a statement respectively. Then, the two models are used jointly to complete the spoken language understanding task. However, since there is also a strong correlation between the intent and slots of the same statement, the two independently trained models cannot learn the correlation between these two tasks, resulting in inaccurate intent detection results and slot recognition results when using two independent models to complete the spoken language understanding task. In response to this, those skilled in the art have proposed a joint training method, that is, integrating an intent detection branch and a slot recognition branch in the same model. When training the model, different-dimensional labels corresponding to the same sample can be used to jointly train these two branches, so that the spoken language recognition model can learn the relationship between these two task branches, and further improve the accuracy of the output of the spoken language understanding model.

[0096] However, the inventors of the present application found in their research that when jointly training a spoken language understanding model in the existing solutions, the grammar and semantic information in the sentences are ignored. The grammar information in the sentences also plays an important role in the process of natural language processing, and there is also a natural semantic structure in the sentences, that is, intent-slot-value, which describes the complex relationship between the intent and the slot. The spoken language understanding model fails to learn the grammar and semantic information of the sentences during joint training, resulting in the accuracy of the trained spoken language understanding model not being high enough. Based on this, the present disclosure provides a training method for a spoken language understanding model, in order to enable the spoken language understanding model to learn the ability to extract the grammar and semantic information of the sentences during the training process, and then improve the ability of the trained spoken language understanding model to understand the sentences, that is, to improve the accuracy of the spoken language understanding model.

[0097] System architecture and scenario description applied in the embodiments of the present disclosure

[0098] Figure 1 It is a system architecture diagram to which the training method of the spoken language understanding model according to the embodiments of the present disclosure is applied. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.

[0099] The terminal 140 includes various forms such as a desktop computer, a laptop computer, a PDA (Personal Digital Assistant), a mobile phone, a vehicle-mounted terminal, a home theater terminal, a dedicated terminal, a smart voice interaction device, a smart home appliance, an aircraft, etc. In addition, it can be a single device or a set of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0100] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with an ordinary terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc.

[0101] The gateway 120 is also called an internetwork connector and a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion function. Between two systems that use different communication protocols, data formats or languages, and even have completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 needs to be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 also needs to be sent to the corresponding terminal 140 through the gateway 120.

[0102] The training method of the spoken language understanding model provided by the embodiments of the present disclosure can be implemented independently in the aforementioned terminal 140, can be implemented independently in the aforementioned server 110, or can be partially implemented in the terminal 140 and partially implemented in the server 110.

[0103] When the training method of the spoken language understanding model provided by the embodiments of the present disclosure is implemented in the terminal 140, the terminal 140 can obtain training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and a slot label corresponding to the intent label; the terminal 140 inputs the sample sentences into the spoken language understanding model for intent prediction and slot recognition processing, and calculates a joint training loss according to the intent prediction result, slot recognition result, and label data output by the spoken language understanding model; the terminal 140 constructs a syntax tree of the words in the sample sentences, and calculates a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, where the word features are the features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model; the terminal 140 constructs a semantic tree corresponding to each sample sentence, and calculates a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree; the terminal 140 calculates a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and updates the parameters of the spoken language understanding model based on the target loss.

[0104] When the training method of the spoken language understanding model provided by the embodiments of the present disclosure is implemented in the server 110, the server 110 can obtain training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and a slot label corresponding to the intent label; the server 110 inputs the sample sentences into the spoken language understanding model for intent prediction and slot recognition processing, and calculates a joint training loss according to the intent prediction result, slot recognition result, and label data output by the spoken language understanding model; the server 110 constructs a syntax tree of the words in the sample sentences, and calculates a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, where the word features are the features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model; the server 110 constructs a semantic tree corresponding to each sample sentence, and calculates a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree; the server 110 calculates a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and updates the parameters of the spoken language understanding model based on the target loss.

[0105] When a part of the training method of the spoken language understanding model provided by the embodiments of the present disclosure is implemented in the server 110 and another part is implemented in the terminal 140, the terminal 140 can obtain training sample data including a plurality of sample sentences and label data corresponding to the sample sentences. The label data includes at least one intent label and slot labels corresponding to the intent label. Then, the terminal 140 sends the obtained training sample data to the server 110 for training the spoken language understanding model. The server 110 inputs the sample sentences into the spoken language understanding model for intent prediction and slot recognition processing, and calculates a joint training loss according to the intent prediction result, slot recognition result, and label data output by the spoken language understanding model. The server 110 constructs a syntax tree of the words in the sample sentences, and calculates a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words. The word features are the features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model. The server 110 constructs a semantic tree corresponding to each sample sentence, and calculates a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree. The server 110 calculates a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and updates the parameters of the spoken language understanding model based on the target loss.

[0106] The training method of the spoken language understanding model provided by the embodiments of the present disclosure can be specifically applied to training the spoken language understanding model used in different task-based human-computer dialogue scenarios. For example, it can be used to train the spoken language understanding model used in the AI medical guidance scenario, and can also be used to train the spoken language understanding model used in the book search scenario.

[0107] As Figure 2a shown, it is a schematic diagram of the application of the embodiments of the present disclosure in the AI medical guidance scenario. When a patient asks a consultation question using an AI medical guidance application, the spoken language understanding model carried by the application can detect intent information and slot information from the question. Then, according to the knowledge pre-stored in the system, the slot is filled, and the answer corresponding to the question asked by the patient is generated and displayed. For example, when a patient asks "Where is the pediatrics department?", it can be recognized that the intent is to ask for directions and the slot is the location. Then, the location information of the pediatrics department can be searched in the knowledge base, and the answer "The pediatrics department is on the fourth floor of Dongchuan Clinic" is generated, and then the answer can be displayed on the AI medical guidance application interface.

[0108] As Figure 2bAs shown, it is a schematic diagram of the application of the embodiments of the present disclosure in the book search scenario. When a reader asks a question using a book search application, the spoken language understanding model incorporated in the book search application can detect intent information and slot information from the question. Then, based on the knowledge pre-stored in the system, the slot is filled, and an answer corresponding to the reader's consultation question is generated and displayed. For example, when a reader consults "Where is Three of Us?", it can be recognized that the intent is to find the storage location of the book, and the slot is the location. Then, the storage location of the book "Three of Us" can be found in the book storage record, and an answer "Three of Us is on the third layer of Shelf A in the Literature Area on the Second Floor of the Library" is generated. Then, this answer can be displayed on the interface of the book search application.

[0109] General description of the embodiments of the present disclosure

[0110] According to an embodiment of the present disclosure, a method for training a spoken language understanding model is provided. This method can be used in the aforementioned AI medical guidance scenario or book search scenario, and can also be used in other task-based human-machine dialogue scenarios to train the spoken language understanding models used in these scenarios.

[0111] As Figure 3 shown, it is a schematic flowchart of a method for training a spoken language understanding model provided by the present disclosure. This method can be applied to a training device for a spoken language understanding model, and this training device for a spoken language understanding model can be integrated in a computer device, and the computer device can specifically be a terminal or a server. The method for training a spoken language understanding model may include:

[0112] Step 310, obtaining training sample data including a plurality of sample sentences and label data corresponding to the sample sentences.

[0113] Among them, the method for training a spoken language understanding model provided in the embodiments of the present disclosure is used to train a spoken language understanding model in a joint training paradigm. A spoken language understanding model in a joint training paradigm generally includes several modules such as a sentence encoder, a sentence decoder, an intent detection branch, and a slot value recognition branch. As Figure 4 shown, it is a schematic diagram of the model structure of the spoken language understanding model to be trained in the present disclosure. As shown in the figure, the spoken language understanding model provided in the embodiments of the present disclosure includes a sentence encoder 410, a sentence decoder 420, an intent detection layer 430, and a slot recognition layer 440.

[0114] The sentence encoder 410 is used to encode the input sentence to obtain a sentence-level representation, that is, to obtain sentence features. Among them, since a sentence is composed of multiple words, the word features corresponding to each word can be further determined according to the position information of each word in the sentence and the sentence features. Among them, in the embodiments of the present disclosure, the sentence encoder may specifically adopt a self-attention encoder, that is, a self-attention encoder can be used to generate a series of hidden states of the input sentence. Then, H can be input into the max pooling layer to obtain the sentence-level representation t of the input sentence. Alternatively, in the embodiments of the present disclosure, a pre-trained language model encoder (such as Transformer or BERT) can also be used as the sentence encoder. In this embodiment, the input sentence can be input into the pre-trained language model encoder to obtain the hidden state H, and then the hidden state corresponding to the cls token in the hidden state H is output as the sentence-level representation t. Among them, the input sentence can be an input sentence in speech form or an input sentence in text form. When the input sentence is an input sentence in speech form, speech recognition can be performed on the input sentence first to convert it into the corresponding text.

[0115] The sentence decoder 420 is used to decode the sentence-level representation t, and then input the decoding result into the intent detection layer 430 and the slot recognition layer 440 for intent detection and slot recognition.

[0116] When training the spoken language understanding model with the above joint training paradigm, the training sample data for training the model can be obtained first. Among them, the training sample data includes multiple sample sentences and the label data corresponding to each sample sentence. Since the joint training paradigm is adopted, the label data of each sample sentence can include at least one intent label and the slot labels corresponding to each intent label. In the embodiments of the present disclosure, the model structure of the trained spoken language understanding model and the training sample data adopted are the same as those of the joint training paradigm. The training method of the spoken language understanding model provided by the present disclosure enables the spoken language understanding model to learn the ability to recognize the grammatical information and semantic information of sentences during the training process by optimizing the algorithm during the training process, thereby improving the accuracy of the spoken language understanding model in the intent detection task and the slot recognition task. In the following, the training method of the spoken language understanding model provided by the present disclosure will be further introduced in detail.

[0117] Step 320: Input the sample sentence into the spoken language understanding model for intent prediction and slot recognition processing, and calculate the joint training loss according to the intent prediction result, slot recognition result and label data output by the spoken language understanding model.

[0118] In the embodiments of the present disclosure, after obtaining the training sample data for training the spoken language understanding model, the sample sentences in the training sample data can be first input into the spoken language understanding model to be trained for intention prediction processing and slot recognition processing. The spoken language understanding model will output an intention prediction result and a slot recognition result. At this time, the joint training loss can be constructed based on the intention prediction result and the intention label included in the label data, as well as the slot recognition result and the slot label included in the label data. The role of the joint training loss is to make the intention prediction result output by the spoken language recognition model closer to the corresponding intention label and the slot recognition result output by the spoken language understanding model closer to the slot label in the label data during the continuous training process of the spoken language understanding model.

[0119] In some embodiments, inputting the sample sentence into the spoken language understanding model for intention prediction and slot recognition processing, and calculating the joint training loss according to the intention prediction result, slot recognition result output by the spoken language understanding model, and label data, includes:

[0120] Input the sample sentence into the spoken language understanding model for intention prediction and slot recognition processing to obtain the intention prediction result and slot recognition result output by the spoken language understanding model;

[0121] Calculate the intention prediction loss according to the intention prediction result and the intention label corresponding to the sample sentence, and calculate the slot recognition loss according to the slot recognition result and the slot label corresponding to the sample sentence;

[0122] Calculate the sum of the intention prediction loss and the slot recognition loss to obtain the joint training loss.

[0123] In the embodiments of the present disclosure, when calculating the joint training loss, the intention prediction loss can be calculated based on the intention detection result and the intention label respectively, and the slot recognition loss can be calculated based on the slot recognition result and the slot label. Then, calculate the sum of the intention prediction loss and the slot recognition loss to obtain the joint training loss. That is, L U = L I + L S , where L U is the joint training loss, L I is the intention prediction loss, and L S is the slot recognition loss.

[0124] In some embodiments, when calculating the joint training loss, certain weights can also be set for the intention prediction loss and the slot recognition loss to adjust the importance of the intention prediction and slot recognition tasks in different application scenarios.

[0125] In some embodiments, an intent prediction loss is calculated based on the intent prediction result and the intent label corresponding to the sample statement, and a slot recognition loss is calculated based on the slot recognition result and the slot label corresponding to the sample statement, including:

[0126] Calculate at least one first cross-entropy between the intent prediction result and at least one intent label, and sum up the at least one first cross-entropy to obtain the intent prediction loss;

[0127] Calculate at least one second cross-entropy between the slot recognition result and at least one slot label, and sum up the at least one second cross-entropy to obtain the slot recognition loss.

[0128] Specifically, in the embodiments of the present disclosure, calculating the intent prediction loss based on the intent prediction result and the corresponding intent label can be achieved by calculating the cross-entropy between the intent prediction result and the corresponding intent label. And in the disclosed embodiments, since the intent prediction result may include multiple intents, and the intent label may also be multiple labels. When there are multiple intents in the intent prediction result, the cross-entropy between each intent and the corresponding intent label can be calculated to obtain multiple cross-entropies, which are referred to as the first cross-entropies here. Then, the sum of these multiple first cross-entropies can be calculated to obtain the intent prediction loss.

[0129] Similarly, in the embodiments of the present disclosure, the slot recognition result output by the spoken language understanding model may also include multiple slots, and then the second cross-entropy between each slot and the corresponding slot label can be calculated to obtain multiple second cross-entropies. Further, the sum of these multiple second cross-entropies can be calculated to obtain the slot recognition loss.

[0130] Step 330, construct a syntax tree of the words in the sample statement, and calculate a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words.

[0131] In the embodiments of the present disclosure, in addition to calculating the joint training loss based on the intent detection result and the slot recognition result output by inputting the training sample data into the spoken language understanding model and the label data, a first contrast loss for enabling the spoken language understanding model to learn to extract the syntax information of the input statement can be further constructed. Specifically, a syntax tree corresponding to the words in the sample statement can be constructed, and then the syntax features of each word can be generated according to the syntax tree of the word, and the loss is calculated based on the syntax features of the word and the word features corresponding to the word in the sentence representation input to the sentence encoder. In the embodiments of the present disclosure, the loss calculated based on the syntax features of the word and the word features corresponding to the word in the sentence representation input to the sentence encoder is specifically a contrast loss, which can be referred to as the first contrast loss here.

[0132] In some embodiments, a syntax tree of words in a sample sentence is constructed, and a first contrastive loss is calculated based on the syntactic features corresponding to the syntax tree and the word features corresponding to the words, including:

[0133] Construct the main syntax tree of the sample sentence;

[0134] Generate at least one positive syntax tree and at least one negative syntax tree of the words in the sample sentence according to the connection relationships between the nodes in the main syntax tree;

[0135] Calculate the first contrastive loss based on the syntactic features determined by at least one positive syntax tree and at least one negative syntax tree and the word features corresponding to the words.

[0136] In the embodiments of the present disclosure, when constructing the syntax tree of the words in the sample sentence, the overall main syntax tree of the sample sentence can be constructed first. Then, according to the connection relationships between the nodes in the main syntax tree of the sample sentence, the syntax tree of each word in the sample sentence is generated. Among them, in the embodiments of the present disclosure, a well-trained Stanza (a natural language analysis model) can be used to construct the main syntax tree corresponding to the sample sentence. The main syntax tree includes multiple nodes corresponding to multiple words in the sample sentence and the connection relationships between the nodes.

[0137] In the embodiments of the present disclosure, in order to enable the spoken language understanding model to learn the ability to understand the syntactic information of sentences by using the method of contrastive learning, when generating the syntax tree of each word according to the main syntax tree of the sample sentence, at least one positive syntax tree and at least one negative syntax tree corresponding to each word can be generated, and then the first contrastive loss is constructed by maximizing the similarity between the features corresponding to the positive syntax tree and the word features of the word, and minimizing the similarity between the features corresponding to the negative syntax tree and the word features of the word.

[0138] In some embodiments, generating at least one positive syntax tree and at least one negative syntax tree of the words in the sample sentence according to the connection relationships between the nodes in the main syntax tree includes:

[0139] Extract at least one subtree with any target word in the sample sentence as the root node according to the connection relationships between the nodes in the main syntax tree to obtain at least one positive syntax tree of the target word;

[0140] Randomly replace the nodes in at least one positive syntax tree to obtain at least one negative syntax tree of the target word.

[0141] In the embodiments of the present disclosure, in the specific process of constructing at least one positive syntax tree and at least one negative syntax tree corresponding to each word in the sample sentence according to the main syntax tree of the sample sentence, it may be first to use each word in the sample sentence as the root node of the node in the main syntax tree, extract other nodes associated with the node to construct at least one subtree, and the subtree can be called the positive syntax tree corresponding to the word.

[0142] Further, partial random replacement can be performed on the nodes in each positive syntax tree. For example, the word corresponding to the node can be replaced with other words, so as to obtain at least one negative syntax tree corresponding to each positive syntax tree.

[0143] In some embodiments, calculating the first contrast loss based on the syntax features determined by at least one positive syntax tree and at least one negative syntax tree and the word features corresponding to the words includes:

[0144] Generating at least one positive syntax feature based on at least one positive syntax tree;

[0145] Generating at least one negative syntax feature according to at least one negative syntax tree;

[0146] Performing text encoding on the sample sentence based on the sentence encoder in the speech understanding model to obtain the sample sentence feature, and determining the word feature corresponding to the word in the sample sentence feature;

[0147] Calculating the first contrast loss according to at least one positive syntax feature, at least one negative syntax feature and the word feature.

[0148] As introduced above, in the embodiments of the present disclosure, the first contrast loss can be constructed by maximizing the similarity between the feature corresponding to the positive syntax tree and the word feature of the word, and minimizing the similarity between the feature corresponding to the negative syntax tree and the word feature of the word. Therefore, after generating at least one positive syntax tree and at least one negative syntax tree corresponding to each word in the sample sentence, the features corresponding to each positive syntax tree and each negative syntax tree can be further generated. Specifically, a positive syntax feature can be generated according to each positive syntax tree, and a negative syntax feature can be generated according to each negative syntax tree.

[0149] Specifically, to generate the positive syntax feature corresponding to the word according to the positive syntax tree, the word features of the words corresponding to each node in the positive syntax tree can be obtained first. The word features can be specifically determined from the aforementioned sentence representation, and then the corresponding positive syntax feature can be calculated according to the word features of the words corresponding to each node in the positive syntax tree. Correspondingly, the corresponding method can also be used to calculate the negative syntax feature corresponding to each negative syntax tree.

[0150] Further, a first contrast loss is constructed based on the word features of each word in the sample sentence, the corresponding at least one positive syntactic feature, and the at least one negative syntactic feature. The first contrast loss aims to control the maximization of the similarity between the word features of each word in the sample sentence and the corresponding at least one positive syntactic feature, and the minimization of the similarity between the word features of each word in the sample sentence and the corresponding at least one negative syntactic feature, so that the spoken language understanding model can learn the ability to extract syntactic information from the sentence.

[0151] In some embodiments, calculating the first contrast loss according to at least one positive syntactic feature, at least one negative syntactic feature, and word features includes:

[0152] Calculating a first cosine similarity between at least one positive syntactic feature and the word feature, and calculating a second cosine similarity between at least one negative syntactic feature and the word feature;

[0153] Calculating the first contrast loss according to the first cosine similarity and the second cosine similarity.

[0154] Wherein, in the embodiments of the present disclosure, the similarity between the word feature and the syntactic feature can specifically be evaluated by using the cosine similarity between the features. That is, the first pre-similarity between at least one positive syntactic feature and the word feature can be calculated first, and then the second cosine similarity between at least one negative syntactic feature and the word feature can be calculated. Then, taking the first cosine similarity as the numerator and the second cosine similarity as the denominator and adding them to a preset or general contrast loss function to obtain the first contrast loss.

[0155] Step 340, constructing a semantic tree corresponding to each sample sentence, and calculating a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree.

[0156] In the embodiments of the present disclosure, in addition to constructing the first contrast loss guided by the syntax tree to improve the ability of the spoken language understanding model to extract syntactic information from the sentence, a semantic tree can be further constructed, and a second contrast loss guided by semantics can be constructed based on the similarity between the semantic trees corresponding to different sentences. By controlling the training process of the spoken language understanding model based on the second contrast loss, the spoken language understanding model can further learn the ability to extract the semantics of the sentence.

[0157] Specifically, a semantic tree corresponding to each sample sentence can be constructed first, then the similarity between the sample sentences can be calculated according to the semantic tree corresponding to each sample sentence, and further the second contrast loss can be calculated according to the similarity between the sample sentences.

[0158] In some embodiments, constructing a semantic tree corresponding to each sample sentence, and calculating a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree includes:

[0159] Construct a semantic tree corresponding to each sample statement;

[0160] Calculate the similarity between sample statements according to the similarity relationship between the node combination relationships and node connection relationships in the semantic tree;

[0161] Calculate a second contrast loss according to the similarity between sample statements.

[0162] In the embodiments of the present disclosure, after constructing a semantic tree corresponding to each sample statement, multiple node combinations and multiple node connection relationships can be determined according to the nodes in the semantic tree of the sample statement, and then the similarity relationship between the node combinations and node connection relationships is obtained, and the similarity between sample statements is calculated according to the similarity relationship between the node combinations or node connection relationships.

[0163] In some embodiments, constructing a semantic tree corresponding to each sample statement includes:

[0164] Construct a semantic tree for each sample statement according to the label data corresponding to the sample statement, the semantic tree includes at least one semantic branch, and each branch includes an intent node, a slot node, and a value node;

[0165] Calculating the similarity between sample statements according to the similarity relationship between the node combination relationships and node connection relationships in the semantic tree includes:

[0166] Generate multiple node sets corresponding to each sample statement according to the combination and connection relationships between the intent nodes, slot nodes, and value nodes;

[0167] Calculate the similarity of the corresponding node sets between sample statements to obtain multiple similarities corresponding to the multiple node sets between sample statements;

[0168] Calculating a second contrast loss according to the similarity between sample statements includes:

[0169] Calculate a second contrast loss according to multiple similarities.

[0170] In the embodiments of the present disclosure, a semantic tree for each sample statement can be constructed according to the label data corresponding to the sample statement. The constructed semantic tree of the sample statement can include multiple semantic branches, and each semantic branch contains three layers, namely an intent layer, a slot layer, and a value layer. That is, each semantic branch includes an intent node, a slot node (or slot position node), and a value node. Among them, the intent node, slot node, and value node can form various combinations with connection relationships, and the connection relationship here can specifically be a node path. Specifically, it can include three paths: the path from the intent node to the slot node, the path from the slot node to the value node, and the path from the intent node to the slot node and then to the value node.

[0171] Then, multiple node sets corresponding to the sample sentences can be generated according to the combination and connection relationships among the intent nodes, slot nodes, and value nodes, specifically including an intent node set, a slot node set, a value node set, a node set corresponding to the path from the intent node to the slot node, a node set corresponding to the path from the slot node to the value node, and a node set corresponding to the path from the intent node to the slot node and then to the value node. That is, node sets in six dimensions corresponding to each sample sentence can be generated.

[0172] Furthermore, when calculating the similarity between sample sentences, the six-dimensional similarity between sample sentences can be calculated based on these six-dimensional node sets. Then, the second contrast loss is calculated based on the six-dimensional similarity.

[0173] In some embodiments, calculating the second contrast loss according to multiple similarities includes:

[0174] Obtain the sample sentence features of each sample sentence;

[0175] Calculate the second contrast loss according to the sample sentence features and multiple similarities.

[0176] Among them, in the embodiments of the present disclosure, after calculating the similarity between sample sentences, the positive samples similar to the sample sentences and the negative samples dissimilar to them in the training batch can be determined. When there are multiple similarities, the positive samples similar to the sample sentences in some dimensions and the negative samples dissimilar to them in some dimensions in the training batch can be determined. For the similar positive samples, it is necessary to control the maximum similarity of the sentence representations of these two sample sentences in this dimension; for the dissimilar negative samples, it is necessary to control the minimum similarity of the sentence representations of these two sample sentences in this dimension.

[0177] Therefore, before calculating the second contrast loss based on multiple similarities, the sample sentence features of each sample sentence can be obtained first, and then the second contrast loss is calculated according to the sample sentence features and multiple similarities.

[0178] Step 350, calculate the target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and update the parameters of the spoken language understanding model based on the target loss.

[0179] Among them, after calculating the joint training loss, the first contrast loss, and the second contrast loss, the target loss can be calculated based on the first contrast loss, the second contrast loss, and the joint training loss, and further the model parameters of the spoken language understanding model are updated based on the target loss.

[0180] In some embodiments, a target loss is calculated based on a first contrastive loss, a second contrastive loss, and a joint training loss, and the parameters of the spoken language understanding model are updated based on the target loss, including:

[0181] Calculate the sum of the first contrastive loss and the second contrastive loss to obtain a contrastive loss sum;

[0182] Obtain a first weight corresponding to the contrastive loss sum and a second weight corresponding to the joint training loss;

[0183] Perform weighted calculation on the contrastive loss sum and the joint training loss based on the first weight and the second weight to obtain a target loss;

[0184] Update the parameters of the spoken language understanding model based on the target loss.

[0185] In the embodiments of the present disclosure, after calculating the first contrastive loss and the second contrastive loss, the sum of the first contrastive loss and the second contrastive loss can be further calculated to obtain a contrastive loss sum. Then, a first weight corresponding to the contrastive loss sum and a second weight corresponding to the joint training loss can be obtained, and weighted calculation is performed on the above contrastive loss sum and the joint training loss based on the first weight and the second weight, so as to obtain a target loss for training the spoken language understanding model.

[0186] After calculating the target loss, the backpropagation gradient can be further calculated according to the target loss, and the model parameters of the spoken language understanding model are updated based on the backpropagation gradient. Then, a new batch of training sample data can be re-obtained, and the above method is repeated to perform cyclic training on the spoken language understanding model until the parameters of the spoken language understanding model converge, and a trained spoken language understanding model is obtained.

[0187] In the embodiments of the present disclosure, by obtaining training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, the label data includes at least one intent label and a slot label corresponding to the intent label; inputting the sample sentences into the spoken language understanding model for intent prediction and slot recognition processing, and calculating a joint training loss according to the intent prediction result, the slot recognition result, and the label data output by the spoken language understanding model; constructing a syntax tree of the words in the sample sentences, and calculating a first contrastive loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, where the word features are the features corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model; constructing a semantic tree corresponding to each sample sentence, and calculating a second contrastive loss according to the similarity between the sample sentences calculated based on the semantic tree; calculating a target loss according to the first contrastive loss, the second contrastive loss, and the joint training loss, and updating the parameters of the spoken language understanding model based on the target loss.

[0188] In the embodiments of the present disclosure, when jointly training a spoken language understanding model, by designing a first contrast loss guided by a syntax tree and a second contrast loss guided by a semantic tree, and combining the first contrast loss and the second contrast loss with the joint training loss to train the spoken language understanding model, the spoken language understanding model can learn the ability to analyze the syntax and semantics of sentences. A large amount of guiding information for intentions and slot values is also hidden in the syntax and semantic information of sentences. Thus, the method in this case can improve the understanding ability of the spoken language understanding model for sentences, that is, improve the ability to detect intentions and identify slots in sentences, and can greatly improve the accuracy of the spoken language understanding model.

[0189] Detailed description of the embodiments of the present disclosure in combination with specific application scenarios

[0190] As Figure 5 shown, it is another schematic flowchart of the training method of the spoken language understanding model provided by the present disclosure. In this embodiment, the training method of the spoken language understanding model will be introduced in detail in combination with the execution subject of each step. The method specifically includes the following steps:

[0191] Step 501, the computer device obtains batch sample data for training the spoken language understanding model.

[0192] Among them, the embodiments of the present disclosure specifically provide a contrastive learning method for a spoken language understanding model based on syntax tree guidance and semantic tree guidance. Specifically, as Figure 6 shown, it is a schematic diagram of the main idea of the present disclosure for the spoken language understanding model. As shown in the figure, for each sample sentence in the batch sample data, in addition to inputting it into the joint training model to calculate the intention detection loss and the slot recognition loss, the embodiments of the present disclosure can also calculate the syntax contrast loss and the semantic contrast loss respectively. Then, the model parameters of the speech recognition model are adjusted as a whole by combining the intention detection loss, the slot recognition loss, the syntax contrast loss, and the semantic contrast loss. Specifically, in the process of calculating the syntax contrast loss, the syntax tree of the sample sentence can be generated first, and then the word-level positive and negative syntax trees can be generated according to the syntax tree. Further, the syntax contrast loss is calculated by combining the word features in the sentence representation with the word-level positive and negative syntax trees. In the process of calculating the semantic contrast loss, the semantic tree of each sample sentence can be constructed first, and then a multi-perspective scoring set can be generated according to the relationship between the nodes in the semantic tree. Further, the semantic contrast loss can be calculated according to the sentence representation of the sample sentence and the multi-perspective scoring set.

[0193] Next, based on Figure 6The framework introduces in detail the training method of the spoken language understanding model provided by the present disclosure. First, before training the spoken language understanding model, the computer device can first obtain batch sample data for training the spoken language understanding model. The batch sample data includes multiple sample data, and each sample data contains a sample sentence and label data corresponding to the sample sentence. The label data can specifically include an intent label and a slot label.

[0194] Step 502, the computer device inputs the sample sentence into the spoken language understanding model to obtain the sentence feature of the sample sentence, the word feature of each word in the sample sentence, as well as the intent prediction result and the slot recognition result.

[0195] After obtaining the batch sample data, the computer device can start model training based on the batch sample data. Specifically, the sample sentence in each sample data can be input into the spoken language understanding model, and the spoken language understanding model extracts features from each sample sentence to obtain the sentence feature of each sample sentence. Since the sample sentence is composed of multiple words, the word feature corresponding to each word can also be determined. For example, when there is a sample sentence "what is fare code h and show me the cheapest fare" (What is the fare code and show me the cheapest fare). Then when this sample sentence is input into the spoken language understanding model, the spoken language understanding model can not only extract the sentence feature of this sentence, but also extract the word feature of each word (such as show) in it. The English sample sentence here is just an example, and sample sentences in Chinese or other languages can also be used, which does not limit the solution of the present disclosure.

[0196] In addition, the spoken language understanding model can also perform intent prediction on the input sample sentence to obtain an intent prediction result, and can also perform slot recognition to obtain a slot recognition result.

[0197] Step 503, the computer device calculates the intent prediction loss and the slot recognition loss according to the intent prediction result and the slot recognition result, as well as the intent label and the slot label.

[0198] After the computer device obtains the intent prediction result output by the spoken language understanding model, it can calculate the intent prediction loss according to the intent prediction result and the intent label. The specific calculation formula is as follows:

[0199]

[0200] where n I represents the number of intent labels or intent prediction results, represents the intent prediction result, represents the intent label.

[0201] Further, after the computer device obtains the slot recognition results output by the spoken language understanding model, it can calculate the slot recognition loss based on the slot recognition results and slot labels. The specific calculation formula is as follows:

[0202]

[0203] where n S represents the number of slot labels or slot recognition results, represents the slot recognition result, represents the slot label.

[0204] Step 504: The computer device constructs the syntax tree and semantic tree for each sample sentence.

[0205] As previously introduced, in the embodiments of the present disclosure, it is not only necessary to calculate the intent prediction loss and slot recognition loss, but also necessary to calculate the syntax comparison loss and semantic comparison loss. Therefore, the computer device needs to further construct the syntax tree and semantic tree corresponding to each sample sentence.

[0206] Among them, the well-trained parsing model Stanza can be used to generate the syntax tree of each sample sentence. As Figure 7 shown, it is a schematic diagram of the syntax tree of the foregoing example sample sentence. In this syntax tree, the word "what" is used as the root node and jointly forms the syntax tree of the sample sentence with all other words dominated by this word.

[0207] In addition, in the embodiments of the present disclosure, when constructing the semantic tree of the sample sentence, the label data corresponding to the sample sentence is used. The semantic tree includes multiple semantic branches, and each semantic branch includes three layers of nodes: the intent layer, the slot layer, and the value layer. The intent node, slot node, and value node are nodes determined according to the corresponding label data. As Figure 8 shown, it is a schematic diagram of the semantic tree of the foregoing example sample sentence. As shown in the figure, the semantic tree of the foregoing example sample sentence includes two semantic branches, and each semantic branch contains three layers of nodes. The intent node of the first semantic branch is atis_abbreviation (abbreviation), the slot node is h, and the value node is B-fare_basic_code (fare code).

[0208] Step 505: The computer device constructs the positive tree and negative tree for each word according to the syntax tree of each sample sentence.

[0209] Further, for each word in the sample sentence, its subtree can be derived from the syntax tree of the sample sentence and regarded as the positive tree of this word, denoted as T + . As Figure 9As shown, it is a schematic diagram of the subtree corresponding to a word in the syntax tree of the foregoing sample sentence. Specifically, this subtree is the one in the syntax tree with the word "show" as the root node and includes all other words dominated by this word as the nodes of this subtree.

[0210] Then, the computer device can randomly replace some nodes in the positive tree T + to obtain the negative tree corresponding to this word. As Figure 10 shown, it is a schematic diagram of the negative tree corresponding to the subtree shown in Figure 9 . As shown in the figure, according to the syntax tree in Figure 7 , h is not a word dominated by "show" and should not appear in its corresponding subtree. Therefore, this tree is the negative tree of the word "show". Among them, in the embodiments of the present disclosure, when replacing the node labels in the positive tree corresponding to a word, no more than 3 node labels can be replaced to ensure that at least one node label in the positive tree and the negative tree of the word is the same.

[0211] Step 506, the computer device calculates the syntax comparison loss according to the positive tree and negative tree of each word and the word feature of each word.

[0212] After constructing the positive tree and negative tree corresponding to each word in the sample sentence and obtaining the word feature of each word in the sample sentence, the computer device can calculate the syntax comparison loss according to the positive tree and negative tree of each word and the word feature of each word. The specific calculation formula is as follows:

[0213]

[0214] Among them, h i is the word feature of the current word (the i-th), t j is the word in the positive tree of the current word, and h j is the word feature of the word t j . M represents the total number of the positive tree and negative tree of the current word, and T m represents any one of the positive tree or negative tree of the current word. t′j 表 represents the word in the tree T m of the current word. Among them, e i,j is calculated according to the following formula:

[0215]

[0216] Among them, h k is the other word in the tree T m of the current word except the current word. e i,j’ can be calculated using the same formula as e i,j .

[0217] Step 507, the computer device constructs a set of scoring nodes corresponding to each sample statement according to the semantic tree of each sample statement.

[0218] Among them, after constructing the semantic tree corresponding to each sample statement, different sets of nodes can be constructed according to the nodes at different layers. For example, the set of intent layer nodes I = {atis_abbreviation, atis_cheapest}, the set of slot layer nodes S = {h, cheapest}, and the set of value layer nodes V = {B - fare_basis_code, B - cost_relative}.

[0219] In addition, sets can also be constructed for all possible paths between two grounding nodes, denoted as IS, SV, and ISV. For example, for the aforementioned example, ISV = {atis_abbreviation -> h -> B - fare_basis_code, atis_cheapest -> cheapest -> B - cost_relative}.

[0220] Then, all the above sets are combined into the set of scoring nodes s of the sample statement. The size of the set of scoring nodes s (i.e., the number of sets included) is K. According to the foregoing introduction, the value of K is 6.

[0221] Step 508, the computer device calculates the similarity between the sets of scoring nodes among the sample statements to obtain multiple similarities between the sample statements.

[0222] After generating the set of scoring nodes s corresponding to each sample statement, for any two semantic trees T i and T j in the batch of sample data, the Jaccard similarity of each group of s k in s can be calculated. The specific formula is as follows:

[0223]

[0224] Among them, is the Jaccard similarity calculation formula, which means calculating the quotient of the intersection and union of two sets. s i,k is the k - th set in the set of scoring nodes s corresponding to the semantic tree T i , s j,k is the k - th set in the set of scoring nodes s corresponding to the semantic tree T j , and k is a positive integer from 1 to 6. In this way, 6 similarities between any two sample statements in the batch of sample data can be calculated.

[0225] Step 509, the computer device calculates a semantic contrast loss based on multiple similarities between sample statements and the sentence features of multiple sample statements.

[0226] Further, the semantic contrast loss can be calculated based on multiple similarities between sample statements and the sentence features of multiple sample statements. The specific calculation formula is as follows:

[0227]

[0228] Where B is a batch of samples, and B\{i} represents other sample data in the batch of samples except the i-th sample data. σ k (x) = Norm(W k x + b k ) represents a linear mapping with different parameters for different sets in s, where represents the Norm(·) normalization operation. t i and t j respectively represent the sentence features of the i-th and j-th sample statements.

[0229] Step 510, the computer device calculates an objective loss based on the intent detection loss, slot recognition loss, syntactic contrast loss, and semantic contrast loss.

[0230] After calculating the syntactic contrast loss and semantic contrast loss, the computer device can further calculate the objective loss based on the intent detection loss, slot recognition loss, syntactic contrast loss, and semantic contrast loss. The specific calculation formula is as follows:

[0231]

[0232] Where λ is a trade-off hyperparameter.

[0233] Step 511, the computer device updates the model parameters of the spoken language understanding model according to the objective loss.

[0234] Among them, after calculating the objective loss, the model parameters of the spoken language understanding model can be updated based on the objective loss. Then, a batch of sample data can be re-obtained, and the spoken language understanding model can be trained based on the re-obtained batch of sample data until the parameters of the spoken language understanding model converge to obtain the trained spoken language understanding model.

[0235] According to an embodiment of the present disclosure, a question answering method is provided. As Figure 11 shown, it is a flowchart of the question answering method provided by the present disclosure. The method includes the following steps;

[0236] Step 1110, obtain a target question.

[0237] Among them, the target question can be any question received when an object asks in any scenario of the spoken language understanding model trained by the foregoing training method of the spoken language understanding model. These scenarios can include, but are not limited to, the foregoing AI medical guidance scenario and book search scenario. For example, the target question can specifically be a question received in intelligent medical guidance and intelligent hospital service Q&A assistants, such as "Where is the pediatrics department?". The intelligent assistant can determine the target question by receiving the text "Where is the pediatrics department?" input by the user; it can also receive the voice signal input by the user, and then perform speech recognition on the voice signal to obtain the target question "Where is the pediatrics department?"; or the intelligent assistant can provide a video image acquisition function, can perform video image acquisition on the user, to acquire the sign language video of the user, and then perform recognition on the sign language video to obtain the target question. This method can greatly improve the problem input efficiency of special populations with language barriers.

[0238] Step 1120, use the spoken language understanding model to perform intent detection and slot recognition on the target question to obtain the target intent and the target slot.

[0239] After receiving the target question, the intelligent assistant can send the received target question to the spoken language understanding model, and use the spoken language understanding model to perform intent detection and slot recognition on the target question to obtain the target intent and the target slot. Among them, the spoken language understanding model here can specifically be a model trained by using the training method of the spoken language understanding model provided by the present disclosure.

[0240] For example, the spoken language understanding model can perform intent detection on the question "Where is the pediatrics department?" to determine that its intent is "asking for directions", and perform slot recognition on the target question to obtain the slot as "location".

[0241] In some embodiments, the target question can also include multiple intents and multiple slots to be filled. Thus, after inputting the target question into the spoken language understanding model, the spoken language understanding model can detect multiple intents and the slots to be filled corresponding to each intent. For example, when the target question is "Where is the pediatrics department? How to get there?", then performing intent detection on this target question can obtain two intents, "asking for directions" and "navigation", and at the same time, the slot information obtained by performing slot recognition on this target question can be "location" and "route".

[0242] Since the spoken language understanding model used in the embodiments of the present disclosure is a model trained by using the training method of the spoken language understanding model provided in the foregoing disclosure, and this model is a model that learns the grammar and semantic information of spoken language sentences and then combines the grammar and semantic information of spoken language sentences to perform spoken language understanding, the model can have better spoken language understanding ability, that is, the model has higher accuracy in detecting the intent of the input target sentence and higher accuracy in slot recognition. Therefore, using the spoken language understanding model to detect the intent and recognize the slots of the target question in the embodiments of the present disclosure can obtain more accurate intent and slots.

[0243] Step 1130, perform slot filling on the target slot based on the target intent, and generate answer data corresponding to the target question according to the slot filling result.

[0244] After detecting the accurate intent of the target question and identifying the accurate slot information corresponding to the target question by using the spoken language understanding model, slot filling can be performed on the target slot according to the pre-configured information corresponding to the target slot, and answer data corresponding to the target question can be generated according to the slot filling result.

[0245] Among them, when performing slot filling according to the slots identified by the spoken language understanding model, the corresponding knowledge graph in this usage scenario can be obtained first. The knowledge graph (Knowledge Graph, KG) is a modern theory that combines the theories and methods of disciplines such as applied mathematics, graphics, information visualization technology, and information science with methods such as bibliometric citation analysis and co-occurrence analysis, and uses a visual graph to vividly display the core structure, development history, frontier fields, and overall knowledge architecture of the discipline to achieve the purpose of multi-disciplinary integration. It displays complex knowledge fields through data mining, information processing, knowledge measurement, and graph drawing, reveals the dynamic development law of the knowledge field, and provides practical and valuable references for disciplinary research. The knowledge graph can be divided into a general knowledge graph and a special knowledge graph according to the application field. In some specific fields and scenarios, a special knowledge graph matching the field and scenario can be used to improve the accuracy of knowledge use in this field. For example, in the embodiments of the present disclosure, when the question answering method is applied to intelligent medical guidance and intelligent hospital service question answering assistants, a special knowledge graph corresponding to the hospital can be constructed according to the relevant knowledge of the current hospital. After detecting the intent result and slot result of the target question input by the user by using the spoken language understanding model provided in this case, a special knowledge graph can be further obtained to perform slot filling on the recognized slot result to obtain an accurate slot filling result.

[0246] After obtaining the accurate slot filling result, corresponding answer data for answering can be further generated based on the detected intent result, the recognized slot result, and the slot filling result obtained by slot filling.

[0247] Furthermore, the answer data can also be output. Among them, when outputting the answer data, the answer data can be output in the default output form, for example, by displaying the answer text on the display screen of the intelligent assistant. Or, in some embodiments, the intelligent assistant can also automatically adjust the output form of the answer data according to the input mode of the target question. For example, when the target question is input in text form, the answer data can be displayed in text form on the display interface of the intelligent assistant; when the target question is input in voice form, when outputting the answer data, the answer data can be output in the form of voice broadcast; or, when the target question is input in the form of a sign language video, then the answer data can be converted into a sign language video, and then the answer data can be output by displaying the sign language video on the display interface of the intelligent assistant. Or, in some embodiments, after obtaining the answer data, multi-modal output data can be generated according to the answer data, that is, including multi-form outputs such as text output, voice broadcast output, and sign language video output.

[0248] Device and equipment description of the embodiments of the present disclosure

[0249] It can be understood that although the steps in the above various flowcharts are sequentially displayed according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0250] It should be noted that in each specific embodiment of the present disclosure, when it comes to performing relevant processing based on data related to the target object characteristics such as target object attribute information or attribute information set, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant region. In addition, when the embodiment of the present application needs to obtain the target object attribute information, the separate permission or separate consent of the target object will be obtained by means of a pop-up window or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.

[0251] Figure 12Schematic diagram of the training device 1200 for the spoken language understanding model provided by the embodiments of the present disclosure. The device includes:

[0252] A first acquisition unit 1210, configured to acquire training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and slot labels corresponding to the intent label;

[0253] A first calculation unit 1220, configured to input the sample sentence into the spoken language understanding model for intent prediction and slot recognition processing, and calculate a joint training loss according to the intent prediction result, slot recognition result, and label data output by the spoken language understanding model;

[0254] A second calculation unit 1230, configured to construct a syntax tree of words in the sample sentence, and calculate a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, where the word features are the features corresponding to the words obtained by encoding the sample sentence by the spoken language understanding model;

[0255] A third calculation unit 1240, configured to construct a semantic tree corresponding to each sample sentence, and calculate a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree;

[0256] An update unit 1250, configured to calculate a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and update the parameters of the spoken language understanding model based on the target loss.

[0257] Optionally, in some embodiments, the second calculation unit includes:

[0258] A first construction subunit, configured to construct a main syntax tree of the sample sentence;

[0259] A generation subunit, configured to generate at least one positive syntax tree and at least one negative syntax tree of the words in the sample sentence according to the connection relationship between the nodes in the main syntax tree;

[0260] A first calculation subunit, configured to calculate a first contrast loss according to the syntax features determined by at least one positive syntax tree and at least one negative syntax tree, and the word features corresponding to the words.

[0261] Optionally, in some embodiments, the generation subunit includes:

[0262] An extraction module, configured to extract at least one subtree with any target word in the sample sentence as the root node according to the connection relationship between the nodes in the main syntax tree, to obtain at least one positive syntax tree of the target word;

[0263] A replacement module, configured to randomly replace the nodes in at least one positive syntax tree to obtain at least one negative syntax tree of the target word.

[0264] Optionally, in some embodiments, the first calculation subunit includes:

[0265] A first generation module, configured to generate at least one positive grammar feature based on at least one positive grammar tree;

[0266] A second generation module, configured to generate at least one negative grammar feature according to at least one negative grammar tree;

[0267] A determination module, configured to perform text encoding on a sample statement based on a statement encoder in a spoken language understanding model to obtain a sample statement feature, and determine a word feature corresponding to a word in the sample statement feature;

[0268] A first calculation module, configured to calculate a first contrast loss according to at least one positive grammar feature, at least one negative grammar feature, and the word feature.

[0269] Optionally, in some embodiments, the first calculation module includes:

[0270] A first calculation sub-module, configured to calculate a first cosine similarity between at least one positive grammar feature and the word feature, and calculate a second cosine similarity between at least one negative grammar feature and the word feature;

[0271] A second calculation sub-module, configured to calculate a first contrast loss according to the first cosine similarity and the second cosine similarity.

[0272] Optionally, in some embodiments, the third calculation unit includes:

[0273] A second construction sub-unit, configured to construct a semantic tree corresponding to each sample statement;

[0274] A second calculation sub-unit, configured to calculate a similarity between sample statements according to a similarity relationship between node combination relationships and node connection relationships in the semantic tree;

[0275] A third calculation sub-unit, configured to calculate a second contrast loss according to the similarity between sample statements.

[0276] Optionally, in some embodiments, the second construction sub-unit is further configured to:

[0277] Construct a semantic tree for each sample statement according to label data corresponding to the sample statement, where the semantic tree includes at least one semantic branch, and each branch includes an intent node, a slot node, and a value node;

[0278] The second calculation sub-unit is further configured to:

[0279] Generate a plurality of node sets corresponding to each sample statement according to combination and connection relationships between the intent node, the slot node, and the value node;

[0280] Calculate the similarity of the corresponding node sets between sample statements to obtain multiple similarities corresponding to multiple node sets between sample statements;

[0281] The third calculation subunit is further configured to:

[0282] Calculate a second contrast loss according to the multiple similarities.

[0283] Optionally, in some embodiments, the third calculation subunit includes:

[0284] An acquisition module for acquiring the sample statement features of each sample statement;

[0285] A second calculation module for calculating a second contrast loss according to the sample statement features and the multiple similarities.

[0286] Optionally, in some embodiments, the first calculation unit includes:

[0287] A processing subunit for inputting a sample statement into a spoken language understanding model for intent prediction and slot recognition processing to obtain an intent prediction result and a slot recognition result output by the spoken language understanding model;

[0288] A fourth calculation subunit for calculating an intent prediction loss according to the intent prediction result and the intent label corresponding to the sample statement, and calculating a slot recognition loss according to the slot recognition result and the slot label corresponding to the sample statement;

[0289] A fifth calculation subunit for calculating the sum of the intent prediction loss and the slot recognition loss to obtain a joint training loss.

[0290] Optionally, in some embodiments, the fourth calculation subunit includes:

[0291] A third calculation module for calculating at least one first cross-entropy between the intent prediction result and at least one intent label, and performing a summation calculation on the at least one first cross-entropy to obtain an intent prediction loss;

[0292] A fourth calculation module for calculating at least one second cross-entropy between the slot recognition result and at least one slot label, and performing a summation calculation on the at least one second cross-entropy to obtain a slot recognition loss.

[0293] Optionally, in some embodiments, the update unit includes:

[0294] A sixth calculation subunit for calculating the sum of the first contrast loss and the second contrast loss to obtain a contrast loss sum;

[0295] A second acquisition subunit, configured to acquire a contrast loss and a corresponding first weight, and a second weight corresponding to a joint training loss;

[0296] A seventh calculation subunit, configured to perform weighted calculation on the contrast loss and the joint training loss based on the first weight and the second weight to obtain a target loss;

[0297] An update subunit, configured to update parameters of the speech understanding model based on the target loss.

[0298] Figure 13 FIG. 1300 is a schematic structural diagram of a question answering device provided in an embodiment of the present disclosure. The device includes:

[0299] A second acquisition unit 1310, configured to acquire a target question;

[0300] A processing unit 1320, configured to perform intent detection and slot recognition on the target question by using a speech understanding model trained based on the training method of the speech understanding model provided in the present disclosure to obtain a target intent and a target slot;

[0301] A generation unit 1330, configured to perform slot filling on the target slot based on the target intent, and generate answer data corresponding to the target question according to the slot filling result.

[0302] In an embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit including the function of the module or unit.

[0303] Referring to Figure 14 , Figure 14 , FIG. 140 is a partial structural block diagram of a terminal for implementing the training method of the speech understanding model in an embodiment of the present disclosure. The terminal 140 includes components such as a radio frequency (RF) circuit 1410, a memory 1415, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490. Those skilled in the art can understand that Figure 14 the shown structure of the terminal 140 does not constitute a limitation on a mobile phone or a computer, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0304] The RF circuit 1410 can be used for receiving and transmitting information during information reception and call processes. Specifically, after receiving the downlink information from the base station, it is processed by the processor 1480. Additionally, the data designed for uplink is sent to the base station.

[0305] The memory 1415 can be used to store software programs and modules. The processor 1480 executes various functional applications and document editing of the terminal by running the software programs and modules stored in the memory 1415.

[0306] The input unit 1430 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1430 can include a touch panel 1431 and other input devices 1432.

[0307] The display unit 1440 can be used to display the input information or provided information, as well as various menus of the terminal. The display unit 1440 can include a display panel 1441.

[0308] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface.

[0309] In this embodiment, the processor 1480 included in the terminal 140 can execute the training method of the spoken language understanding model in the previous embodiment.

[0310] The terminal 140 in the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc.

[0311] Figure 15 It is a structural block diagram of a part of the server 110 for implementing the training method of the spoken language understanding model in the embodiments of the present disclosure. The server 110 can vary significantly due to configuration or performance differences, and can include one or more central processing units (CPUs) 1522 (for example, one or more processors) and a storage device 1532, and one or more storage media 1530 (for example, one or more mass storage devices) that store application programs 1542 or data 1544. Among them, the storage device 1532 and the storage media 1530 can be transient storage or persistent storage. The programs stored in the storage media 1530 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations on the server 110. Further, the central processor 1522 can be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the server 110.

[0312] The server 110 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0313] The central processing unit 1522 in the server 110 may be used to execute the training method of the spoken language understanding model according to the embodiments of the present disclosure.

[0314] The embodiments of the present disclosure also provide a storage medium for storing program codes for executing the training method of the spoken language understanding model in the foregoing various embodiments.

[0315] The embodiments of the present disclosure also provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes the training method of the above-mentioned spoken language understanding model.

[0316] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprise" and "include" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0317] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c may be single or multiple.

[0318] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality of (or multiple)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.

[0319] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0320] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0321] In addition, the functional units in each embodiment of the present disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0322] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present disclosure. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs that can store program codes.

[0323] It should also be understood that the various implementation manners provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.

[0324] The above is a specific description of the embodiments of the present disclosure. However, the present disclosure is not limited to the above embodiments. Those skilled in the art can still make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A training method for a spoken language understanding model, characterized in that The method includes: Obtaining training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and slot labels corresponding to the intent label; Inputting the sample sentences into a spoken language understanding model for intent prediction and slot recognition processing, and calculating a joint training loss according to the intent prediction result, slot recognition result output by the spoken language understanding model, and the label data; Constructing a syntax tree of words in the sample sentence, and calculating a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words, where the word features are the features corresponding to the words obtained by encoding the sample sentence by the spoken language understanding model; Constructing a semantic tree corresponding to each sample sentence, and calculating a second contrast loss according to the similarity between the sample sentences calculated based on the semantic tree; Calculating a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and updating the parameters of the spoken language understanding model based on the target loss.

2. The method according to claim 1, wherein The constructing a syntax tree of words in the sample sentence, and calculating a first contrast loss according to the syntax features corresponding to the syntax tree and the word features corresponding to the words includes: Constructing a main syntax tree of the sample sentence; Generating at least one positive syntax tree and at least one negative syntax tree of the words in the sample sentence according to the connection relationship between nodes in the main syntax tree; Calculating a first contrast loss according to the syntax features determined based on the at least one positive syntax tree and the at least one negative syntax tree, and the word features corresponding to the words.

3. The method according to claim 2, wherein The generating at least one positive syntax tree and at least one negative syntax tree of the words in the sample sentence according to the connection relationship between nodes in the main syntax tree includes: Extracting at least one subtree with any target word in the sample sentence as the root node according to the connection relationship between nodes in the main syntax tree, to obtain at least one positive syntax tree of the target word; Randomly replacing the nodes in the at least one positive syntax tree to obtain at least one negative syntax tree of the target word.

4. The method according to claim 2, wherein The calculating a first contrast loss according to the syntax features determined based on the at least one positive syntax tree and the at least one negative syntax tree, and the word features corresponding to the words includes: Generating at least one positive syntax feature based on the at least one positive syntax tree; Generating at least one negative syntax feature according to the at least one negative syntax tree; Encoding the sample sentence by the sentence encoder in the spoken language understanding model to obtain a sample sentence feature, and determining the word features corresponding to the words in the sample sentence feature; Calculating a first contrast loss according to the at least one positive syntax feature, the at least one negative syntax feature, and the word features.

5. The method according to claim 4, characterized in that, The calculating a first contrast loss according to the at least one positive syntax feature, the at least one negative syntax feature, and the word features includes: Calculating a first cosine similarity between the at least one positive syntax feature and the word features, and calculating a second cosine similarity between the at least one negative syntax feature and the word features; Calculate a first contrast loss based on the first cosine similarity and the second cosine similarity.

6. The method according to claim 1, wherein The step of inputting the sample statement into the spoken language understanding model for intent prediction and slot recognition processing, and calculating a joint training loss according to the intent prediction result, slot recognition result, and label data output by the spoken language understanding model includes: Input the sample statement into the spoken language understanding model for intent prediction and slot recognition processing to obtain the intent prediction result and slot recognition result output by the spoken language understanding model; Calculate an intent prediction loss according to the intent prediction result and the intent label corresponding to the sample statement, and calculate a slot recognition loss according to the slot recognition result and the slot label corresponding to the sample statement; Calculate the sum of the intent prediction loss and the slot recognition loss to obtain the joint training loss.

7. The method according to claim 6, wherein The step of calculating an intent prediction loss according to the intent prediction result and the intent label corresponding to the sample statement, and calculating a slot recognition loss according to the slot recognition result and the slot label corresponding to the sample statement includes: Calculate at least one first cross-entropy between the intent prediction result and the at least one intent label, and sum the at least one first cross-entropy to obtain the intent prediction loss; Calculate at least one second cross-entropy between the slot recognition result and the at least one slot label, and sum the at least one second cross-entropy to obtain the slot recognition loss.

8. The method according to claim 1, wherein The step of constructing a semantic tree corresponding to each sample statement and calculating a second contrast loss according to the similarity between the sample statements calculated based on the semantic tree includes: Construct a semantic tree corresponding to each sample statement; Calculate the similarity between the sample statements according to the similarity relationship between the node combination relationship and the node connection relationship in the semantic tree; Calculate a second contrast loss according to the similarity between the sample statements.

9. The method according to claim 8, characterized in that, The step of constructing a semantic tree corresponding to each sample statement includes: Construct a semantic tree of each sample statement according to the label data corresponding to the sample statement, where the semantic tree includes at least one semantic branch, and each branch includes an intent node, a slot node, and a value node; The step of calculating the similarity between the sample statements according to the similarity relationship between the node combination relationship and the node connection relationship in the semantic tree includes: Generate a plurality of node sets corresponding to each sample statement according to the combination and connection relationship between the intent node, the slot node, and the value node; Calculate the similarity between the corresponding node sets of the sample statements to obtain a plurality of similarities corresponding to the plurality of node sets between the sample statements; The step of calculating a second contrast loss according to the similarity between the sample statements includes: Calculate a second contrast loss according to the plurality of similarities.

10. The method according to claim 9, characterized in that, The step of calculating a second contrast loss according to the plurality of similarities includes: Obtain the sample statement feature of each sample statement; Calculate a second contrast loss according to the sample statement feature and the plurality of similarities.

11. The method according to claim 1, characterized in that, Calculating a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and updating parameters of the spoken language understanding model based on the target loss, includes: Calculating a sum of the first contrast loss and the second contrast loss to obtain a contrast loss sum; Obtaining a first weight corresponding to the contrast loss sum and a second weight corresponding to the joint training loss; Performing weighted calculation on the contrast loss sum and the joint training loss based on the first weight and the second weight to obtain a target loss; Updating parameters of the spoken language understanding model based on the target loss.

12. A method for answering questions, characterized in that, The method includes: Obtaining a target question; Performing intent detection and slot recognition on the target question by using a spoken language understanding model trained by the training method of the spoken language understanding model according to any one of claims 1 to 11 to obtain a target intent and target slots; Performing slot filling on the target slots based on the target intent, and generating answer data corresponding to the target question according to a slot filling result.

13. A training device for a spoken language understanding model, characterized in that, The apparatus includes: A first obtaining unit, configured to obtain training sample data including a plurality of sample sentences and label data corresponding to the sample sentences, where the label data includes at least one intent label and a slot label corresponding to the intent label; A first calculating unit, configured to input the sample sentences into a spoken language understanding model to perform intent prediction and slot recognition processing, and calculate a joint training loss according to an intent prediction result, a slot recognition result, and the label data output by the spoken language understanding model; A second calculating unit, configured to construct a syntax tree of words in the sample sentences, and calculate a first contrast loss according to a syntax feature corresponding to the syntax tree and a word feature corresponding to the words, where the word feature is a feature corresponding to the words obtained by encoding the sample sentences by the spoken language understanding model; A third calculating unit, configured to construct a semantic tree corresponding to each sample sentence, and calculate a second contrast loss according to a similarity between the sample sentences calculated based on the semantic tree; An updating unit, configured to calculate a target loss according to the first contrast loss, the second contrast loss, and the joint training loss, and update parameters of the spoken language understanding model based on the target loss.

14. A question answering device, characterized in that, The apparatus includes: A second obtaining unit, configured to obtain a target question; A processing unit, configured to perform intent detection and slot recognition on the target question by using a spoken language understanding model trained by the training method of the spoken language understanding model according to any one of claims 1 to 11 to obtain a target intent and target slots; A generating unit, configured to perform slot filling on the target slots based on the target intent, and generate answer data corresponding to the target question according to a slot filling result.

15. A storage medium storing a computer program, characterized in that, When being executed by a processor, the computer program implements the training method of the spoken language understanding model according to any one of claims 1 to 11 or the question answering method according to claim 12.

16. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the training method of the spoken language understanding model according to any one of claims 1 to 11 or the question answering method according to claim 12.

17. A computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the training method of the spoken language understanding model according to any one of claims 1 to 11 or the question answering method according to claim 12.