Text understanding model training method and device, equipment and storage medium

By performing semantic relationship encoding and decoding on the preset text understanding model, the problem of inaccurate intent and slot prediction in the existing technology is solved, and more efficient spoken language understanding is achieved.

CN120633675APending Publication Date: 2025-09-12腾讯医疗健康(深圳)有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410275054.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies are unable to accurately predict the intent and slots of utterances, resulting in poor spoken language understanding performance.

Method used

By obtaining sample text and sample label sets, using the preset text understanding model to perform semantic relationship encoding processing, combined with intent decoding and slot decoding, iterative training is carried out to improve prediction accuracy.

Benefits of technology

Capturing the semantic relationship information between intent labels and utterances in the embedding space improves the accuracy of predicting intent and slots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633675A_ABST
    Figure CN120633675A_ABST
Patent Text Reader

Abstract

The invention discloses a text understanding model training method and device, equipment and a storage medium. Comprising the steps of obtaining a sample text and a sample label set; inputting the plurality of sample intention tags and the plurality of sample slot tags into a preset text understanding model for semantic relationship coding processing to obtain an intention tag relationship representation and a slot tag relationship representation; performing intention decoding processing on the sample text representation and the intention label relation representation based on a preset text understanding model to obtain a predicted intention result, and performing slot decoding processing on the sample text representation and the slot label relation representation to obtain a predicted slot result; determining target loss according to the prediction intention result, the target intention label, the prediction slot position result and the target slot position label; and performing iterative training on a preset text understanding model based on the target loss to obtain a trained text understanding model. The method can improve the prediction accuracy of intention prediction and slot prediction, and can be used for various scenes such as artificial intelligence, spoken language understanding, task-based dialogues and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and more specifically, to a training method for a text understanding model and related methods, devices, equipment, and storage media. Background Art

[0002] Spoken Language Understanding (SLU), a core component of a task-oriented dialogue system (TADS), aims to capture the semantic representation of query statements. This information is then used by the dialogue state labeling module and the natural language generation module. It can be applied to task-based scenarios such as intelligent customer service and food ordering.

[0003] The task of spoken language understanding is to interpret object utterances based on semantic meaning representations so that they can be used with backend actions or knowledge providers. Simply put, this involves converting natural language into domain, intent, and slot-value pairs that computers can process. To this end, spoken language understanding primarily consists of two subtasks: intent detection and slot filling, which can share information.

[0004] Intent detection is typically a sentence-level text classification task, while slot filling is typically a word-level sequence labeling task, aiming to convert the subject utterance into a label sequence of equivalent length. Currently, various deep neural network-based approaches have been proposed to address intent detection and slot filling. However, these approaches cannot accurately predict the intent of an utterance and the slot labels, resulting in poor performance in spoken language understanding. Summary of the Invention

[0005] The present invention provides a text comprehension model training method and related methods, devices, equipment, and storage media to address the problem in related technologies of poor spoken language comprehension performance caused by the inability to accurately predict the intent and slot of an utterance.

[0006] According to one aspect of the present application, a training method for a text understanding model is provided, the method comprising: obtaining sample text and a sample label set, the sample label set comprising a plurality of sample intent labels and a plurality of sample slot labels; inputting the plurality of sample intent labels and the plurality of sample slot labels into a preset text understanding model for respective semantic relationship encoding processing to obtain corresponding intent label relationship representations and slot label relationship representations; performing intent decoding processing on the sample text representation and the intent label relationship representation of the sample text based on the preset text understanding model to obtain a predicted intent result, and performing slot decoding processing on the sample text representation and the slot label relationship representation based on the preset text understanding model to obtain a predicted slot result; determining a target loss based on the predicted intent result and a target intent label corresponding to the sample text, as well as the predicted slot result and the target slot label corresponding to the sample text; iteratively training the preset text understanding model based on the target loss until a training end condition is reached to obtain a trained text understanding model.

[0007] According to one aspect of the present application, a method for spoken language understanding is provided, which includes: obtaining a target text and a tag set, the tag set including multiple intent tags and multiple slot tags; inputting the target text, the multiple intent tags, and the multiple slot tags into a trained text understanding model for spoken language understanding, and obtaining text intent and text slots corresponding to the target text; wherein the text understanding model is obtained according to the training method of the above-mentioned trained text understanding model.

[0008] According to one aspect of the present application, a device for training a text comprehension model is provided, the device comprising:

[0009] A sample acquisition module, configured to acquire sample text and a sample label set, wherein the sample label set includes a plurality of sample intent labels and a plurality of sample slot labels;

[0010] An encoding processing module, configured to input the plurality of sample intent labels and the plurality of sample slot labels into a preset text understanding model for semantic relationship encoding processing, respectively, to obtain corresponding intent label relationship representations and slot label relationship representations;

[0011] a decoding processing module, configured to perform intent decoding processing on the sample text representation and the intent label relationship representation of the sample text based on the preset text understanding model to obtain a predicted intent result, and perform slot decoding processing on the sample text representation and the slot label relationship representation based on the preset text understanding model to obtain a predicted slot result;

[0012] a loss determination module, configured to determine a target loss based on the predicted intent result and the target intent label corresponding to the sample text, and the predicted slot result and the target slot label corresponding to the sample text;

[0013] The model training module is used to iteratively train the preset text understanding model based on the target loss until the training end condition is met to obtain a trained text understanding model.

[0014] Optionally, the preset text understanding model includes an intent decoder and a slot decoder; the decoding processing module may include:

[0015] A text representation unit, configured to obtain a sample text representation corresponding to the sample text;

[0016] an intent prediction unit, configured to input the sample text representation and the intent label relationship representation into the intent decoder for intent decoding processing to obtain a predicted intent result; wherein the intent label relationship representation indicates a first semantic relationship between a plurality of sample intent labels, so that the intent decoder performs intent decoding processing on the sample text based on the first semantic relationship and a second semantic relationship between the sample text and the sample intent label;

[0017] The slot prediction unit is used to input the sample text representation and the slot label relationship representation into the slot decoder for slot decoding processing to obtain a predicted slot result; wherein, the slot label relationship representation indicates a third semantic relationship between multiple sample slot labels, so that the slot decoder performs slot decoding processing on the sample text according to the third semantic relationship and the fourth semantic relationship between the sample text and the sample slot label.

[0018] Optionally, the intention prediction unit can be specifically used to: perform representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation; input each sample word segmentation enhanced representation and the intention label relationship representation into the intention decoder for intention decoding processing to obtain a sample intention score vector corresponding to each sample word segmentation in the sample text; perform voting processing on each of the sample intention score vectors to obtain a predicted intention result corresponding to the sample text.

[0019] Optionally, the slot prediction unit can be specifically used to: perform representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation; input each sample word segmentation enhanced representation and the slot label relationship representation into the slot decoder for slot decoding processing to obtain a sample slot score vector; perform slot global prediction on the sample slot score vector corresponding to each sample word segmentation enhanced representation based on the first activation function to obtain a slot global prediction result corresponding to each sample word segmentation enhanced representation; perform slot decoding processing on each of the slot global prediction results according to the second activation function to obtain a predicted slot result.

[0020] Optionally, the preset text understanding model includes a label encoder and a prompt learner, and the encoding processing module may include:

[0021] an intent encoding unit, configured to input the plurality of sample intent labels into the label encoder for encoding the intent labels to obtain sample intent label representations;

[0022] an intention learning unit, configured to input the sample intention label representation into the prompt learner for intention prompt learning to obtain an intention label relationship representation;

[0023] a label encoding unit, configured to input the plurality of sample slot labels into the label encoder for slot label encoding to obtain a sample slot label representation;

[0024] The label learning unit is used to input the sample slot label representation into the prompt learner to perform slot prompt learning to obtain a slot label relationship representation.

[0025] Optionally, the intention learning unit can be specifically used to: input the sample intention label representation into the prompt learner for intention mapping processing to generate an intention mapping probability distribution; perform matrix multiplication calculation based on the prompt vector and the intention mapping probability distribution to generate an intention label relationship representation, and the prompt vector is obtained by random initialization.

[0026] Optionally, the label learning unit can be specifically used to: input the sample slot label representation into the prompt learner for slot mapping processing to generate a slot mapping probability distribution; perform matrix multiplication calculation based on the prompt vector and the slot mapping probability distribution to generate a slot label relationship representation, and the prompt vector is obtained by random initialization.

[0027] Optionally, the preset text understanding model includes a text encoder; the text representation unit can be specifically used to: perform word segmentation processing on the sample text to obtain a sample word segmentation sequence corresponding to the sample text; input the sample word segmentation sequence into the text encoder for text encoding to obtain a sample text representation.

[0028] Optionally, the model training module can be specifically used to iteratively train the text encoder, prompt learner, intent decoder and slot decoder of the preset text understanding model based on the target loss until the training end condition is reached, thereby obtaining a trained text understanding model.

[0029] Optionally, the loss determination module may include:

[0030] A first determining unit, configured to determine a first loss based on the predicted intent result and a target intent label corresponding to the sample text;

[0031] A second determining unit, configured to determine a second loss based on the predicted slot result and a target slot label corresponding to the sample text;

[0032] A weighted determination unit is used to perform weighted calculation on the first loss and the second loss to obtain a target loss.

[0033] According to one aspect of the present application, a spoken language understanding device is provided, the device comprising:

[0034] A data acquisition module, configured to acquire a target text and a tag set, wherein the tag set includes a plurality of intent tags and a plurality of slot tags;

[0035] The spoken language comprehension module is used to input the target text, the multiple intent labels and the multiple slot labels into the trained text comprehension model for spoken language comprehension, and obtain the text intent and text slots corresponding to the target text; wherein the trained text comprehension model is trained according to the training method of the above-mentioned text comprehension model.

[0036] Optionally, the trained text understanding model includes a text encoder, a label encoder, a prompt learner, an intent decoder, and a slot decoder; and the spoken language understanding module may include:

[0037] A word segmentation label unit, used to input the word segmentation sequence of the target text into the text encoder to obtain a text representation;

[0038] a first encoding unit, configured to input the intent label into the label encoder for intent encoding to obtain a first representation of the intent;

[0039] a second encoding unit, configured to input the slot label into the label encoder for slot encoding to obtain a first slot representation;

[0040] A first learning unit is configured to input the first intention representation into the prompt learner to perform intention prompt learning to obtain a second intention representation;

[0041] A second learning unit is configured to input the first slot representation into the prompt learner to perform slot prompt learning to obtain a second slot representation;

[0042] A first decoding unit is configured to input the text representation and the second representation of the intention into the intention decoder for intention decoding to obtain a text intention corresponding to the target text;

[0043] The second decoding unit is used to input the text representation and the second slot representation into the slot decoder for slot decoding processing to obtain the text slot corresponding to the target text.

[0044] According to one aspect of the present application, a computer-readable storage medium is provided, which stores a computer program, wherein when the computer program is executed by a processor, the above-mentioned text understanding model training method or spoken language understanding method is executed.

[0045] According to one aspect of the present application, a computer device is provided, which includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is called by the processor, the computer program executes the above-mentioned text understanding model training method or spoken language understanding method.

[0046] According to one aspect of the present application, a computer program product is provided, which includes a computer program stored in a storage medium; a processor of a computer device reads the computer program from the storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned text understanding model training method or spoken language understanding method.

[0047] The embodiment of the present application can obtain sample text and a sample label set, wherein the sample label set includes a plurality of sample intent labels and a plurality of sample slot labels. The plurality of sample intent labels and the plurality of sample slot labels are input into a preset text understanding model for semantic relationship encoding processing respectively, and corresponding intent label relationship representations and slot label relationship representations are obtained. In this way, a plurality of sample intent labels can be encoded into an embedding space, so that the semantic relationship information between the plurality of sample intent labels can be captured in the embedding space as supervision information for supervised learning of intent prediction. Similarly, the semantic relationship information between the plurality of sample slot labels can also be captured in the embedding space as supervision information for supervised learning of slot prediction.

[0048] Furthermore, based on the preset text understanding model, the sample text representation and the intent label relationship representation of the sample text are subjected to intent decoding processing to obtain a predicted intent result, and based on the preset text understanding model, the sample text representation and the slot label relationship representation are subjected to slot decoding processing to obtain a predicted slot result. Then, based on the predicted intent result and the target intent label corresponding to the sample text, as well as the predicted slot result and the target slot label corresponding to the sample text, a target loss is determined, and the preset text understanding model is iteratively trained based on the target loss until the training end condition is met, thereby obtaining a trained text understanding model. Thus, the semantic relationship information between intent labels and utterances, as well as the semantic relationship information between slot labels and utterances, can be effectively captured in the embedding space. Thus, by using the semantic relationship information between multiple sample intent labels, the semantic relationship information between multiple sample slot labels, the semantic relationship information between intent labels and utterances, and the semantic relationship information between slot labels and utterances as supervision information, rich supervision information is provided for the supervised learning of the preset text understanding model, thereby improving the accuracy of the preset text understanding model in predicting the intent and slots of utterances.

[0049] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0051] Figure 1 A schematic diagram of a system architecture provided in an embodiment of the present application is shown.

[0052] Figure 2 A schematic diagram of system deployment provided in an embodiment of the present application is shown.

[0053] Figure 3 An application scenario diagram of a text understanding model training method provided in an embodiment of the present application is shown.

[0054] Figure 4 An application scenario diagram of another text comprehension model training method provided in an embodiment of the present application is shown.

[0055] Figure 5A flow chart of a method for training a text comprehension model provided in an embodiment of the present application is shown.

[0056] Figure 6 A schematic diagram of a discourse, its intention and slot provided in an embodiment of the present application is shown.

[0057] Figure 7 A schematic diagram of the architecture of a preset text understanding model provided in an embodiment of the present application is shown.

[0058] Figure 8 A schematic diagram of the architecture of a prompt learner provided in an embodiment of the present application is shown.

[0059] Figure 9 A schematic diagram of a data set provided in an embodiment of the present application is shown.

[0060] Figure 10 A schematic diagram showing an experimental result provided in an embodiment of the present application is shown.

[0061] Figure 11 A schematic diagram showing another experimental result provided in an embodiment of the present application is shown.

[0062] Figure 12 A flow chart of a spoken language comprehension method provided in an embodiment of the present application is shown.

[0063] Figure 13 A flow chart of a spoken language comprehension method provided in an embodiment of the present application is shown.

[0064] Figure 14 This is a module block diagram of a text comprehension model training device provided in an embodiment of the present application.

[0065] Figure 15 This is a module block diagram of a spoken language understanding device provided in an embodiment of the present application.

[0066] Figure 16 This is a module block diagram of a computer device provided in an embodiment of the present application.

[0067] Figure 17 It is a module block diagram of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0068] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The implementation methods described with reference to the drawings are exemplary and are only used to explain the present application, and should not be understood as limitations on the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0069] Some processes described in the specification, claims, and accompanying figures include multiple steps that appear in a specific order. However, it should be understood that these steps may be performed in a different order or in parallel. Step numbers are used solely to distinguish between different steps and do not inherently indicate an order of execution. Furthermore, terms such as "first" and "second" are used to distinguish similar items and are not necessarily used to describe a specific order, sequence, or quantity.

[0070] It is worth noting that in the specific implementation of the present application, data related to sample texts, for example, sample texts and target texts, etc., involve data related to the privacy of the object. When the above embodiments of the present application are applied to specific products or technologies, it is necessary to obtain permission or consent from the object, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards. For example, when the embodiment of the present application needs to obtain relevant data such as the target text, a separate permission or separate consent for the relevant data such as the sample text can be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining a separate permission or separate consent, the necessary relevant data for the normal operation of the embodiment of the present application can be obtained.

[0071] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:

[0072] The training method for the text understanding model proposed in this application involves artificial intelligence (AI) technology. Artificial intelligence technology is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, giving them the functions of perception, reasoning, and decision-making.

[0073] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0074] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; it also involves computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, is developed from the large language model (Large Language Model) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies. For example, in the embodiment of the present application, it is used for spoken language understanding of the task-based question answering system.

[0075] Among them, intent recognition is a subtask of spoken language understanding, which is usually regarded as a classification problem, that is, identifying the intent of the input sentence by classifying it into predefined intent categories. These categories can be various tasks, queries and requests, such as search, purchase, consultation and command. For example, the object input is: "I want to book a plane ticket from Beijing to Shanghai", and the corresponding intent recognition is: booking a plane ticket. Slot filling is another subtask of spoken language understanding. It can extract the values ​​of well-defined attributes of a given entity from a large-scale corpus. Simply put, it is to mark certain words in the sentence. For example, the object input is: "Help me book a plane ticket from Hangzhou", and correspondingly, [air ticket] is filled in the slot named [transportation], and [Hangzhou] is filled in the slot named [destination].

[0076] Spoken language understanding, a core component of task-oriented dialogue systems, can be used to capture the semantics of object queries. It involves intent detection and slot filling. Traditional approaches treat these two tasks as separate. For example, intent detection typically utilizes convolutional neural networks and recurrent neural networks for sentence classification. Slot filling typically utilizes conditional random field algorithms and long short-term memory networks for sequence labeling. Consequently, traditional approaches ignore the shared knowledge between the two tasks.

[0077] Intuitively, intent detection and slot filling are not independent tasks, but rather highly interconnected. To this end, joint models are currently being employed to leverage the shared knowledge between the two tasks, performing both intent detection and slot filling simultaneously. Related techniques utilize Graph Attention Networks (GANs) to model the interaction between intent and slots, leveraging the correlation between the two subtasks. Furthermore, graph structures are further enhanced to enable bidirectional or heterogeneous guidance between intent and slots.

[0078] For example, a local slot-aware layer and a global intent-slot interaction layer are proposed to achieve the simultaneous generation of intent and slot sequences without autoregression. However, when jointly performing intent detection and slot filling, the related technologies ignore the rich semantics in the labels. Specifically, the labels in the two subtasks are now converted into orthogonal one-hot vectors in the embedding space, so that the intrinsic semantic relationship between the labels is not utilized. In addition, the intent or slot itself may be semantically related to the user's utterance, but the semantic relationship between them is still not fully utilized. Therefore, this type of method of the related technology cannot accurately predict the intent and slot of the utterance, resulting in poor performance in spoken language understanding.

[0079] In order to solve the above problems, the inventors have proposed a text comprehension model training method and a spoken language comprehension method provided in the embodiments of the present application after research. The following first describes the system architecture and related application scenarios of the above methods involved in this application.

[0080] See also Figure 1 , Figure 1 A schematic diagram of the system architecture is shown in FIG. Figure 1As shown, the above method provided in the embodiment of the present application can be applied in the system 100. The data acquisition device 110 is used to acquire training data. The training data can be a training sample used for model training. The training sample can be a text in the form of a sentence and a ground truth label (Ground Truth) corresponding to the text. The ground truth label includes an intent label and a slot label, for example, the intent label "search" and the slot label "B-destination". These labels can be manually labeled or labeled by a label labeling algorithm. After acquiring the training data, the data acquisition device 110 can store the training data in the database 120. The training device 130 can perform network training on the preset neural network based on the training data in the database 120 until the preset neural network meets the training end condition, thereby obtaining the trained target model 101.

[0081] Optionally, the training data maintained in the database 120 does not necessarily all come from the data acquisition device 110, but may also be received from other devices. For example, the execution device 140 may also serve as a data acquisition terminal, directly using the acquired data as new training data and storing it in the database 120. In addition, the training device 130 does not necessarily train the preset neural network entirely based on the training data maintained in the database 120, but may also train the preset neural network based on training data obtained from the cloud or other devices. For example, the execution device 140 may transmit the speech of the object collected in real time to the training device 130 as training data. The above description should not be construed as limiting the embodiments of the present application.

[0082] Among them, the training end condition can be: the total loss value of the target loss is less than a preset value, the total loss value of the target loss is within a threshold range, or the number of training times reaches a preset number, etc. The target model 101 can be a deep neural network or a network model composed of multiple neural networks. For example, the encoder (Encoder), decoder (Decoder), pre-trained language model (PLM) of the self-attention (Self-Attention) can be applied to the network framework for performing spoken language comprehension tasks, which is not limited here.

[0083] During the process of executing calculations and other related processing by the processing module 141 of the execution device 140, the execution device 140 may call data, programs, etc. in the data storage system 150 for the corresponding calculations and processing, and store data and instructions such as processing results obtained from the calculations in the data storage system 150. For example, the execution device 140 may store the intent and slot prediction results generated by the text understanding model in the data storage system 150.

[0084] The training device 130 and execution device 140 can be computer devices such as servers or terminals. A server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), blockchain, and big data and artificial intelligence platforms. A terminal can be a smartphone, tablet computer, desktop computer, or in-vehicle terminal.

[0085] It should be noted that Figure 1 This is only a schematic diagram of a system architecture provided by the embodiment of the present application. The system architecture and application scenarios described in the embodiment of the present application are intended to more clearly illustrate the technical solutions of the embodiment of the present application and do not constitute a limitation on the technical solutions provided by the embodiment of the present application. For example, Figure 1 The data storage system 150 in the embodiment is an internal memory relative to the execution device 140. In other cases, the data storage system 150 may also be placed outside the execution device 140.

[0086] See also Figure 2 , Figure 2 A schematic diagram of system deployment is shown. For example, the text understanding model training method and spoken language understanding method provided in the embodiment of the present application can also be deployed in Figure 2 In the system shown in FIG. , the system includes a terminal 240 , the Internet 230 , a gateway 220 , a server 210 , and the like.

[0087] Terminal 240 can include various forms, such as desktop computers, laptops, personal digital assistants (PDAs), mobile phones, in-vehicle terminals, home theater terminals, dedicated terminals, and intelligent customer service terminals. Furthermore, it can be a single device or a collection of multiple devices. Terminal 240 can communicate with Internet 230 via wired or wireless means to exchange data. For example, terminal 240 can communicate with Internet 230 via wireless router 250.

[0088] Server 210 is a computer system that provides certain services to terminal 240. Compared to ordinary terminal 240, server 210 has higher requirements in terms of stability, security, and performance. Server 210 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a high-performance computer (e.g., a virtual machine), or a combination of portions of multiple high-performance computers (e.g., virtual machines).

[0089] Gateway 220, also known as a gateway or protocol converter, implements network interconnection at the transport layer and is a computer system or device that acts as a converter. It acts as a translator between two systems using different communication protocols, data formats, languages, or even completely different architectures. It can also provide filtering and security functions. Messages sent from terminal 240 to server 210 can be transmitted to the corresponding server 210 via gateway 220. Messages sent from server 210 to terminal 240 can also be transmitted to the corresponding terminal 240 via gateway 220.

[0090] The text understanding model training method and spoken language understanding method of the embodiment of the present application can be implemented entirely on the terminal 240 , entirely on the server 210 , or partially on the terminal 240 and the other partially on the server 210 .

[0091] As an implementation method, the text understanding model training method and the spoken language understanding method can be completely implemented in the terminal 240. For example, the terminal 240 can be used as a training device to train the text understanding model locally. Then, the terminal 240 can also be used as an execution device to locally use the text understanding model to perform spoken language understanding on the text to be understood, that is, to perform intent recognition and slot filling on the text to be understood. Further, the intent recognition results and slot filling results are applied to downstream tasks, such as task-based dialogue systems such as intelligent customer service. This implementation method can achieve local intelligence without the help of the server 210, and only the terminal 240 can achieve local intelligence.

[0092] As another embodiment, the training method of the text understanding model and the spoken language understanding method can be completely implemented on the server 210. For example, the server 210 can be used as a training device to train a text understanding model. Furthermore, the server 210 can also be used as an execution device to perform subsequent tasks using the text understanding model. For example, the terminal 240 has a need to perform spoken language understanding tasks. The intelligent dialogue system deployed on the terminal 240 needs to perform spoken language understanding on the user's input speech. At this time, the terminal 240 can send the input speech to the server 210, and then the server 210 uses the text understanding model to perform spoken language understanding on the input speech, and returns the understanding result obtained by the spoken language understanding to the terminal 240, so that the intelligent dialogue system of the terminal 240 performs dialogue-related tasks based on the received understanding result.

[0093] As another embodiment, the text comprehension model training method and the spoken language comprehension method can be implemented partially on terminal 240 and partially on server 210. For example, server 210 can serve as a training device for the text comprehension model, and terminal 240 can serve as an execution device for using the text comprehension model. For example, server 210 can train a text comprehension model based on training samples. Terminal 240 can then deploy the trained text comprehension model locally and directly use the text comprehension model to perform spoken language comprehension on the text to be understood.

[0094] It should be noted that Figure 2 This is only a schematic diagram of a system deployment provided by the embodiment of the present application. The system deployment scheme described in the embodiment of the present application is only for the purpose of more clearly illustrating the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. For example, terminal 240 can generally refer to one of multiple terminals. This embodiment only refers to Figure 2 Those skilled in the art will appreciate that, as system deployment solutions evolve, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0095] The embodiments of the present application can be applied in various scenarios, such as Figure 3 The scenario of intelligent customer service shown in the figure, Figure 4 The AI ​​medical guidance scenario shown, etc.

[0096] (1) Application scenarios of intelligent customer service

[0097] Intelligent customer service leverages artificial intelligence (AI) technology to provide automated customer service through a conversational interface. It interacts with users via voice or text, answering their questions, providing relevant services, and transferring them to a human agent when needed. The emergence of intelligent customer service can significantly improve customer service efficiency and reduce costs, while also enhancing the customer experience and strengthening customer retention. Consequently, an increasing number of companies are adopting intelligent customer service. Currently, intelligent customer service is powered by spoken language understanding technology provided by task-based dialogue systems. These systems can meet the needs of users with specific goals, such as checking data usage, checking phone bills, ordering meals, booking tickets, and providing consultations.

[0098] Because user needs are relatively complex, multiple rounds of human-computer interaction are usually possible. Users may also continuously modify and improve their needs during the conversation. Intelligent customer service can help users clarify their goals by asking, clarifying, and confirming. In this process, intelligent customer service needs to understand the user's spoken words, including identifying the intent of the words and filling in the slots. The training method of the text understanding model and the spoken language understanding method provided in the embodiment of the present application can identify the intent of the user's words and fill in the slots, so that the intelligent customer can identify the user's speaking intentions and the slots in the speech content.

[0099] like Figure 3 As shown, a user logs into the XX travel platform on computer terminal 240 and uses the platform's intelligent client function to book air tickets. This intelligent client function enables interactive question-and-answer interaction between users and machines, providing intelligent services to users. Specifically, computer terminal 240 receives the dialogue "I want to fly to Taipei on December 3rd" input by user A and then uses text understanding model 301 to identify the intent of the dialogue and fill in the slots. Specifically, computer terminal 240 uses text understanding model 301 to understand the spoken language based on the word segmentation sequence and tag set of the input dialogue, identifying the intent of the dialogue as "booking a flight ticket" and extracting "Taipei" from the dialogue to fill in the "destination" slot and "December 3rd" to fill in the "departure time" slot. Computer terminal 240 can then perform downstream tasks based on the obtained dialogue intent and the corresponding slot content. For example, it can generate the question "Where are you departing from?" to obtain the user's departure location, thereby completing the subsequent flight booking service for the user.

[0100] (2) Application scenarios of AI medical guidance

[0101] The full-process AI medical guidance is based on artificial intelligence technologies such as speech recognition, speech synthesis, natural language understanding, and face recognition. It can realize the construction of application services in hospital outpatient clinics with intelligent pre-examination and triage, appointment registration, department query, recharge and payment, and health education as core functions, thereby fully meeting the patient's medical consultation and outpatient business processing needs, providing patients with timely and accurate medical information, and improving the patient's medical experience. The training method of the text understanding model and the spoken language understanding method provided in the embodiment of the present application can be applied to AI medical guidance. For example, when a patient seeks medical treatment in an outpatient clinic, he can input the department he wants to see a doctor through voice input or text on the AI ​​medical guidance terminal, and then the AI ​​medical guidance terminal can perform oral understanding based on the patient's input content, find the location of the department he wants to see a doctor, and display it to the patient.

[0102] like Figure 4As shown, the patient can input the dialogue "Where is the pediatric department" into the AI ​​medical guidance terminal 240, and then the AI ​​medical guidance terminal 240 can perform spoken language understanding based on the input dialogue word segmentation sequence and label set through the text understanding model 301, and recognize that the intention of the dialogue is "asking for directions", and then it can be judged that the patient's purpose is to know the location of a certain department in the hospital, and obtain "pediatrics" from the dialogue to fill in the slot "department", so that the AI ​​medical guidance terminal 240 gives a correct reply according to the pre-configured corresponding answer, that is, fills "Dongchuan Clinic Fourth Floor" in the slot "Location" and displays the reply sentence "Dongchuan Clinic Fourth Floor", thus completing the task of querying the department for the patient.

[0103] It should be noted that Figure 3 and Figure 4 These are only two schematic diagrams of application scenarios provided by the embodiments of the present application. The application scenarios described in the embodiments of the present application are only for the purpose of more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, in addition to the intelligent customer service application scenario and the AI ​​medical guidance application scenario, the training method of the text understanding model proposed in this application can also be used in application scenarios such as chatbots and intelligent retrieval. It is known to those skilled in the art that with the evolution of application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0104] According to one embodiment of the present application, a method for training a text comprehension model is provided. The method can be executed by a computer device (which can be a server, edge computing device, or other terminal device with certain processing capabilities) that has at least storage, computing, and communication functions. Figure 5 FIG. 1 shows a flow chart of the training method of the text understanding model provided in this embodiment. Figure 5 As shown, the training method of the text understanding model may specifically include:

[0105] Step 110: Obtain sample text and a sample label set, where the sample label set includes multiple sample intent labels and multiple sample slot labels.

[0106] In the spoken language understanding task of the task-based dialogue system, the relevant technology only performs intent recognition and slot filling based on the textual utterances input by the user. This way of performing spoken language understanding tasks ignores the rich semantic information in the intent labels and slot labels. In order to improve the accuracy of spoken language understanding, the training method of the text understanding model of this application proposes to make full use of the rich semantic information in the intent labels and slot labels. In this way, in addition to obtaining sample text, sample intent labels and sample slot labels are also required during the model training process.

[0107] The sample text refers to the utterance text, including questions, requirements, and dialogues in human-computer interaction. The sample intent label refers to the label used to describe the intent of the utterance. Optionally, a sentence can have one intent, or two or more intents. That is, an utterance text can correspond to at least one intent label. The sample slot label refers to the label used to describe the explicit instructions in the utterance. Figure 6 , Figure 6 A diagram showing an utterance and its intentions and slots. Figure 6 The utterance “I want to arrive Taipei on November 2” corresponds to the intent tags [Travel] and [Ticket booking], and the slot tags are [B-destination] and [I-arrive_time].

[0108] Usually, the slot filling task in spoken language comprehension is treated as a sequence labeling problem, so the format of the slot label is determined by the specific sequence labeling method. For example, Figure 6 The slot labels in the example are based on the BIO (B-begin, I-inside, O-outside) notation format, where "begin" indicates the beginning of an entity, "inside" indicates the middle or end of an entity, and "outside" indicates that it does not belong to the entity. The format of the sample slot labels can optionally be configured based on the specific sequence annotation method and is not specified here.

[0109] As an embodiment, when the training method of the text understanding model as described above is completely implemented in the terminal 240, the method for obtaining the sample text and the sample label set may include receiving object input, obtaining from a database, etc. Among them, receiving object input refers to receiving the speech input by the object on the terminal 240. For example, when a user uses a chatbot on the terminal 240, the terminal 240 obtains the text record generated by the user's voice input and the human-computer interaction with the chatbot, and then a large number of sentences can be extracted from the text record, and a large number of sample texts can be obtained by performing data preprocessing on the text record. Furthermore, a large number of sample slot labels can be manually set based on different sequence annotation methods, and a large number of sample intent labels can be manually set, thereby forming a sample label set. Obtaining from a database means that the terminal 240 obtains sample text and a sample label set from a local database or a cloud database. For example, a laptop obtains text records generated by the user's interaction with the platform from the database in the cloud server of the intelligent retrieval platform, and then a large number of sample texts can be extracted from the text record, and a large number of sample slot labels and sample intent labels can be manually set.

[0110] As another embodiment, when the training method of the text understanding model as described above is completely implemented on the server 210, the method of obtaining the sample text and the sample label set may include receiving it from downstream tasks, directly obtaining it from the Internet, etc. Herein, the downstream task refers to the task to which the text understanding model of the embodiment of the present application is applied, such as intelligent customer service, AI medical guidance, chatbots, data intelligent retrieval, etc. Taking the chatbot as an example, the server 210 can receive a large amount of utterances input by the terminal user sent by the terminal 240, and then generate a large amount of sample text and sample label sets through data processing. Obtaining it from the Internet means that the server 210 can query the public training data sets for spoken language comprehension tasks from the Internet, such as the MixATIS data set and the MixSNIPS data set.

[0111] In this way, the preset text understanding model can be trained based on the obtained sample text and sample label set. Figure 7 , Figure 7 A schematic diagram of the architecture of a preset text understanding model is shown. The preset text understanding model may include a text encoder, a label encoder, a prompt learner, an intent decoder, and a slot decoder. The text encoder may be used to encode sample text to obtain a corresponding text representation. The label encoder and the prompt learner may be used together to encode sample intent labels and sample slot labels to obtain corresponding label representations, and the weight parameters of the label encoder are fixed. The intent decoder may be used for intent decoding processing to obtain a predicted intent result. The slot decoder may be used for slot decoding processing to obtain a predicted slot result.

[0112] Step 120: Input multiple sample intent labels and multiple sample slot labels into a preset text understanding model for semantic relationship encoding processing to obtain corresponding intent label relationship representation and slot label relationship representation.

[0113] Considering that the relevant technologies ignore the rich semantic information in intent labels and slot labels when identifying and filling in intent in a joint modeling multi-task framework, and that the intent identification and slot filling tasks in spoken language comprehension are essentially supervised learning that uses the semantic information of the utterance as supervisory information, the joint modeling of the relevant technologies has the problem of inaccurate intent identification and slot filling due to the lack of label semantic information. To this end, this application proposes a method for training a text understanding model by combining sample text and sample label sets.

[0114] Among them, the label encoder and prompt learner in the preset text understanding model are used to encode multiple sample intent labels and multiple sample slot labels. In an embodiment of the present application, when the preset text understanding model is trained, the weight parameters of the label encoder are fixed. This is because the scale of the weight parameters of the label encoder is large, which easily leads to overfitting in model training. Therefore, the present application proposes to replace the model training method of directly using the label encoder with a larger parameter scale for label encoding by setting a prompt learner with a smaller parameter scale.

[0115] In some embodiments, the label encoder and prompt learner of the preset text understanding model can respectively perform semantic relationship encoding processing on multiple sample intent labels and multiple sample slot labels to obtain corresponding intent label relationship representations and slot label relationship representations, which may include:

[0116] (1) Input multiple sample intent labels into the label encoder for intent label encoding to obtain sample intent label representation;

[0117] (2) Input the sample intent label representation into the prompt learner for intent prompt learning to obtain the intent label relationship representation;

[0118] (3) Inputting multiple sample slot labels into a label encoder for slot label encoding to obtain a sample slot label representation;

[0119] (4) Input the sample slot label representation into the prompt learner for slot prompt learning to obtain the slot label relationship representation.

[0120] The label encoder can be used to encode the sample intent label and the sample slot label into a feature representation in the embedding space. The network structure of the label encoder can be a self-attention encoder (Self-Attention Encoder), a recurrent neural network (RNN), etc., without limitation here. For example, the label encoder can be an encoder based on self-attention and a bidirectional long-short term memory (BiLSTM) network.

[0121] The hint learner can be used to further learn the output of the label encoder. The setting of the hint learner is to solve the problem of the large parameter scale of the label encoder. Therefore, the weight parameter scale of the hint learner should be relatively smaller than the weight parameter scale of the label encoder. For example, Figure 8The prompt learner shown in FIG. 1 may have a network structure of a multilayer perceptron (MLP). After the output of the label encoder is processed by the MLP, the output of the MLP may be matrix-multiplied with the prompt vector to obtain the final label representation.

[0122] In step (1), multiple sample intent labels Input to tag encoder Φ intent , and then the label encoder can encode multiple sample intent labels to obtain the sample intent label representation corresponding to multiple sample intent labels Among them, T int is a matrix composed of the sample intent label representation of each sample intent label, d refers to the feature dimension of the sample intent label representation of each sample intent label, and |δ| refers to the number of categories of the sample intent label.

[0123] In step (2), the sample intent label representation is input into the prompt learner for intent mapping processing to generate the intent mapping probability distribution, and matrix multiplication is performed based on the prompt vector and the intent mapping probability distribution to generate the intent label relationship representation, where the prompt vector is obtained by random initialization.

[0124] For example, the sample intent label is represented by T int Input to the multi-layer perceptron MLP(·) in the prompt learner, and the multi-layer perceptron generates the intention mapping probability distribution Among them, N δ′ Refers to the number of sample intent labels.

[0125] Furthermore, the prompt vector is obtained by random initialization And the prompt vector h and the intention mapping probability distribution P int Perform matrix multiplication to obtain the intent label relationship representation k int =P int h, where

[0126] In step (3), multiple sample slot labels are added Input to tag encoder Φ intent , and then the label encoder can encode multiple sample slot labels to obtain the sample slot label representation corresponding to multiple sample slot labels Among them, T slo is a matrix composed of the sample slot label representation of each sample slot label, d refers to the feature dimension of the sample slot label representation of each sample slot label, and |δ| refers to the number of categories of the sample slot label.

[0127] In step (4), the sample slot label representation is input into the prompt learner for slot mapping processing to generate a slot mapping probability distribution, and matrix multiplication is performed based on the prompt vector and the slot mapping probability distribution to generate a slot label relationship representation, where the prompt vector is obtained by random initialization.

[0128] For example, the sample slot label is represented by T slo Input to the multilayer perceptron MLP(·) in the prompt learner, and the multilayer perceptron generates the slot mapping probability distribution Among them, N δ′ Refers to the quantity indicated by the sample slot label.

[0129] Furthermore, the prompt vector is obtained by random initialization And the prompt vector h and slot mapping probability distribution P slo Perform matrix multiplication to obtain the slot label relationship representation k slo =P slo h, where

[0130] Through the label encoder and the prompt learner, multiple sample intent labels and multiple sample slot labels can be encoded in a progressive and hierarchical manner respectively to obtain corresponding intent label relationship representations and slot label relationship representations. In this way, multiple sample intent labels are encoded into the embedding space, so that the supervised learning of spoken language comprehension can capture the semantic relationship information between multiple sample intent labels in the embedding space as supervision information. Similarly, the semantic relationship information between multiple sample slot labels can also be captured in the embedding space as supervision information, thereby improving the prediction accuracy of the preset text understanding model for the intent and slot of the discourse during supervised learning.

[0131] Step 130: Based on the preset text understanding model, the sample text representation and the intention label relationship representation of the sample text are subjected to intent decoding processing to obtain a predicted intention result, and based on the preset text understanding model, the sample text representation and the slot label relationship representation are subjected to slot decoding processing to obtain a predicted slot result.

[0132] In addition to the rich semantic information contained in intent and slot labels, when performing intent recognition and slot filling tasks on utterances, the intent or slot itself is semantically related to the user's utterance. However, related technologies still ignore the semantic relationship between them. To this end, this application proposes a method for joint representation decoding of utterances and labels to utilize the semantic information between intent / slot and utterance.

[0133] The intent label relationship indicates the semantic relationship between multiple sample intent labels. Therefore, based on the semantic relationship between multiple sample intent labels and the semantic relationship between the sample text and the sample intent label, the sample text can be decoded for intent to obtain a predicted intent result. Similarly, the slot label relationship indicates the semantic relationship between multiple sample slot labels. Therefore, based on the semantic relationship between multiple sample slot labels and the semantic relationship between the sample text and the sample slot label, the sample text can be slot decoded to obtain a predicted slot result.

[0134] In some embodiments, the preset text understanding model includes an intent decoder and a slot decoder. The intent decoder can be used to perform intent decoding processing on the sample text representation and the intent label relationship representation of the sample text, and the slot decoder can be used to perform slot decoding processing on the sample text representation and the slot label relationship representation. Specifically, it may include:

[0135] (1) Obtaining a sample text representation corresponding to the sample text;

[0136] (2) Input the sample text representation and the intent label relationship representation into the intent decoder for intent decoding processing to obtain the predicted intent result;

[0137] (3) The sample text representation and the slot label relationship representation are input into the slot decoder for slot decoding processing to obtain the predicted slot result.

[0138] Among them, the sample text representation includes the word segmentation representation corresponding to each word segment in the sample text. The intent decoding processing performed by the intent decoder can be a single-label single-intent prediction or a multi-label multi-intent prediction. Among them, the predicted intent result of the single-intent prediction can be understood as an intent label, and the predicted intent result of the multi-intent prediction can be understood as multiple intent labels. The slot decoding processing performed by the slot decoder is to find specific key information from the sample text and fill it into the preset slot, that is, the slot label.

[0139] The preset text understanding model may include a text encoder. The structure of the text encoder may be the same as the structure of the label encoder for intent / slot encoding. Step (1) may include:

[0140] (1.1) Perform word segmentation processing on the sample text to obtain a sample word segmentation sequence corresponding to the sample text;

[0141] The word segmentation process in text processing takes semantics into consideration. Different word segmentation granularities can be used to segment the sample text according to the needs of semantic analysis. The word segmentation granularity can include character granularity, word granularity and sub-word granularity. Word segmentation based on word granularity is the most intuitive word segmentation, which is similar to the process of human understanding of natural language. The division of word granularity can retain more semantic information. Word segmentation based on character granularity is an extreme word segmentation. This type of method directly decomposes the text into the smallest granularity. For English, it is to decompose it into letters; for Chinese, it is to decompose it into Chinese characters. In order to cover as many words as possible with the smallest possible vocabulary, the embodiment of the present application can use sub-word granularity for word segmentation.

[0142] For example, given the sample text "he is likely to be unfriendly to me", we perform word segmentation on the sample text based on subword granularity, obtaining the corresponding sample word sequence ['he', 'is', 'like', 'ly', 'to', 'be', 'un', 'friend', 'ly', 'to', 'me']. Subword granularity word segmentation uses small tokens in the vocabulary to form larger words in the sentence, similar to the combination of affixes in English. For example, given the word "unfriendly" in the above example sentence, it can be split into the following tokens: un, friend, and ly. These tokens themselves have semantics, such as "un" indicating negation, "friend" indicating friendliness, and "ly" indicating quality. Furthermore, tokens can be combined with each other to form new words. For example, "like" and "ly" can be combined into "likely". This helps the model learn the forms and variations between words and also makes the use of word segmentation more generalizable across different tasks.

[0143] (1.2) Input the sample word segmentation sequence into the text encoder for text encoding to obtain the sample text representation.

[0144] As an implementation method, a self-attention encoder with BiLSTM and self-attention mechanism can be used to obtain a sample text representation of a word segmentation sequence. The sample text representation can incorporate temporal features into word order and context information, so that the word segmentation representation of each word in the sample word segmentation sequence can carry context information to improve the accuracy of the text representation.

[0145] For example, the sample word segmentation sequence can be input into the self-attention encoder in two rounds to obtain the sample text representation E δ ,δ∈{I,S}, where E I refers to the sample text representation used for intent decoding, E S refers to the sample text representation used for slot decoding, E I and E Sis the representation matrix composed of all word segmentation representations. In addition, the sample word segmentation sequence can also be input into the pre-trained language model encoder to obtain the sample text representation E δ ,δ∈{I,S}, where E I =E S , that is, the sample text representation can be used for both intent decoding and slot decoding.

[0146] In step (2), to improve the accuracy of intent recognition, token-level intent recognition can be used, that is, identifying the intent of each token in the sentence and obtaining the sentence intent prediction result by voting on all tokens. Specifically, it may include:

[0147] (2.1) Perform representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation;

[0148] Among them, the enhanced representation of sample word segmentation refers to encoding the word segmentation representation under the intent prediction task to obtain a representation vector specifically used for intent decoding processing. In this way, each word segmentation representation in the enhanced sample text representation can more accurately represent the characteristics of the text data during intent prediction, thereby improving the accuracy of the intent decoding processing.

[0149] As an implementation method, the sample text encoded by the context can be represented by E I The input is fed into an intent-aware BiLSTM for representation enhancement to enhance its task-specific representation, for example, intent prediction, and then the sample word segmentation enhanced representation is obtained. The calculation formula can be shown as follows:

[0150]

[0151] in, is the enhanced representation of the word segmentation of the i-th sample, W I is the weight parameter, b I is the regularization term, E is the sample text representation I The segmentation representation corresponding to the i-th segmentation in .

[0152] (2.2) Input the enhanced representation of each sample word segmentation and the intent label relationship representation into the intent decoder for intent decoding processing, and obtain the sample intent score vector corresponding to each sample word segmentation in the sample text;

[0153] Each element in the sample intent score vector is the probability value of the likelihood that the word in the sample text belongs to each preset intent. Preset intent refers to a specific number of pre-set intent labels, where the preset intent can be determined by the specific intent prediction scenario. For example, the sample intent score vector corresponding to each sample word in the sample text can be calculated as follows:

[0154]

[0155] in, is the sample intention score vector corresponding to the i-th word in the sample text, W I is the weight parameter of the intent decoder, b I is the regularization term, k int It represents the intent label relationship.

[0156] (2.3) Voting is performed on each sample intent score vector to obtain the predicted intent result corresponding to the sample text.

[0157] As an implementation method, the intent scores corresponding to all predicted intents in the sample intent score vector of each sample segmentation can be obtained, and voting can be performed based on the voting threshold and the intent score corresponding to each predicted intent to obtain the voting result corresponding to each predicted intent, and then, based on the voting result corresponding to each predicted intent, the predicted intent result corresponding to the sample text can be determined. Among them, each element (that is, the intent score) in the sample intent score vector corresponding to the sample segmentation is the predicted probability that the sample segmentation belongs to a certain intent (that is, the predicted intent). By voting on the intent scores at the same position in each sample intent score vector, the voting result corresponding to each predicted intent can be obtained. The voting process can be voting for predicted intents with intent scores greater than 0.5, and by voting on the intent scores corresponding to the same predicted intent in each sample intent score vector, the number of votes corresponding to each predicted intent can be obtained.

[0158] For example, the sample text includes three segmentations {I1, I2, I3}, and the sample intent score vectors corresponding to segmentations I1, I2, and I3 are The sample intent score vector of each word segment includes four elements, which represent the possibility of four predicted intents. For example, The first element in "0.9", The first element "0.8" in The first element "0.9" in the ___ represents the possibility of the intention being "booking a ticket". The voting threshold is 0.5, and the voting process is to vote for each predicted intention with a probability greater than 0.5, such as voting for the intention of "booking a ticket". The first element in "0.9", The first element "0.8" in If the first element "0.9" in the list is greater than 0.5, then 3 votes can be cast for the intention of "booking tickets". The same voting process can be used for The second, third, and fourth elements in the sentence are voted for, and the corresponding predicted intentions (taking a car, traveling, and climbing) are obtained as follows: 1, 2, and 0. Furthermore, the predicted intentions with more than half the number of word segments are used as the predicted intention results corresponding to the sample text. That is, the predicted intentions "booking tickets" with 3 votes and "travel" with 2 votes in the voting results are used as the predicted intention results y corresponding to the sample text. I .

[0159] In step (3), the sample text representation and the slot label relationship representation are input into the slot decoder for slot decoding processing to obtain the predicted slot result, which may specifically include:

[0160] (3.1) Perform representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation;

[0161] Among them, the enhanced representation of sample word segmentation refers to encoding the word segmentation representation under the slot prediction task to obtain a representation vector specifically used for slot decoding processing, thereby enhancing each word segmentation representation in the sample text representation to more accurately represent the characteristics of the text data during slot prediction and improve the accuracy of the slot decoding processing.

[0162] As an implementation method, the sample text encoded by the context can be represented by E S The input is fed into an intent-aware BiLSTM for representation enhancement to enhance its task-specific representation, such as slot prediction, and then the sample segmentation enhanced representation is obtained. The calculation formula can be shown as follows:

[0163]

[0164] in, is the enhanced representation of the word segmentation of the i-th sample, W S is the weight parameter, b S is the regularization term, E is the sample text representation S The segmentation representation corresponding to the i-th segmentation in .

[0165] (3.2) Input each sample word segmentation enhanced representation and slot label relationship representation into the slot decoder for slot decoding processing to obtain the sample slot score vector;

[0166] Each element in the sample slot score vector is the probability value of the possibility that the word in the sample text belongs to each preset slot. The preset slot refers to a specific number of slot labels that are pre-set, where the preset slot can be determined by a specific slot prediction scenario. For example, the sample slot score vector corresponding to each sample word in the sample text can be calculated as follows:

[0167]

[0168] in, is the sample slot score vector corresponding to the i-th word in the sample text, W S is the weight parameter of the slot decoder, b S is the regularization term, k slo It represents the slot label relationship.

[0169] (3.3) performing a slot global prediction on the sample slot score vector corresponding to each sample word segmentation enhanced representation based on the first activation function, and obtaining a slot global prediction result corresponding to each sample word segmentation enhanced representation;

[0170] (3.4) Slot decoding is performed on the global prediction result of each slot according to the second activation function to obtain the predicted slot result.

[0171] As an implementation method, the sample slot score vector corresponding to each sample word segmentation enhanced representation can be input into the Softmax activation function for slot global prediction to obtain the slot global prediction result corresponding to each sample word segmentation enhanced representation, and each slot global prediction result is input into the Argmax activation function for slot decoding to obtain the predicted slot result. The calculation formula can be as follows:

[0172]

[0173] Among them, G i Perform global slot prediction for the mortal corresponding to the i-th sample segmentation, O i The predicted slot result corresponding to the i-th sample word segmentation, W g are trainable parameters.

[0174] In an embodiment of the present application, the intent label relationship representation indicates a first semantic relationship between a plurality of sample intent labels, so that the intent decoder can perform intent decoding processing on the sample text based on the first semantic relationship and the second semantic relationship between the sample text and the sample intent label. Similarly, the slot label relationship representation indicates a third semantic relationship between a plurality of sample slot labels, so that the slot decoder can perform slot decoding processing on the sample text based on the third semantic relationship and the fourth semantic relationship between the sample text and the sample slot label.

[0175] Compared with related technologies, the labels in the two subtasks of predicting intent and slot are converted into orthogonal one-hot vectors in the embedding space, which makes the intrinsic relationship between labels unused, and even more so, the semantic relationship between the intent or slot itself and the semantic relationship between the utterance cannot be utilized. However, this application introduces a label encoder with the same architecture as the text encoder, which can explicitly encode the intent / slot labels.

[0176] In this way, in the process of predicting intent and slots, the semantic relationship information between intent / slot labels in the embedding space, as well as the semantic relationship information between intent / slot labels and discourse can be effectively captured, thereby providing strong supervision information for the training of the preset text understanding model, thereby improving the prediction accuracy of the preset text understanding model for the intent labels and slot labels of the discourse.

[0177] Secondly, this application sets up the calculation of sample intent / slot score vectors specific to the intent prediction task and slot prediction task, so as to accurately determine the intent prediction and slot prediction corresponding to the entire text based on the intent prediction and slot prediction of each word segment, making the intent prediction task and slot prediction task explainable.

[0178] Step 140: Determine the target loss based on the target intent label corresponding to the predicted intent result and the sample text, and the target slot label corresponding to the predicted slot result and the sample text.

[0179] As an implementation method, the intent prediction task and the slot prediction task can be jointly trained through multi-task learning. Specifically, a first loss is determined based on the predicted intent result and the target intent label corresponding to the sample text, and a second loss is determined based on the predicted slot result and the target slot label corresponding to the sample text. Furthermore, the first loss and the second loss are weighted to obtain the target loss. The calculation formula can be as follows:

[0180]

[0181]

[0182] Among them, Loss1 is the first loss, N I is the number of sample intent labels, is the target intent label corresponding to the sample text, To predict the intent result. Similarly, determine the second loss:

[0183]

[0184] Among them, Loss2 is the second loss, N S is the number of sample slot labels, is the target slot label corresponding to the sample text, To predict the slot result. Further, the first loss and the second loss are weighted to obtain the target loss:

[0185] L=αLoss1+βLoss2 Formula (9)

[0186] Among them, L is the target loss, and α and β are hyperparameters.

[0187] Step 150: Iteratively train the preset text understanding model based on the target loss until the training end condition is reached to obtain a trained text understanding model.

[0188] As an implementation method, the text encoder, prompt learner, intent decoder, and slot decoder of a preset text understanding model can be iteratively trained based on the target loss until the training end condition is reached to obtain a trained text understanding model. The preset conditions may be: the total loss value of the target loss is less than a preset value, the total loss value of the target loss no longer changes, or the number of training times reaches a preset number. Optionally, an optimizer can be used to optimize the target loss, and the learning rate, batch size during training, and epoch of training can be set based on experimental experience.

[0189] To further evaluate the performance of the trained text understanding model, you can set up specific experiments for multiple intent detection and slot filling. For evaluation metrics, you can use Accuracy (Acc) to evaluate multiple intent detection, F1 score to evaluate slot filling, and Overall Acc to evaluate sentence-level semantic frame parsing. Overall Acc represents the proportion of utterances for which both intent and slot are correctly predicted.

[0190] The datasets used in the experiment are MixATIS and MixSNIPS. Figure 9 As shown in the dataset diagram, in the MixATIS dataset, all utterances can be split into training / development / testing, with the corresponding number of utterances being 13,162 / 756 / 828. In the MixSNIPS dataset, all utterances can be split into training / development / testing, with the corresponding number of utterances being 39,776 / 2,198 / 2,199.

[0191] As an implementation method, the text prediction model proposed in this application (hereinafter referred to as "L A P A ”) and compared with several state-of-the-art methods in related technologies. The experimental results are shown in Figure 10 The table shown, from which we can make the following observations:

[0192] (1) The L proposed in this application A P A On all tasks and datasets, it consistently improves the performance of several baselines compared to related techniques. A P A The best variant (i.e., Co-guiding Net+L A P A ) achieved a new optimal result. This is because the method of this application utilizes the rich semantic information between labels and between labels and utterances, providing strong supervision information for the model.

[0193] (2) It is worth noting that the improvement of w / oPLMs is more obvious than that of w / PLMs. This is because PLM can already provide relatively rich semantic information, but the combination of PLM and the L A P A Combining can further improve L A P A performance.

[0194] (3) The improvement in overall accuracy is more obvious. This is because the L A P A The semantics of intent and slots are utilized to support the mutual transmission of information between the two, thus significantly improving the overall performance of the model.

[0195] Furthermore, the present application conducted an ablation study, and the experimental results are as follows: Figure 11 The table shown is composed of Figure 11 It can be observed that after removing each submodule in the present application, the indicators all decrease, proving that each submodule in the present application has a certain effect. s and k I They are the intention prompt learner and the slot prompt learner. The full model refers to the one with L A P A The baseline model.

[0196] This embodiment can obtain sample text and a sample label set, wherein the sample label set includes a plurality of sample intent labels and a plurality of sample slot labels. The plurality of sample intent labels and the plurality of sample slot labels are input into a preset text understanding model for semantic relationship encoding processing respectively, to obtain corresponding intent label relationship representations and slot label relationship representations. In this way, the plurality of sample intent labels can be encoded into an embedding space, so that the semantic relationship information between the plurality of sample intent labels can be captured in the embedding space as supervision information for supervised learning of intent prediction and slot prediction. Similarly, the semantic relationship information between the plurality of sample slot labels can also be captured in the embedding space as supervision information for supervised learning of intent prediction and slot prediction.

[0197] Furthermore, based on the preset text understanding model, the sample text representation and the intent label relationship representation of the sample text are subjected to intent decoding processing to obtain a predicted intent result, and based on the preset text understanding model, the sample text representation and the slot label relationship representation are subjected to slot decoding processing to obtain a predicted slot result. Then, based on the predicted intent result and the target intent label corresponding to the sample text, as well as the predicted slot result and the target slot label corresponding to the sample text, the target loss is determined, and the preset text understanding model is iteratively trained based on the target loss until the training end condition is reached to obtain a trained text understanding model. In this way, the semantic relationship information between intent / slot labels in the embedding space, as well as the semantic relationship information between intent / slot labels and discourse, can be effectively captured, thereby providing strong supervision information for the training of the preset text understanding model, so that the preset text understanding model's prediction accuracy for the intent label and slot label of the discourse is improved.

[0198] See also Figure 12 , Figure 12 The following is a flow chart of a method for understanding spoken language provided by one embodiment of the present application. The method can be executed by a computer device having at least storage, computing, and communication functions. The method may specifically include the following steps:

[0199] For example, a customer service provider's intelligent customer service platform can provide intelligent customer service to users. The customer service provider's cloud service terminal can perform steps 210 to 250 to train a text understanding model and develop a corresponding intelligent customer service terminal based on the text understanding model. Then, the user can download and install the intelligent customer service terminal on the smartphone through the app store, log in to the intelligent customer service terminal, and use the corresponding service functions.

[0200] For example, by talking with the customer service robot of the intelligent customer service end, you can make an air ticket reservation. Then, the intelligent customer service end recognizes the intent of the conversation text and fills the slots based on the text understanding model. Specifically, you can execute steps 260 and 270 to complete the air ticket reservation service. It should be noted that the above-mentioned intelligent customer service scenario is only used as an application scenario for explaining the technical solution of this application. The implementation scenario and deployment method of the solution are not fixed, and there may be other application scenarios and deployment methods, which are not limited here. For example, the text understanding model can also be directly deployed on the cloud service end, and the intelligent customer service end on the smartphone is only used for human-computer interaction and data transmission. The application scenario can also be an intelligent navigation service for an in-vehicle terminal.

[0201] Step 210: The computer device obtains sample text and a sample label set, where the sample label set includes multiple sample intent labels and multiple sample slot labels.

[0202] Step 220: The computer device inputs multiple sample intent labels and multiple sample slot labels into the text understanding model to perform semantic relationship encoding processing respectively, and obtains corresponding intent label relationship representation and slot label relationship representation.

[0203] Step 230: The computer device performs intent decoding processing on the sample text representation and the intent label relationship representation of the sample text based on the text understanding model to obtain a predicted intent result, and performs slot decoding processing on the sample text representation and the slot label relationship representation based on the text understanding model to obtain a predicted slot result.

[0204] Step 240: The computer device determines the target loss based on the predicted intent result and the target intent label corresponding to the sample text, and the predicted slot result and the target slot label corresponding to the sample text.

[0205] Step 250: The computer device iteratively trains the text understanding model based on the target loss until the training end condition is reached, thereby obtaining a trained text understanding model.

[0206] As an implementation, the execution process of steps 210 to 250 is the same as the execution process of steps 110 to 150 in the above embodiment, and is not further described here. The computer device can train a text understanding model based on steps 210 to 250 for intent recognition and slot filling.

[0207] Step 260: The computer device obtains the target text and a tag set, where the tag set includes multiple intent tags and multiple slot tags.

[0208] For example, a computer device can obtain a sentence U, also known as the target text, and then jointly perform the multi-task of multi-intent detection and slot filling. The output of the task is the text intent and text slot corresponding to the target text U. For example, a user enters the sentence "I would like to fly to Taipei on November 2" on the customer service interface of the intelligent customer service terminal. Then, the customer service robot responsible for customer service obtains the sentence and obtains a customer service-related tag set from the cloud server. This tag set may include multiple intent tags and multiple slot tags.

[0209] Step 270: The computer device inputs the target text, multiple intent tags, and multiple slot tags into the text understanding model for spoken language understanding, and obtains the text intent and text slots corresponding to the target text.

[0210] When the computer device obtains the target text and the tag set, it can use the text understanding model to perform spoken language understanding on the target text, that is, intent recognition and slot filling. Figure 12 , Figure 12 A processing flow chart of a text understanding model is shown. Figure 13 As shown in the figure, the text understanding model can include a text encoder, a label encoder, a prompt learner, an intent decoder, and a slot decoder. The specific processing of spoken language understanding can include:

[0211] (1) Input the target text's word segmentation sequence into the text encoder to obtain the text representation;

[0212] For example, the computer device may first perform word segmentation processing on the acquired speech U = [I would like to fly to Taipei on November 2] to obtain the corresponding word segmentation sequence V = {v1, v2, ..., v x}, and input the word segmentation sequence L into the text encoder for representation learning to obtain the corresponding text representation E T .

[0213] (2) inputting multiple intent tags into the tag encoder for intent encoding to obtain a first representation of the intent;

[0214] For example, a computer device may store multiple intent tags L I Input to the label encoder for intent encoding to obtain the first representation of the intent F I , where the first intention is F I is the representation matrix corresponding to multiple intent labels.

[0215] (3) Inputting multiple slot labels into a label encoder for slot encoding to obtain a first slot representation;

[0216] For example, a computer device may have multiple slot labels L S Input to the tag encoder for slot encoding to obtain the slot first representation F S , where the first slot represents F S It is the representation matrix corresponding to multiple slot labels.

[0217] (4) Inputting the first representation of the intention into the prompt learner for intention prompt learning to obtain the second representation of the intention;

[0218] For example, the computer device may first express the intent F I The multilayer perceptron in the prompt learner is input to generate the intention mapping probability distribution. Furthermore, the prompt vector is obtained by random initialization, and the prompt vector and the intention mapping probability distribution are matrix multiplied to obtain the second representation of the intention K. I , K I It is a representation matrix composed of representations corresponding to multiple intent labels.

[0219] (5) inputting the first representation of the slot into the prompt learner to perform slot prompt learning to obtain the second representation of the slot;

[0220] For example, a computer device may first represent the slot F S The input is sent to the multi-layer perceptron in the prompt learner, which generates the slot mapping probability distribution. Furthermore, the prompt vector is obtained by random initialization, and the prompt vector and the slot mapping probability distribution are matrix multiplied to obtain the slot second representation K S , K S It is a representation matrix composed of the representations corresponding to multiple slot labels.

[0221] (6) Inputting the text representation and the second representation of the intent into the intent decoder for intent decoding to obtain the text intent corresponding to the target text;

[0222] For example, a computer device may represent a text E T and the second representation of intention K I Input to the intent decoder for intent decoding processing to obtain the text intent corresponding to the target text Where m is the number of predicted intents. For example, the intent scores corresponding to all predicted intents in the intent score vector of each word segment can be obtained through intent decoding, and voting can be performed based on the voting threshold and the intent score corresponding to each predicted intent to obtain the voting result corresponding to each predicted intent. Then, the text intent corresponding to the target text can be determined based on the voting result corresponding to each predicted intent. I .

[0223] The target text U includes multiple word segments {v1, v2, ..., v9}, and the intent score vectors corresponding to word segments v1 to v9 are The sample intent score vector of each word segment includes 5 possible prediction intents. Voting is performed on each prediction intent with a probability greater than 0.5. For example, by voting on the 5 prediction intents separately, the voting results are {5, 4, 1}. Furthermore, the prediction intent with more than half of the votes is used as the prediction intent result corresponding to the sample text. That is, the prediction intents with votes of 5 and 4 in the voting results are used as the text intent corresponding to the target text. I ={Ticket booking, Travel}.

[0224] (7) inputting the text representation and the second slot representation into a slot decoder for slot decoding processing to obtain a text slot corresponding to the target text;

[0225] For example, a computer device may represent a text E T and the second representation of intention K S Input to the intent decoder for intent decoding processing to obtain the text intent corresponding to the target text Specifically, the output of the slot decoding process can be input into the Sigmoid activation function to perform slot global prediction, and the slot global prediction result corresponding to each word segmentation can be obtained. Each slot global prediction result can then be input into the Argmax activation function to obtain the text slot O. S ={B-destination:Tapei,I-arrive_time:November 2}.

[0226] See also Figure 14 , which shows a structural block diagram of a text comprehension model training device 400 provided in an embodiment of the present application. The device 400 may include:

[0227] A sample acquisition module 410 is configured to acquire sample text and a sample label set, wherein the sample label set includes a plurality of sample intent labels and a plurality of sample slot labels;

[0228] An encoding processing module 420 is configured to input the plurality of sample intent labels and the plurality of sample slot labels into a preset text understanding model for semantic relationship encoding processing, respectively, to obtain corresponding intent label relationship representations and slot label relationship representations;

[0229] A decoding processing module 430 is configured to perform intent decoding processing on the sample text representation and the intent label relationship representation of the sample text based on the preset text understanding model to obtain a predicted intent result, and to perform slot decoding processing on the sample text representation and the slot label relationship representation based on the preset text understanding model to obtain a predicted slot result;

[0230] A loss determination module 440 is configured to determine a target loss based on the predicted intent result and the target intent label corresponding to the sample text, and the predicted slot result and the target slot label corresponding to the sample text;

[0231] The model training module 450 is used to iteratively train the preset text understanding model based on the target loss until the training end condition is met, thereby obtaining a trained text understanding model.

[0232] In some embodiments, the preset text understanding model includes an intent decoder and a slot decoder; the decoding processing module 430 may include:

[0233] A text representation unit, configured to obtain a sample text representation corresponding to the sample text;

[0234] an intent prediction unit, configured to input the sample text representation and the intent label relationship representation into the intent decoder for intent decoding processing to obtain a predicted intent result; wherein the intent label relationship representation indicates a first semantic relationship between a plurality of sample intent labels, so that the intent decoder performs intent decoding processing on the sample text based on the first semantic relationship and a second semantic relationship between the sample text and the sample intent label;

[0235] The slot prediction unit is used to input the sample text representation and the slot label relationship representation into the slot decoder for slot decoding processing to obtain a predicted slot result; wherein, the slot label relationship representation indicates a third semantic relationship between multiple sample slot labels, so that the slot decoder performs slot decoding processing on the sample text according to the third semantic relationship and the fourth semantic relationship between the sample text and the sample slot label.

[0236] In some embodiments, the intention prediction unit can be specifically used to: perform representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation; input each sample word segmentation enhanced representation and the intention label relationship representation into the intention decoder for intention decoding processing to obtain a sample intention score vector corresponding to each sample word segmentation in the sample text; perform voting processing on each of the sample intention score vectors to obtain a predicted intention result corresponding to the sample text.

[0237] In some embodiments, the slot prediction unit can be specifically used to: perform representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation; input each sample word segmentation enhanced representation and the slot label relationship representation into the slot decoder for slot decoding processing to obtain a sample slot score vector; perform slot global prediction on the sample slot score vector corresponding to each sample word segmentation enhanced representation based on the first activation function to obtain a slot global prediction result corresponding to each sample word segmentation enhanced representation; perform slot decoding processing on each of the slot global prediction results according to the second activation function to obtain a predicted slot result.

[0238] In some embodiments, the preset text understanding model includes a label encoder and a prompt learner, and the encoding processing module may include:

[0239] an intent encoding unit, configured to input the plurality of sample intent labels into the label encoder for encoding the intent labels to obtain sample intent label representations;

[0240] an intention learning unit, configured to input the sample intention label representation into the prompt learner for intention prompt learning to obtain an intention label relationship representation;

[0241] a label encoding unit, configured to input the plurality of sample slot labels into the label encoder for slot label encoding to obtain a sample slot label representation;

[0242] The label learning unit is used to input the sample slot label representation into the prompt learner to perform slot prompt learning to obtain a slot label relationship representation.

[0243] In some embodiments, the intent learning unit can be specifically used to: input the sample intent label representation into the prompt learner for intent mapping processing to generate an intent mapping probability distribution; perform matrix multiplication calculation based on the prompt vector and the intent mapping probability distribution to generate an intent label relationship representation, and the prompt vector is obtained by random initialization.

[0244] In some embodiments, the label learning unit can be specifically used to: input the sample slot label representation into the prompt learner for slot mapping processing to generate a slot mapping probability distribution; perform matrix multiplication calculation based on the prompt vector and the slot mapping probability distribution to generate a slot label relationship representation, and the prompt vector is obtained by random initialization.

[0245] In some embodiments, the preset text understanding model includes a text encoder; the text representation unit can be specifically used to: perform word segmentation processing on the sample text to obtain a sample word segmentation sequence corresponding to the sample text; input the sample word segmentation sequence into the text encoder for text encoding to obtain a sample text representation.

[0246] In some embodiments, the model training module 450 can be specifically used to iteratively train the text encoder, prompt learner, intent decoder and slot decoder of the preset text understanding model based on the target loss until the training end condition is reached, thereby obtaining the trained preset text understanding model.

[0247] In some embodiments, the loss determination module may include:

[0248] A first determining unit, configured to determine a first loss based on the predicted intent result and a target intent label corresponding to the sample text;

[0249] A second determining unit, configured to determine a second loss based on the predicted slot result and a target slot label corresponding to the sample text;

[0250] A weighted determination unit is used to perform weighted calculation on the first loss and the second loss to obtain a target loss.

[0251] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, modules, etc. can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0252] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.

[0253] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0254] This application can obtain sample text and sample label sets, wherein the sample label set includes multiple sample intent labels and multiple sample slot labels. And the multiple sample intent labels and multiple sample slot labels are input into the preset text understanding model for semantic relationship encoding processing respectively, to obtain corresponding intent label relationship representation and slot label relationship representation. In this way, multiple sample intent labels can be encoded into the embedding space, so that the semantic relationship information between the multiple sample intent labels can be captured in the embedding space as supervision information for supervised learning of intent prediction. Similarly, the semantic relationship information between the multiple sample slot labels can also be captured in the embedding space as supervision information for supervised learning of intent prediction.

[0255] Furthermore, based on the preset text understanding model, the sample text representation and the intent label relationship representation of the sample text are subjected to intent decoding processing to obtain a predicted intent result, and based on the preset text understanding model, the sample text representation and the slot label relationship representation are subjected to slot decoding processing to obtain a predicted slot result. Then, based on the predicted intent result and the target intent label corresponding to the sample text, as well as the predicted slot result and the target slot label corresponding to the sample text, the target loss is determined, and the preset text understanding model is iteratively trained based on the target loss until the training end condition is reached to obtain a trained text understanding model. In this way, the semantic relationship information between intent / slot labels in the embedding space, as well as the semantic relationship information between intent / slot labels and discourse, can be effectively captured, thereby providing strong supervision information for the training of the preset text understanding model, so that the preset text understanding model's prediction accuracy for the intent label and slot label of the discourse is improved.

[0256] See also Figure 15 , which shows a structural block diagram of a spoken language understanding device 500 provided in an embodiment of the present application. The device 500 may include:

[0257] A data acquisition module 510 is configured to acquire a target text and a tag set, wherein the tag set includes a plurality of intent tags and a plurality of slot tags;

[0258] The spoken language comprehension module 520 is used to input the target text, the multiple intent labels and the multiple slot labels into the trained text comprehension model for spoken language comprehension, and obtain the text intent and text slots corresponding to the target text; wherein the trained text comprehension model is obtained by the training device according to the above-mentioned text comprehension model.

[0259] In some embodiments, the trained text understanding model includes a text encoder, a label encoder, a prompt learner, an intent decoder, and a slot decoder; the spoken language understanding module 520 may include:

[0260] A word segmentation label unit, used to input the word segmentation sequence of the target text into the text encoder to obtain a text representation;

[0261] a first encoding unit, configured to input the intent label into the label encoder for intent encoding to obtain a first representation of the intent;

[0262] a second encoding unit, configured to input the slot label into the label encoder for slot encoding to obtain a first slot representation;

[0263] A first learning unit is configured to input the first intention representation into the prompt learner to perform intention prompt learning to obtain a second intention representation;

[0264] A second learning unit is configured to input the first slot representation into the prompt learner to perform slot prompt learning to obtain a second slot representation;

[0265] A first decoding unit is configured to input the text representation and the second representation of the intention into the intention decoder for intention decoding to obtain a text intention corresponding to the target text;

[0266] The second decoding unit is used to input the text representation and the second slot representation into the slot decoder for slot decoding processing to obtain the text slot corresponding to the target text.

[0267] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, modules, etc. can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0268] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.

[0269] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0270] like Figure 16As shown, the embodiment of the present application further provides a computer device 600, which includes a processor 610, a memory 620, a power supply 630 and an input unit 640. The memory 620 stores a computer program. When the computer program is called by the processor 610, the various method steps provided in the above embodiments can be executed. It can be understood by those skilled in the art that the structure of the computer device shown in the figure does not constitute a limitation of the computer device, and it can include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:

[0271] The processor 610 may include one or more processing cores. The processor 610 utilizes various interfaces and circuits to connect various components within the battery management system. By running or executing instructions, programs, instruction sets, or program sets stored in the memory 620 and accessing data stored in the memory 620, the processor 610 performs various functions of the battery management system and processes data, as well as various functions of the computer device and processes data, thereby providing overall control of the computer device. Optionally, the processor 610 may be implemented in the form of at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 610 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 610 and may be implemented separately via a communication chip.

[0272] The memory 620 may include a random access memory 620 (Random Access Memory, RAM), and may also include a read-only memory 620 (Read-Only Memory). The memory 620 may be used to store instructions, programs, instruction sets, or program sets. The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the various method embodiments described above, etc. The data storage area may also store data created by the computer device during use (such as a phone book and audio and video data), etc. Accordingly, the memory 620 may also include a memory controller to provide the processor 610 with access to the memory 620.

[0273] The power supply 630 can be logically connected to the processor 610 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. The power supply 630 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0274] The input unit 640 can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0275] Although not shown, the computer device 600 may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 610 in the computer device loads the executable files corresponding to one or more computer program processes into the memory 620 according to the following instructions. The processor 610 then executes the data stored in the memory 620, such as the phone book and audio and video data, thereby implementing the various method steps provided in the aforementioned embodiments.

[0276] like Figure 17 As shown, the embodiment of the present application further provides a computer-readable storage medium 700, in which a computer program 710 is stored. The computer program 710 can be called by a processor to execute various method steps provided in the embodiment of the present application.

[0277] The computer-readable storage medium may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium 700 has storage space for a computer program that executes any of the method steps in the above embodiments. These computer programs can be read from or written to one or more computer program products. The computer program can be compressed in an appropriate form.

[0278] According to one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the various method steps provided in the above embodiments.

[0279] The terms "comprises" and "includes" in the specification of this application and the above-mentioned drawings, as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. The term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the functions of the module or unit.

[0280] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0281] It should be understood that in the description of the embodiments of the present application, multiple (or multiple items) means more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0282] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0283] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0284] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. It should also be understood that the various implementation methods provided in the embodiments of the present application can be arbitrarily combined to achieve different technical effects.

[0285] The above is only a preferred embodiment of the present application and does not constitute any form of limitation to the present application. Although the present application has been disclosed as above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to equivalent embodiments using the technical contents disclosed above without departing from the scope of the technical solution of the present application. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A training method for a text comprehension model, characterized in that: The method comprises: Obtaining sample text and a sample label set, wherein the sample label set includes multiple sample intent labels and multiple sample slot labels; Inputting the multiple sample intent labels and the multiple sample slot labels into a preset text understanding model for semantic relationship encoding processing, respectively, to obtain corresponding intent label relationship representation and slot label relationship representation; Performing intent decoding processing on the sample text representation and the intent label relationship representation of the sample text based on the preset text understanding model to obtain a predicted intent result, and performing slot decoding processing on the sample text representation and the slot label relationship representation based on the preset text understanding model to obtain a predicted slot result; Determine a target loss based on the predicted intent result and the target intent label corresponding to the sample text, and the predicted slot result and the target slot label corresponding to the sample text; The preset text comprehension model is iteratively trained based on the target loss until a training end condition is reached, thereby obtaining a trained text comprehension model.

2. The method according to claim 1, characterized in that The preset text understanding model includes an intent decoder and a slot decoder; performing intent decoding processing on the sample text representation of the sample text and the intent label relationship representation based on the preset text understanding model to obtain a predicted intent result, and performing slot decoding processing on the sample text representation and the slot label relationship representation based on the preset text understanding model to obtain a predicted slot result, including: Obtaining a sample text representation corresponding to the sample text; Inputting the sample text representation and the intent label relationship representation into the intent decoder for intent decoding processing to obtain a predicted intent result; The intent label relationship indicates a first semantic relationship between a plurality of sample intent labels, so that the intent decoder performs intent decoding processing on the sample text according to the first semantic relationship and a second semantic relationship between the sample text and the sample intent label; Inputting the sample text representation and the slot label relationship representation into the slot decoder for slot decoding processing to obtain a predicted slot result; Among them, the slot label relationship indicates a third semantic relationship between multiple sample slot labels, so that the slot decoder performs slot decoding processing on the sample text according to the third semantic relationship and the fourth semantic relationship between the sample text and the sample slot label.

3. The method according to claim 2, characterized in that The step of inputting the sample text representation and the intent label relationship representation into the intent decoder for intent decoding to obtain a predicted intent result includes: Performing representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation; Inputting the enhanced representation of each sample word segmentation and the intention label relationship representation into the intention decoder for intention decoding processing to obtain a sample intention score vector corresponding to each sample word segmentation in the sample text; Voting is performed on each of the sample intent score vectors to obtain a predicted intent result corresponding to the sample text.

4. The method according to claim 2, characterized in that The step of inputting the sample text representation and the slot label relationship representation into the slot decoder for slot decoding to obtain a predicted slot result includes: Performing representation enhancement on each word segmentation representation in the sample text representation to obtain a sample word segmentation enhanced representation corresponding to each sample word segmentation representation; Inputting the enhanced word segmentation representation of each sample and the slot label relationship representation into the slot decoder for slot decoding processing to obtain a sample slot score vector; Performing a slot global prediction on the sample slot score vector corresponding to each sample word segmentation enhanced representation based on the first activation function to obtain a slot global prediction result corresponding to each sample word segmentation enhanced representation; Slot decoding processing is performed on each of the slot global prediction results according to the second activation function to obtain a predicted slot result.

5. The method according to any one of claims 2 to 4, characterized in that The preset text understanding model includes a label encoder and a prompt learner, and the weight parameters of the label encoder are fixed; the inputting of the multiple sample intent labels and the multiple sample slot labels into the preset text understanding model for semantic relationship encoding processing respectively to obtain corresponding intent label relationship representation and slot label relationship representation, including: Inputting the plurality of sample intent labels into the label encoder for intent label encoding to obtain sample intent label representation; Inputting the sample intention label representation into the prompt learner to perform intention prompt learning to obtain the intention label relationship representation; Inputting the plurality of sample slot labels into the label encoder for slot label encoding to obtain a sample slot label representation; The sample slot label representation is input into the prompt learner to perform slot prompt learning to obtain a slot label relationship representation.

6. The method according to claim 5, characterized in that The step of inputting the sample intent label representation into the prompt learner for intent prompt learning to obtain the intent label relationship representation includes: Inputting the sample intent label representation into the prompt learner for intent mapping processing to generate an intent mapping probability distribution; Matrix multiplication is performed based on the prompt vector and the intent mapping probability distribution to generate an intent label relationship representation, where the prompt vector is obtained by random initialization.

7. The method according to claim 5, characterized in that The step of inputting the sample slot label representation into the prompt learner for slot prompt learning to obtain a slot label relationship representation includes: Inputting the sample slot label representation into the prompt learner for slot mapping processing to generate a slot mapping probability distribution; A slot label relationship representation is generated by performing matrix multiplication based on a hint vector and the slot mapping probability distribution, wherein the hint vector is obtained by random initialization.

8. The method according to claim 6 or 7, characterized in that The preset text understanding model includes a text encoder; and obtaining a sample text representation corresponding to the sample text includes: Performing word segmentation processing on the sample text to obtain a sample word segmentation sequence corresponding to the sample text; The sample word segmentation sequence is input into the text encoder for text encoding to obtain a sample text representation.

9. The method according to claim 8, characterized in that The iterative training of the preset text comprehension model based on the target loss until a training end condition is reached to obtain a trained text comprehension model includes: The text encoder, prompt learner, intent decoder and slot decoder of the preset text understanding model are iteratively trained based on the target loss until the training end condition is reached to obtain a trained text understanding model.

10. The method according to any one of claims 1 to 9, characterized in that Determining a target loss based on the predicted intent result and the target intent label corresponding to the sample text, and the predicted slot result and the target slot label corresponding to the sample text, includes: Determining a first loss according to the predicted intent result and the target intent label corresponding to the sample text; Determining a second loss based on the predicted slot result and the target slot label corresponding to the sample text; The first loss and the second loss are weightedly calculated to obtain a target loss.

11. A method for spoken language comprehension, characterized in that: The method comprises: Obtaining a target text and a tag set, wherein the tag set includes multiple intent tags and multiple slot tags; The target text, the multiple intent labels, and the multiple slot labels are input into a trained text comprehension model for spoken language comprehension to obtain text intent and text slots corresponding to the target text; wherein the trained text comprehension model is trained according to the training method of the text comprehension model described in any one of claims 1 to 10 above.

12. The method according to claim 11, characterized in that The trained text understanding model includes a text encoder, a label encoder, a prompt learner, an intent decoder, and a slot decoder; the target text, the multiple intent labels, and the multiple slot labels are input into the text understanding model for spoken language understanding to obtain the text intent and text slot corresponding to the target text, including: Inputting the word segmentation sequence of the target text into the text encoder to obtain a text representation; Inputting the intent label into the label encoder for intent encoding to obtain a first representation of the intent; Inputting the slot label into the label encoder for slot encoding to obtain a first slot representation; Inputting the first intention representation into the prompt learner to perform intention prompt learning to obtain a second intention representation; Inputting the first slot representation into the prompt learner to perform slot prompt learning to obtain a second slot representation; Inputting the text representation and the second representation of the intention into the intention decoder for intention decoding processing to obtain the text intention corresponding to the target text; The text representation and the second slot representation are input into the slot decoder for slot decoding processing to obtain the text slot corresponding to the target text.

13. A training device for a text comprehension model, characterized in that: The device comprises: A sample acquisition module, configured to acquire sample text and a sample label set, wherein the sample label set includes a plurality of sample intent labels and a plurality of sample slot labels; An encoding processing module, configured to input the plurality of sample intent labels and the plurality of sample slot labels into a preset text understanding model for semantic relationship encoding processing, respectively, to obtain corresponding intent label relationship representations and slot label relationship representations; a decoding processing module, configured to perform intent decoding processing on the sample text representation and the intent label relationship representation of the sample text based on the preset text understanding model to obtain a predicted intent result, and perform slot decoding processing on the sample text representation and the slot label relationship representation based on the preset text understanding model to obtain a predicted slot result; a loss determination module, configured to determine a target loss based on the predicted intent result and the target intent label corresponding to the sample text, and the predicted slot result and the target slot label corresponding to the sample text; The model training module is used to iteratively train the preset text understanding model based on the target loss until the training end condition is met to obtain a trained text understanding model.

14. A spoken language comprehension device, characterized in that: The device comprises: A data acquisition module, configured to acquire a target text and a tag set, wherein the tag set includes a plurality of intent tags and a plurality of slot tags; A spoken language comprehension module is used to input the target text, the multiple intent labels and the multiple slot labels into a trained text comprehension model for spoken language comprehension, and obtain the text intent and text slots corresponding to the target text; wherein the trained text comprehension model is trained according to the training method of the text comprehension model described in any one of claims 1 to 10 above.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the text understanding model training method according to any one of claims 1 to 10 or the spoken language understanding method according to claim 11 or 12 is implemented.

16. A computer device, characterized in that: include: Memory; A processor, wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the training method of the text understanding model according to any one of claims 1 to 10 or the spoken language understanding method according to claim 11 or 12 is implemented.

17. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the text understanding model training method according to any one of claims 1 to 10 or the spoken language understanding method according to claim 11 or 12 is implemented.

Citation Information

Cited By

  • Text image and formula image unified identification method and system, storage medium and equipment

    CN121366423A