An intention slot recognition method and device

By using the encoding and intention recognition model in the intelligent dialogue system, combined with the slot recognition model, the connection between intention and slot is established, the problem of low semantic understanding accuracy in multi-intention queries is solved, and more efficient intention and slot recognition is achieved.

CN115713086BActive Publication Date: 2025-06-20HISENSE VISUAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211222613.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2025-06-20
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

When multi-intention query in existing intelligent dialogue systems, the semantic understanding accuracy is low, and the traditional model fails to effectively establish the connection between intention and slots, resulting in insufficient slot prediction and fill functions.

Method used

By obtaining the text sequence of the statement to be identified, using the coding model for encoding processing and average pooling, combining the intention recognition model and the slot recognition model, identifying and splicing the intent results, establishing the connection between intention and slots, and achieving accurate identification of multiple intentions and slots.

Benefits of technology

It improves the accuracy of semantic understanding in intelligent dialogue systems, can more accurately identify intentions and slot results in multi-intention queries, and improves the system's understanding ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713086B_ABST
    Figure CN115713086B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for identifying intent slots, which are applied to the field of natural language processing and can improve the semantic understanding accuracy rate in an intelligent dialogue system. The method includes: obtaining a first text sequence corresponding to a sentence to be identified; encoding the first text sequence through a first encoding model to obtain a first word vector corresponding to the first text sequence; performing average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence; performing intent recognition on the target sentence vector through an intent recognition model to recognize at least one intent result; respectively splicing each intent result with the first text sequence to obtain at least one second text sequence; encoding each second text sequence through a second encoding model to obtain at least one second word vector, and each second word vector corresponds to a second text sequence; performing slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intent result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of natural language processing. More specifically, it relates to a method and device for intent slot recognition. Background Art

[0002] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Task-based dialogue belongs to the category of closed domains, where the content answers of the dialogue are unique and restricted. A framework is given to the other party when asking questions, allowing the other party to only choose answers within the framework. In an intelligent dialogue system, accurately identifying the user's intent and then determining the slots under the intent is the key to improving the capabilities of the dialogue system.

[0003] In existing intelligent dialogue systems, the functions of semantic understanding systems mostly focus on single-intent query recognition, that is, a single-sentence user query only involves a single intent. However, in real-world scenarios, user requirements are relatively complex, and there may be multi-intent queries. Many traditional multi-intent semantic understanding models only perform intent recognition and do not perform slot prediction and filling. Obviously, this method cannot fully meet the requirements of language understanding. Therefore, it is necessary to establish the connection between multi-intent and slots and improve the function of slot prediction and filling to solve the problem of low semantic understanding accuracy in existing intelligent dialogue systems. Summary of the Invention

[0004] To solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present application provide a method and device for intent slot recognition, which can improve the semantic understanding accuracy in an intelligent dialogue system.

[0005] In a first aspect, the embodiments of the present application provide a method for intent slot recognition, including:

[0006] Obtain a first text sequence corresponding to the statement to be recognized;

[0007] Perform encoding processing on the first text sequence through a first encoding model to obtain a first word vector corresponding to the first text sequence;

[0008] Perform average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence;

[0009] Perform intent recognition on the target sentence vector through an intent recognition model to recognize at least one intent result;

[0010] Concatenate each intent result with the first text sequence respectively to obtain at least one second text sequence;

[0011] Encoding each second text sequence respectively through a second encoding model to obtain at least one second word vector, where each second word vector corresponds to a second text sequence;

[0012] Identifying slots for each second word vector through a first slot identification model to obtain at least one slot result corresponding to each intent result.

[0013] In some embodiments of the present application, the obtaining of the first text sequence corresponding to the statement to be recognized includes:

[0014] Receiving the statement to be recognized;

[0015] Performing speech recognition on the statement to be recognized to obtain a third text sequence in the statement to be recognized, where the third text sequence includes multiple words;

[0016] Based on an explicit knowledge base, respectively performing knowledge base matching on each word in the multiple words to obtain at least one classification knowledge, where each classification knowledge corresponds to a word;

[0017] Inserting the at least one classification knowledge into the third text sequence respectively to obtain a first text sequence, in which each classification identifier is located behind the corresponding word and is separated by a preset separator.

[0018] In some embodiments of the present application, before the step of identifying slots for each second word vector through a first slot identification model to obtain at least one slot result corresponding to each intent result, the method further includes:

[0019] Marking the positions of each word in the third text sequence in the second text sequence to obtain first position information;

[0020] The step of identifying slots for each second word vector through a first slot identification model to obtain at least one slot result corresponding to each intent result includes:

[0021] Identifying each slot result corresponding to each intent result in the second text sequence through the first slot identification model for each second word vector;

[0022] Based on the first position information, filtering out at least one slot result corresponding to each intent result in the third text sequence from each slot result corresponding to each intent result.

[0023] In some embodiments of the present application, after the step of performing average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence, the method further includes:

[0024] The domain identification model is used to identify the domain of the target sentence vector, and at least one domain result is obtained, and each domain result corresponds to an intention result.

[0025] In some embodiments of the present application, splicing each intention result with the first text sequence respectively to obtain at least one second text sequence includes:

[0026] Splicing each intention result and the corresponding domain result with the first text sequence respectively to obtain the at least one second text sequence.

[0027] In some embodiments of the present application, before splicing each intention result with the first text sequence respectively to obtain at least one second text sequence, the method further includes:

[0028] Determining whether the number of the at least one intention result is greater than 1;

[0029] Splicing each intention result with the first text sequence respectively to obtain at least one second text sequence includes:

[0030] When the number of the at least one intention result is greater than 1, splicing each intention result with the first text sequence respectively to obtain at least one second text sequence;

[0031] The method further includes:

[0032] When the number of the at least one intention result is equal to 1, the second slot identification model is used to identify the slot of the first word vector to obtain at least one slot result corresponding to each intention result.

[0033] In some embodiments of the present application, after the first slot identification model is used to identify the slot of each second word vector to obtain at least one slot result corresponding to each intention result, the method further includes:

[0034] Based on the positional relationship of the specific slots in at least one slot result corresponding to each intention result in the first text sequence, determining the arrangement order of each intention in the first text sequence.

[0035] In a second aspect, an intention slot identification device provided by an embodiment of the present application includes:

[0036] An acquisition module, configured to acquire a first text sequence corresponding to a sentence to be identified;

[0037] An encoding module, configured to perform encoding processing on the first text sequence through a first encoding model to obtain a first word vector corresponding to the first text sequence;

[0038] A pooling module, configured to perform average pooling on the first word vector to obtain a target sentence vector corresponding to the first text sequence;

[0039] An intent recognition module, configured to perform intent recognition on the target sentence vector through an intent recognition model to recognize at least one intent result;

[0040] A splicing module, configured to splice each intent result with the first text sequence respectively to obtain at least one second text sequence;

[0041] The encoding module is further configured to perform encoding processing on each second text sequence through a second encoding model to obtain at least one second word vector, and each second word vector corresponds to a second text sequence;

[0042] A slot recognition module, configured to perform slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intent result.

[0043] In some embodiments of the present application, the obtaining module is configured to receive the sentence to be recognized;

[0044] Perform speech recognition on the sentence to be recognized to obtain a third text sequence in the sentence to be recognized, and the third text sequence includes multiple words;

[0045] Based on an explicit knowledge base, perform knowledge base matching on each of the multiple words respectively to obtain at least one classification knowledge, and each classification knowledge corresponds to a word;

[0046] Insert the at least one classification knowledge into the third text sequence respectively to obtain a first text sequence. In the first text sequence, each classification identifier is located behind the corresponding word and is separated by a preset separator.

[0047] In some embodiments of the present application, the apparatus further includes:

[0048] A marking module, configured to mark the positions of each word in the third text sequence in the second text sequence before performing slot recognition on each second word vector through the first slot recognition model to obtain at least one slot result corresponding to each intent result, so as to obtain first position information;

[0049] The slot recognition module is specifically configured to perform slot recognition on each second word vector through the first slot recognition model to obtain each slot result corresponding to each intent result in the second text sequence;

[0050] Based on the first position information, filter out at least one slot result corresponding to each intent result in the third text sequence from each slot result corresponding to each intent result.

[0051] In some embodiments of the present application, the apparatus further includes:

[0052] A domain recognition module, configured to perform domain recognition on the target sentence vector obtained by performing average pooling processing on the first word vector, to recognize at least one domain result, and each domain result corresponds to an intention result.

[0053] In some embodiments of the present application, the splicing module is specifically configured to splice each intention result and the corresponding domain result with the first text sequence respectively, to obtain the at least one second text sequence.

[0054] In some embodiments of the present application, the apparatus further includes:

[0055] A determination module, configured to determine whether the number of the at least one intention result is greater than 1 before splicing each intention result with the first text sequence respectively to obtain the at least one second text sequence;

[0056] The splicing module is specifically configured to, when the number of the at least one intention result is greater than 1, splice each intention result with the first text sequence respectively to obtain the at least one second text sequence;

[0057] The slot recognition module is further configured to, when the number of the at least one intention result is equal to 1, perform slot recognition on the first word vector through a second slot recognition model, to obtain at least one slot result corresponding to each intention result.

[0058] In some embodiments of the present application, the determination module is configured to, after performing slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intention result, determine the arrangement order of the respective intentions in the first text sequence based on the position relationship of a specific slot in the at least one slot result corresponding to each intention result in the first text sequence.

[0059] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, and when the computer program is executed by a processor, the intention slot recognition method as shown in the second aspect is implemented.

[0060] In a fourth aspect, an embodiment of the present application provides a computer program product, including: when the computer program product runs on a computer, the computer is enabled to implement the intention slot recognition method as shown in the second aspect.

[0061] The technical solution provided by the embodiment of the present application has the following advantages compared with the prior art: In the embodiment of the present application, a first text sequence corresponding to the statement to be recognized is obtained; the first text sequence is encoded by a first encoding model to obtain a first word vector corresponding to the first text sequence; average pooling is performed on the first word vector to obtain a target sentence vector corresponding to the first text sequence; at least one intention result is recognized by an intention recognition model for the target sentence vector; each intention result is respectively concatenated with the first text sequence to obtain at least one second text sequence; each second text sequence is respectively encoded by a second encoding model to obtain at least one second word vector, and each second word vector corresponds to a second text sequence; at least one slot result corresponding to each intention result is obtained by a first slot recognition model for each second word vector. In this solution, at least one intention result in the first text sequence corresponding to the statement to be recognized is first recognized, and then each recognized intention result is respectively concatenated with the first text sequence (to establish the connection between each intention result and its corresponding slot) to obtain at least one second text sequence, and then slot recognition is performed on each second text sequence to obtain at least one slot result corresponding to each intention result, so as to accurately recognize the intention and slot of the statement to be recognized, and the semantic understanding accuracy in the intelligent dialogue system can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present application or the implementation manners in the related art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the related art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0063] Figure 1 FIG. shows a scenario architecture diagram of an intention slot recognition method according to some embodiments;

[0064] Figure 2 FIG. shows a hardware configuration block diagram of an electronic device according to some embodiments;

[0065] Figure 3 FIG. shows an operating system schematic diagram of an electronic device and a server according to some embodiments;

[0066] Figure 4 FIG. shows a schematic diagram of a speech recognition network architecture according to some embodiments;

[0067] Figure 5 FIG. shows a software configuration diagram of an electronic device 200 according to some embodiments;

[0068] Figure 6Shows one of the flow diagrams of the intention slot recognition method according to some embodiments;

[0069] Figure 7 Shows another flow diagram of the intention slot recognition method according to some embodiments;

[0070] Figure 8 Shows a third flow diagram of the intention slot recognition method according to some embodiments;

[0071] Figure 9 Shows a fourth flow diagram of the intention slot recognition method according to some embodiments;

[0072] Figure 10 Shows a fifth flow diagram of the intention slot recognition method according to some embodiments;

[0073] Figure 11 Shows a sixth flow diagram of the intention slot recognition method according to some embodiments;

[0074] Figure 12 Shows the structural diagram of the intention slot recognition model according to some embodiments;

[0075] Figure 13 Shows the schematic diagram of the intention slot recognition framework according to some embodiments;

[0076] Figure 14 Shows the framework diagram of the intention slot recognition device according to some embodiments;

[0077] Figure 15 Shows the schematic diagram of the computer device hardware according to some embodiments. Detailed implementation manners

[0078] To make the objectives and implementation manners of this application clearer, the following will clearly and completely describe the exemplary implementation manners of this application with reference to the accompanying drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only a part, rather than all, of the embodiments of this application.

[0079] It should be noted that the brief descriptions of the terms in this application are only for facilitating the understanding of the subsequent described implementation manners, rather than intending to limit the implementation manners of this application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.

[0080] The terms "first", "second", "third", etc. in the specification, claims, and the above-mentioned drawings of this application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.

[0081] The terms "comprising" and "having" and any variations thereof are intended to cover inclusion without exclusivity. For example, a product or device that comprises a series of components need not be limited to all the components clearly listed, but may include other components not clearly listed or inherent to such products or devices.

[0082] Figure 1 The figure is a schematic diagram of the scenario architecture of an intention slot recognition method provided for the embodiments of this application. As Figure 1 shown, the scenario architecture provided for the embodiments of this application includes: a server 100 and an electronic device 200.

[0083] The electronic device 200 provided for the embodiments of this application can have various implementation forms. For example, it can be a smart speaker, a TV, a refrigerator, a washing machine, an air conditioner, smart curtains, a router, a set-top box, a mobile phone, a personal computer (PC), a smart TV, a laser projection device, a monitor, an electronic bulletin board, a wearable device, a vehicle-mounted device, an electronic table, etc.

[0084] In some embodiments, when the electronic device 200 receives a voice command from a user, it can communicate with the server 100. The electronic device 200 is allowed to communicate and connect with the server 100 through a local area network (LAN) or a wireless local area network (WLAN).

[0085] The server 100 can be a server that provides various services. For example, it is a server that provides support for the audio data collected by the electronic device 200. The server can analyze and process data such as the received audio, and feedback the processing results (such as endpoint information) to the electronic device 200. The server 100 can be a server cluster or multiple server clusters, and can include one or more types of servers.

[0086] The electronic device 200 can be hardware or software. When the electronic device 200 is hardware, it can be various electronic devices with a sound collection function, including but not limited to smart speakers, smart phones, TVs, tablets, e-book readers, smart watches, players, computers, AI devices, robots, smart vehicles, etc. When the electronic device 200 is software, it can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (for example, used to provide a sound collection service), or can be implemented as a single software or software module. No specific limitation is made here.

[0087] Exemplarily, the electronic device 200 receives a statement to be recognized issued by a user, and then sends the statement to be recognized to the server 100. The server 100 processes the statement to be recognized through the intent slot recognition method provided in the embodiments of the present application, obtains at least one intent result corresponding to the statement to be recognized and at least one slot result corresponding to each intent result, and then returns the at least one intent result corresponding to the statement to be recognized and the at least one slot result corresponding to each intent result to the electronic device 200. The electronic device 200 verbally broadcasts or displays the at least one intent result corresponding to the statement to be recognized and the at least one slot result corresponding to each intent result.

[0088] It should be noted that Figure 1 The shown scenario architecture diagram only shows one possible scenario for implementing the intent slot recognition method provided in this embodiment. The execution subject of the intent slot recognition method provided in the embodiments of the present application can be the above-mentioned server or a functional module or functional entity in the server that has the function of implementing the intent slot recognition method. The execution subject of the intent slot recognition method provided in the embodiments of the present application can also be the above-mentioned electronic device or a functional module or functional entity in the electronic device that has the function of implementing the intent slot recognition method.

[0089] Figure 2 Shows a block diagram of the hardware configuration of the electronic device 200 according to an exemplary embodiment. As Figure 2 shown, the electronic device 200 includes at least one of a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, an audio processor, RAM, ROM, and first to nth interfaces for input / output.

[0090] The communicator 220 is a component for communicating with external devices or servers according to various communication protocol types. For example: the communicator may include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module and other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The electronic device 200 can establish the sending and receiving of control signals and data signals with the server 100 through the communicator 220.

[0091] The user interface 280 can be used to receive external control signals.

[0092] The detector 230 is used to collect signals from the external environment or interact with the outside. For example, the detector 230 includes a light receiver, a sensor for collecting the ambient light intensity; or, the detector 230 includes an image collector, such as a camera, which can be used to collect the external environment scene, user attributes or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0093] The sound collector can be a microphone, also known as a "microphone" or "transmitter", which can be used to receive the user's voice and convert the sound signal into an electrical signal. The electronic device 200 can be provided with at least one microphone. In some other embodiments, the electronic device 200 can be provided with two microphones, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 200 can also be provided with three, four or more microphones to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.

[0094] In addition, the microphone can be built into the electronic device 200, or the microphone is connected to the electronic device 200 in a wired or wireless manner. Of course, the embodiments of the present application do not limit the position of the microphone on the electronic device 200. Or, the electronic device 200 may not include a microphone, that is, the above microphone is not provided in the electronic device 200. The electronic device 200 can externally connect a microphone (also called a microphone) through an interface (such as the USB interface 130). The externally connected microphone can be fixed on the electronic device 200 through an external fixing member (such as a camera bracket with a clip).

[0095] The controller 250 controls the operation of the electronic device and responds to the user's operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the electronic device 200.

[0096] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a random access memory (RAM), a read-only memory (ROM), a first interface to an nth interface for input / output, a communication bus, etc.

[0097] In some examples, taking the Android system as an example of the operating system of the electronic device, as Figure 3 shown, the electronic device 200 can be logically divided into an applications layer (abbreviation "application layer") 21, a kernel layer 22, and a hardware layer 23.

[0098] Among them, as Figure 3As shown, the hardware layer 23 may include Figure 2 the controller 250, communicator 220, detector 230, etc. shown. The application layer 21 includes one or more applications. The applications can be system applications or third-party applications. For example, the application layer 21 includes a voice recognition application, which can provide a voice interaction interface and service for the connection between the electronic device 200 and the server 100.

[0099] The kernel layer 22 serves as software middleware between the hardware layer 23 and the application layer 21, and is used to manage and control hardware and software resources.

[0100] For example Figure 3 As shown, the server 100 may include: a communication control module 101, an intent slot recognition module 102, and a data storage module 103, and may also include other modules, which are not limited here. Among them, the communication control module 101 is used to communicate with the communicator 220, and the data storage module 103 is used to store various databases. In the embodiments of the present application, it can be used to store tabular data.

[0101] In some examples, the kernel layer 22 includes a detector driver, which is used to send the voice data collected by the detector 230 to the voice recognition application. Exemplarily, when the voice recognition application in the electronic device 200 is started and the electronic device 200 has established a communication connection with the server 100, the detector driver is used to send the voice data of the user input collected by the detector 230 to the voice recognition application. Then, the voice recognition application sends the query information containing the voice data to the intent slot recognition module 102 in the server. The intent slot recognition module 102 is used to input the voice data sent by the electronic device 200 into the intent recognition model to obtain an intent recognition result, and then transmit the intent recognition result to the electronic device 200.

[0102] To clearly illustrate the embodiments of the present application, the following Figure 4 describes a voice recognition network architecture provided by the embodiments of the present application.

[0103] See Figure 4 , Figure 4 which is a schematic diagram of a voice recognition network architecture provided by the embodiments of the present application. Figure 4In it, the electronic device is used to receive the input information and output the processing result of the information. The speech recognition service device is an electronic device deployed with a speech recognition service, the semantic service device is an electronic device deployed with a semantic service, and the business service device is an electronic device deployed with a business service. The electronic devices here may include servers, computers, etc. The speech recognition service, semantic service (which can also be called a semantic engine), and business service here are web services that can be deployed on electronic devices. Among them, the speech recognition service is used to recognize audio as text, the semantic service is used to perform semantic parsing on the text, and the business service is used to provide specific services such as weather query services of Moji Weather, music query services of QQ Music, etc. In one embodiment, Figure 6 In the architecture shown, there may be multiple entity service devices deployed with different business services, or one or more entity service devices may integrate one or more functional services.

[0104] In some embodiments, the process of processing the information input to the electronic device based on Figure 4 the architecture shown is described by way of example below. Taking the information input to the electronic device as an identification statement input by voice as an example, the above process may include the following three processes:

[0105] [Speech Recognition]

[0106] After receiving the identification statement input by voice, the electronic device may upload the audio of the identification statement to the speech recognition service device, so that the speech recognition service device can recognize the audio as text through the speech recognition service and then return it to the electronic device. In one embodiment, before uploading the audio of the identification statement to the speech recognition service device, the electronic device may perform denoising processing on the audio of the identification statement. The denoising processing here may include steps such as removing echo and ambient noise.

[0107] [Semantic Understanding]

[0108] The electronic device uploads the text of the identification statement recognized by the speech recognition service to the semantic service device, so that the semantic service device can perform semantic parsing on the text through the semantic service to obtain the business field, intention, etc. of the text.

[0109] [Semantic Response]

[0110] The semantic service device issues a ╳ query instruction to the corresponding business service device according to the semantic parsing result of the text of the identification statement to obtain the query result given by the business service. The electronic device may obtain the query result from the semantic service device and output it. As an embodiment, the semantic service device may also send the semantic parsing result of the identification statement to the electronic device, so that the electronic device outputs the feedback statement in the semantic parsing result.

[0111] It should be noted that Figure 4The architecture shown is only an example and does not limit the protection scope of this application. In the embodiments of this application, other architectures can also be used to implement similar functions. For example, all or part of the three processes can be completed by a smart terminal, which will not be elaborated here.

[0112] Next, some terms and related technologies involved in this application will be explained to facilitate the understanding of those skilled in the art.

[0113] I. Intent and slot

[0114] (I). Definitions of intent and slot:

[0115] An intent refers to what the computer device recognizes as the actual or potential needs of the user. Fundamentally, an intent is a classifier that divides user needs into a certain type defined in advance.

[0116] An intent and slots together constitute a "user action". Since a computer device cannot directly understand natural language, the role of intent recognition is to map natural language into a structured semantic representation that the machine can understand.

[0117] Intent recognition, also known as SUC (Spoken Utterance Classification), as the name implies, is to classify the natural language conversation input by the user into categories. The corresponding category is the user's intent. For example, for "What's the weather like today", its intent is "inquiring about the weather". Naturally, intent recognition can be regarded as a typical classification problem.

[0118] A slot is the parameter carried by an intent. An intent may correspond to several slots. For example, when asking about the bus route, necessary parameters such as the departure place, destination, and time need to be given. The above parameters are the slots corresponding to the intent of "asking about the bus route".

[0119] For example, the main goal of the semantic slot filling task is to extract the values of the semantic slots predefined in the semantic frame from the input sentence on the premise of knowing the semantic frame of a specific domain or specific intent. The semantic slot filling task can be transformed into a sequence labeling task, that is, using the IOB tagging method to label whether a certain word is the beginning (begin), continuation (inside), or non-semantic slot (outside) of a certain semantic slot. To make a system work properly, intents and slots need to be designed first. Intents and slots can enable the system to know which specific task to execute and give the parameter types required when executing the task.

[0120] (2) Intent recognition and slot filling: After defining the intents and slots, the user intent and the slot values corresponding to the respective slots can be recognized from the user input.

[0121] The goal of intent recognition is to recognize the user intent from the input. A single task can be simply modeled as a binary classification problem. For example, the intent of "inquiring about the weather" can be modeled as a binary classification problem of "yes, inquiring about the weather" or "no, not inquiring about the weather" during intent recognition. When the system needs to handle multiple tasks, it needs to be able to distinguish each intent. In this case, the binary classification problem is transformed into a multi-classification problem. The task of slot filling is to extract information from the data and fill it into the pre-defined slots. For example, if the intent and the corresponding slots have been defined, for the user input "What's the weather like in Shanghai today", the system should be able to extract "today" and "Shanghai" and fill them into the "time" and "location" slots respectively.

[0122] II. BERT (Bidirectional Encoder Representations from Transformers) model: The BERT model is the encoder of the bidirectional Transformer. Among them, Transformer is a method that fully relies on self-attention to calculate the input and output representations. BERT uses the masked model to achieve the bidirectionality of the language model, proving the importance of bidirectionality for language representation pre-training. The BERT model is a truly bidirectional language model, and each word can utilize the context information of that word simultaneously. BERT aims to pre-train deep bidirectional representations by jointly adjusting the context in all layers. Therefore, the pre-trained BERT representation can be fine-tuned through an additional output layer to build state-of-the-art models suitable for a wide range of tasks.

[0123] After adding a fully connected layer to the BERT model for training, the BERT model after removing the fully connected layer can be used for various natural language processing tasks (including sequence labeling tasks, classification tasks, sentence relationship judgment, and generative tasks).

[0124] Figure 5 FIG. is a flowchart of the steps for implementing the intent slot recognition method according to one or more embodiments of the present invention. The execution subject of the intent slot recognition method can be a server or an electronic device, or a functional module or functional entity in the server or electronic device that can implement the intent slot recognition method, which is not limited herein. In the embodiments of the present application, taking the execution subject as a server as an example for illustrative purposes, the intent slot recognition method may include the following S501 to S507.

[0125] S501. Obtain a first text sequence corresponding to the statement to be recognized.

[0126] Among them, the first text sequence may be the text sequence corresponding to the sentence to be recognized. The first text sequence may include the text sequence corresponding to the sentence to be recognized and other text sequences, which can be specifically determined according to the actual situation and are not limited here.

[0127] S502. Encode the first text sequence through the first encoding model to obtain the first word vector corresponding to the first text sequence.

[0128] Among them, the first encoding model may be a pre-trained language model such as a BERT model, a RoBerta model, an Electra model, etc., which can be specifically determined according to the actual situation and are not limited here. In the embodiments of the present application, the first encoding model is taken as an example of a BERT model for illustration.

[0129] Among them, the word vector (sequence output) is the vector representation of each word in the text sequence, and the first word vector is the vector representation of each word in the first text series.

[0130] S503. Perform average pooling processing on the first word vector to obtain the target sentence vector corresponding to the first text sequence.

[0131] Exemplarily, input the first text sequence into the BERT model, and output the first word vector corresponding to the first text sequence. Assume that the first word vector is H = {h1, h2,..., h n}. Perform average pooling processing on the first word vector to obtain a new vector. At this time, it is considered that this new vector synthesizes the information of each word in the sentence, and use it as the sentence vector C representing the sentence (that is, obtain the target sentence vector). Among them, the average pooling processing formula can be:

[0132]

[0133] S504. Perform intent recognition on the target sentence vector through the intent recognition model to recognize at least one intent result.

[0134] Among them, the intent recognition model can be pre-trained in advance, which can be specifically determined according to the actual situation and are not limited here.

[0135] Exemplarily, the intent recognition model includes a neural network linear layer for intent recognition and a sigmoid activation function. For intent recognition, pass the target sentence vector through the neural network linear layer for intent recognition to obtain an intent distribution vector. The neural network linear layer is a learnable network parameter. Assume that the intent distribution vector formula is as follows: I = W I C + b I Among them, I is the intent distribution vector; W I and b IThey are the learnable parameters of the neural network linear layer for intent recognition. Then, for the intent distribution vector, an intent prediction vector is obtained through the sigmoid activation function respectively. Suppose the formula of the intent prediction vector is as follows: I′ = sigmoid(I), where I′ is the intent prediction vector. During the training process, for the intent prediction vector and the final classification result, a loss function of binary cross entropy is used for training. During the prediction process, in a way of threshold screening, those higher than the threshold in the prediction vector are screened out to obtain the final intent classification result.

[0136] S505. Concatenate each intent result with the first text sequence respectively to obtain at least one second text sequence.

[0137] Exemplarily, each intent result is concatenated in front of the first text sequence respectively to obtain at least one second text sequence, and each second text sequence includes an intent result and the first text sequence.

[0138] S506. Encode each second text sequence through a second encoding model respectively to obtain at least one second word vector, and each second word vector corresponds to a second text sequence.

[0139] Among them, the second encoding model can be a pre-trained language model such as a BERT model, a RoBerta model, an Electra model, etc., which can be specifically determined according to the actual situation and is not limited here. In the embodiments of the present application, the second encoding model is taken as an example of the BERT model for illustration.

[0140] Among them, the second encoding model and the first encoding model can be the same or different, which can be specifically determined according to the actual situation and is not limited here.

[0141] Among them, a second word vector is the vector representation of each word in a second text sequence. A second word vector is obtained by encoding a second text sequence.

[0142] S507. Perform slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intent result.

[0143] Among them, the intent recognition model can be pre-trained in advance, which can be specifically determined according to the actual situation and is not limited here.

[0144] It can be understood that since each second text sequence is obtained by concatenating the intent result and the first text sequence, when performing slot recognition on the second word vector, at least one slot result corresponding to the intent result in the first text sequence can be recognized.

[0145] Exemplarily, the first slot recognition model may include a slot recognition neural network linear layer and a softmax function. For the above-mentioned first text sequence, intent recognition is performed to identify at least one intent result. Then, in the slot recognition stage, by respectively splicing each intent result and the first text sequence, at least one second text sequence (i.e., slot recognition sequence) is obtained. Each second text sequence is input into the BERT model to obtain at least one second word vector (i.e., slot word vector). Then, for each second word vector, through the slot recognition neural network linear layer, the slot distribution vector of each word is obtained. For the slot distribution vector of each word, the softmax function is used for normalization to obtain the slot prediction vector of each word. During the training process, for the slot prediction vector and the final classification result, the cross-entropy loss function is used for training. During the prediction process, the highest possible value is directly selected as the slot classification result of the word.

[0146] Example 1, the sentence to be recognized (Query) in the user query is "Play Predator and download the TV version of Wukong Remote Control". For this sentence to be recognized, intent result 1 - "MediaSearch" and the corresponding slot of intent result 1 - "Play: actionPlay; Predator: title" are obtained; intent result 2 - "TvAPP" and the corresponding slot of intent result 1 - "Download: actionDownload; Wukong Remote Control: app". Therefore, the intent results and slot results obtained by recognizing the Query are shown in Table 1 below.

[0147] Table 1

[0148]

[0149] In the embodiments of the present application, at least one intent result in the first text sequence corresponding to the sentence to be recognized is first recognized, and then each recognized intent result is respectively spliced with the first text sequence (to establish the connection between each intent result and its corresponding slot) to obtain at least one second text sequence, and then slot recognition is performed on each second text sequence to obtain at least one slot result corresponding to each intent result, so as to accurately recognize the intent and slot of the sentence to be recognized, and improve the semantic understanding accuracy in the intelligent dialogue system.

[0150] In some embodiments of the present application, in combination with Figure 5 , as Figure 6 shown, the above S501 can be specifically implemented through the following S501a to S501c.

[0151] S501a. Perform speech recognition on the sentence to be recognized to obtain a third text sequence in the sentence to be recognized, and the third text sequence includes multiple words.

[0152] S501b. Based on the explicit knowledge base, perform knowledge base matching on each of the multiple words respectively to obtain at least one classification knowledge.

[0153] Among them, each classification knowledge corresponds to one word.

[0154] Among them, the explicit knowledge base includes classification knowledge of a large number of words, similar to a word classification knowledge dictionary, and can also be understood as a kind of explicit guidance in text.

[0155] Among them, the at least one classification identifier may include the classification knowledge corresponding to each word in the multiple words, or may include the classification knowledge corresponding to some of the words in the multiple words, which can be specifically determined according to the actual situation and is not limited here.

[0156] S501c. Insert the at least one classification knowledge into the third text sequence respectively to obtain the first text sequence.

[0157] Among them, in the first text sequence, each classification identifier is located behind the corresponding word and is separated by a preset delimiter.

[0158] It can be understood that the first text sequence is composed of the third text sequence, at least one classification knowledge, and a preset delimiter.

[0159] Among them, the preset delimiter can be determined according to the actual usage requirements and is not limited here.

[0160] Example 2. Continuing from the above Example 1, the Query is "Play Predator and download the TV version of Wukong Remote Control". First, perform knowledge base matching on the text sequence (the third text sequence) in the input Query, query some possible classification knowledge from the explicit knowledge base for the third text sequence, and insert the queried classification knowledge into the third text sequence. When inserting the classification knowledge, use ":" as the delimiter between the word and the classification knowledge corresponding to the word, that is, use the colon ":" to indicate that the following is the content of the classification knowledge; use a comma "," to separate multiple classification knowledge corresponding to one word; use a semicolon ";" to separate two words. Therefore, the first text sequence obtained after inserting the corresponding at least one classification knowledge for the third text sequence is "Play: actionPlay; Predator: videoName, musicFeeble; and:; Download: actionDownload; Wukong Remote Control: app; TV version:;". Exemplarily, the second text sequence can be represented as W = {w1, w2,..., w n}, where n is the length of the second text sequence.

[0161] For example, in this process, for the entity "Liu XX", the common identities are actor, director, and singer. If it is already known that the domain of this sentence is songs, then in slot recognition, the possibility of correctly labeling Liu XX is greater.

[0162] In the embodiments of the present application, by performing knowledge base matching on each word in the third text sequence based on an explicit knowledge base, at least one classification knowledge is obtained, and then the at least one classification knowledge is inserted into the third text sequence to obtain the first text sequence. Thus, in the process of performing intent and slot recognition based on the first text sequence, since the at least one classification knowledge is included in the first text sequence, the accuracy and efficiency of intent recognition and slot recognition can be improved.

[0163] In some embodiments of the present application, in combination with Figure 6 , as Figure 7 shown, before the above S507, the intent slot recognition method provided by the embodiments of the present application may further include the following S508, and the above S507 may be specifically implemented by the following S507a and S507b.

[0164] S508: Mark the positions of each word in the third text sequence in the second text sequence to obtain the first position information.

[0165] It can be understood that the above S508 is to record the positions of each word in the third text sequence in the second text sequence. The first position information is used to indicate the positions of each word in the third text sequence in the second text sequence.

[0166] S507a: Perform slot recognition on each second word vector through the first slot recognition model to obtain the respective slot results corresponding to each intent result in the second text sequence.

[0167] S507b: Based on the first position information, screen out at least one slot result corresponding to each intent result in the third text sequence from the respective slot results corresponding to each intent result.

[0168] It can be understood that since the second word vector is obtained by encoding the second text sequence, the second text sequence includes the first text sequence and the intent structure, and the first text sequence includes the third text sequence and at least one classification knowledge. Therefore, in the above S507a, the respective slot results corresponding to each intent result include at least one slot result corresponding to each intent result in the third text sequence, and may also include classification knowledge (the classification knowledge is also used as a slot result). Furthermore, based on the first position information, at least one slot result corresponding to each intent result in the third text sequence can be screened out from the respective slot results corresponding to each intent result.

[0169] Example 3, following the above Examples 1 and 2, for slot recognition, this patent adopts the BIO annotation method. That is, for the word "play", the model should output "B-actionPlay" for the position of "bo", and "I-actionPlay" for the position of "fang". For positions of words without meaning, "O" is output for representation.

[0170] In the embodiments of the present application, this solution is based on an explicit knowledge base and introduces classification knowledge in the second text sequence. Therefore, before slot recognition of the second text sequence, it is necessary to first record the positions of each word in the third text sequence in the second text sequence to obtain the first position information, so that when slot results are output after slot recognition of the second text sequence, according to the first position information, only the slot results corresponding to each word in the third text sequence are output, and the slot results corresponding to the classification knowledge positions are not output. In this way, the accuracy of intent recognition and slot recognition can be improved.

[0171] In some embodiments of the present application, in combination with Figure 7 , as Figure 8 shown, after the above S503, the intent slot recognition method provided by the embodiments of the present application may further include the following S509.

[0172] S509: Perform domain recognition on the target sentence vector through a domain recognition model, and recognize at least one domain result, and each domain result corresponds to an intent result.

[0173] Among them, the domain recognition model can be pre-trained, which can be specifically determined according to the actual situation and is not limited here.

[0174] Exemplarily, the domain recognition model includes a neural network linear layer for domain recognition and a sigmoid activation function. For domain recognition, the target sentence vector is passed through the neural network linear layer for domain recognition to obtain a domain distribution vector. The neural network linear layer is a learnable network parameter. Assume the domain distribution vector formula is as follows: D = W D C + b D . Among them, D is the domain distribution vector; W D and b D are the learnable parameters of the neural network linear layer for domain recognition. Then, the sigmoid activation function is used for the domain distribution vector to obtain a domain prediction vector. Assume the domain prediction vector formula is as follows: D′ = sigmoid(D). Among them, D′ is the domain prediction vector. During the training process, for the prediction vector and the final classification result, a loss function of binary cross entropy is used for training. During the prediction process, a threshold screening method is adopted to screen out those higher than the threshold in the prediction vector to obtain the final domain classification result.

[0175] Example 4. Continuing from the above Example 1, after adding domain recognition, the domain results obtained by recognizing the Query are APP and MOVIE. Then the final domain results, intent results, and slot results are shown in Table 2 below.

[0176] Table 2

[0177]

[0178] In the embodiments of the present application, domain recognition is added, so that the domain information included in the statement to be recognized can be displayed to the user.

[0179] In some embodiments of the present application, in combination with Figure 8 , as Figure 9 shown, the above S505 can be specifically implemented by the following S505a.

[0180] S505a: Concatenate each of the intent results and the corresponding domain results with the first text sequence respectively to obtain the at least one second text sequence.

[0181] Example 5. Continuing from the above Example 4, assume that "Domain: MOVIE; Intent: MediaSearch" and "Domain: APP; Intent: TvAPP" are recognized in the domain and intent recognition stage. Then after traversing, two input text sequences corresponding to slot recognition will be obtained, as follows:

[0182] The input text sequence corresponding to the first slot recognition is: "Domain: MOVIE; Intent: MediaSearch; Play: actionPlay; Predator: videoName, musicFeeble; And:; Download: actionDownload; Wukong Remote Control: app; TV side:;".

[0183] The input text sequence corresponding to the first slot recognition is: "Domain: APP; Intent: TvAPP; Play: actionPlay; Predator: videoName, musicFeeble; And:; Download: actionDownload; Wukong Remote Control: app; TV side:;".

[0184] In the embodiments of the present application, after concatenating the intent results and the domain results corresponding to the intent results, and then concatenating them with the first text sequence to obtain the second text sequence, when performing slot recognition based on the second text sequence, due to the addition of the guidance of the domain results and intent results to slot recognition, the at least one slot result corresponding to the intent result can be recognized more accurately, which can improve the accuracy and efficiency of intent and slot recognition.

[0185] In some embodiments of the present application, regardless of whether one or multiple intent results are identified through intent recognition of the first text sequence, at least one slot result corresponding to each intent result is identified according to the steps of S504 to S507 above. This can simplify the entire intent slot recognition model and improve the training efficiency of the intent slot recognition model.

[0186] In some embodiments of the present application, in combination with Figure 9 , as Figure 10 shown, before the above S505, the intent slot recognition method provided by the embodiments of the present application may further include the following S510 and S511, and the above S505a may be specifically implemented through the following S505b.

[0187] S510. Determine whether the number of the at least one intent result is greater than 1.

[0188] S505b. When the number of the at least one intent result is greater than 1, splice each of the intent results and the corresponding domain results with the first text sequence respectively to obtain the at least one second text sequence.

[0189] The above S505 may be specifically: when the number of the at least one intent result is greater than 1, splice each of the intent results with the first text sequence respectively to obtain at least one second text sequence;

[0190] The method further includes:

[0191] S511. When the number of the at least one intent result is equal to 1, perform slot recognition on the first word vector through the second slot recognition model to obtain at least one slot result corresponding to each intent result.

[0192] Wherein, the second slot recognition model and the first slot recognition model may be the same or different, and may be specifically determined according to the actual situation, which is not limited herein.

[0193] In the embodiments of the present application, when multiple intent results are identified through intent recognition of the first text sequence, each intent is spliced with the first text sequence to obtain a second text sequence, and then slot recognition is performed based on the second text sequence; while when one intent result is identified through intent recognition of the first text sequence, slot recognition is directly performed on the first word vector to identify at least one slot corresponding to the one intent result, so that it is no longer necessary to splice the one intent with the first text sequence, and slot recognition is directly performed based on the first text sequence, which can improve the current intent recognition efficiency.

[0194] It should be noted that the intent slot recognition method provided in the embodiments of the present application can be applied to the recognition of single intents and the corresponding slots, and is also applicable to the recognition of multiple intents and the corresponding slots for each intent.

[0195] For the single intent semantic understanding module, the input is the sentence to be recognized, and the output is the domain, intent, and slot information of the sentence to be recognized. Table 3 is an example of single intent semantic understanding data, where Query is the user query; the domain, intent, and slot are the recognition results to be obtained.

[0196] Table 3

[0197] Query Play Predator Intent & Slot MediaSearch: {Play: actionPlay; Predator: videoName} Domain MOVIE

[0198] For the recognition of multiple intents and the order relationship between intents, the input is the sentence to be recognized, and the output is the domain of the sentence to be recognized, multiple intents and the corresponding slots thereunder, and the order between the multiple intents. The above Table 2 is an example of multiple intent semantic understanding data, where Query is the user query; the domain, intent, and slot are the recognition results to be obtained.

[0199] In some embodiments of the present application, in combination with Figure 10 , such as Figure 11 shown, after the above S507b, the intent slot recognition method provided in the embodiments of the present application may further include the following S512.

[0200] S512. Determine the arrangement order of each intent in the first text sequence based on the positional relationship of a specific slot in at least one slot result corresponding to each intent result in the first text sequence.

[0201] Among them, for the order recognition between multiple intents, a heuristic algorithm based on specific slots under the intent is adopted.

[0202] Among them, the specific slot is also called the specific feature slot, which refers to the slot unique to a certain intent. For example, the "refrigerator compartment temperature" slot under the "control refrigerator" intent. For example, the "train number" slot only appears under the "book train ticket" intent.

[0203] It can be understood that due to the limitations of intents and slots, there will be specific slot information under an intent. When it is parsed that the sentence contains these specific slots, the intent order between multiple intents can be determined according to the order in which the specific slots appear in the sentence.

[0204] Exemplarily, when it is parsed that the sentence contains these specific slots, the intent order between multiple intents can be determined according to the order in which the specific slots appear in the sentence. Continuing with the above Examples 1 to 5, if "Predator" is recognized as "videoName", according to the specific attribution, the "MediaSearch" intent can be ranked before the "TvAPP" intent.

[0205] In the embodiments of the present application, the order of identifying intents plays an important role in helping to determine the execution order of downstream services. For example, in the query "Help me search for movies by Liu xx and play them on the TV", the search operation must be executed first before the play operation can be performed.

[0206] It should be noted that the intent slot recognition method provided in the embodiments of the present application can be implemented by an intent slot recognition model. The intent slot recognition model includes multiple small models, and each small model implements different functions in the intent slot recognition method, such as Figure 12 As shown, the intent slot recognition model may include a text conversion model, a text processing model, an encoding model (such as a BERT model), a word vector to sentence vector model, a domain recognition model, an intent recognition model, and a slot recognition model. Among them, the text conversion model can implement the function of converting the query statement into a first text sequence. The text processing model can implement functions such as knowledge base matching, text splicing, or insertion based on the display knowledge base. The encoding model is used to encode the text sequence to obtain word vectors. The word vector to sentence vector model is used to convert the obtained word vectors into sentence vectors. The domain recognition model is used to recognize the domain result in the statement to be recognized. The intent recognition model is used to recognize the intent result in the statement to be recognized. The slot recognition model is used to recognize the slot result in the statement to be recognized. When training the intent slot recognition model, it can be trained based on a loss function for the entire intent slot recognition model, or it can be trained based on the loss functions corresponding to each small model for the entire intent slot recognition model. Specifically, it can be determined according to the actual situation and is not limited here.

[0207] Exemplarily, as Figure 13 shown, it is a schematic diagram of an intent slot recognition framework shown in the embodiments of the present application. The intent slot recognition model includes two stages, one is the domain intent recognition stage, and the other is the slot recognition stage. Among them, the domain intent recognition stage includes text processing, encoding processing, average pooling processing, domain recognition, and intent recognition. The slot recognition stage includes single intent slot recognition and multi-intent slot recognition. Among them, multi-intent recognition includes: text splicing, encoding processing, and slot recognition.

[0208] Figure 14 It is a structural block diagram of an intent slot recognition device shown in the embodiments of the present disclosure. As Figure 14As shown in the figure, the device includes: an acquisition module 1401, configured to acquire a first text sequence corresponding to a statement to be recognized; an encoding module 1402, configured to perform encoding processing on the first text sequence through a first encoding model to obtain a first word vector corresponding to the first text sequence; a pooling module 1403, configured to perform average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence; an intent recognition module 1404, configured to perform intent recognition on the target sentence vector through an intent recognition model to recognize at least one intent result; a splicing module 1405, configured to splice each intent result with the first text sequence respectively to obtain at least one second text sequence; the encoding module 1406 is further configured to perform encoding processing on each second text sequence through a second encoding model to obtain at least one second word vector, and each second word vector corresponds to a second text sequence; a slot recognition module 1407, configured to perform slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intent result.

[0209] In some embodiments of the present application, the acquisition module 1401 is configured to receive the statement to be recognized; perform speech recognition on the statement to be recognized to obtain a third text sequence in the statement to be recognized, where the third text sequence includes multiple words; based on an explicit knowledge base, perform knowledge base matching on each of the multiple words respectively to obtain at least one classification knowledge, and each classification knowledge corresponds to a word; insert the at least one classification knowledge into the third text sequence respectively to obtain a first text sequence, in the first text sequence, each classification identifier is located behind the corresponding word and is separated by a preset separator.

[0210] In some embodiments of the present application, the device further includes: a marking module, configured to mark the positions of each word in the third text sequence in the second text sequence before performing slot recognition on each second word vector through the first slot recognition model to obtain at least one slot result corresponding to each intent result, to obtain first position information; the slot recognition module 1407 is specifically configured to perform slot recognition on each second word vector through the first slot recognition model to obtain each slot result corresponding to each intent result in the second text sequence; based on the first position information, filter out at least one slot result corresponding to each intent result in the third text sequence from each slot result corresponding to each intent result.

[0211] In some embodiments of the present application, the device further includes: a domain recognition module, configured to perform domain recognition on the target sentence vector through a domain recognition model after performing average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence, to recognize at least one domain result, and each domain result corresponds to an intent result.

[0212] In some embodiments of the present application, the splicing module 1405 is specifically configured to splice each of the intent results and the corresponding domain results with the first text sequence respectively to obtain the at least one second text sequence.

[0213] In some embodiments of the present application, the device further includes: a determination module, configured to determine whether the number of the at least one intent result is greater than 1 before splicing each intent result with the first text sequence to obtain the at least one second text sequence; the splicing module 1405 is specifically configured to, when the number of the at least one intent result is greater than 1, splice each intent result with the first text sequence respectively to obtain the at least one second text sequence; the slot recognition module 1407 is further configured to, when the number of the at least one intent result is equal to 1, perform slot recognition on the first word vector through a second slot recognition model to obtain at least one slot result corresponding to each intent result.

[0214] In some embodiments of the present application, the determination module is configured to, after performing slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intent result, determine the arrangement order of the respective intents in the first text sequence based on the positional relationship of a specific slot in at least one slot result corresponding to each intent result in the first text sequence.

[0215] As Figure 15 shown, an embodiment of the present application further provides a computing device 1500, which may be the above-mentioned electronic device or server. The computing device 1500 includes: a processor 1501, a memory 1502, and a computer program stored on the memory 1502 and executable on the processor 1501. When the computer program is executed by the processor 1501, it implements each process executed by the above-mentioned intent slot recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0216] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process executed by the above-mentioned intent slot recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0217] Among them, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0218] The present invention provides a computer program product, comprising: when the computer program product runs on a computer, enabling the computer to implement the above-mentioned intention slot recognition method.

[0219] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

[0220] For the sake of convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A method for identifying intent slots, characterized in that, Including: Obtain a first text sequence corresponding to the statement to be recognized; Perform encoding processing on the first text sequence through a first encoding model to obtain a first word vector corresponding to the first text sequence; Perform average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence; Perform intent recognition on the target sentence vector through an intent recognition model to recognize at least one intent result; Concatenate each intent result with the first text sequence respectively to obtain at least one second text sequence; Perform encoding processing on each second text sequence through a second encoding model to obtain at least one second word vector, and each second word vector corresponds to a second text sequence; Perform slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intent result.

2. The method according to claim 1, characterized in that, The obtaining of the first text sequence corresponding to the statement to be recognized includes: Receive the statement to be recognized; Perform speech recognition on the statement to be recognized to obtain a third text sequence in the statement to be recognized, and the third text sequence includes multiple words; Based on an explicit knowledge base, perform knowledge base matching on each of the multiple words respectively to obtain at least one classification knowledge, and each classification knowledge corresponds to a word; Insert the at least one classification knowledge into the third text sequence respectively to obtain the first text sequence. In the first text sequence, each classification identifier is located behind the corresponding word and is separated by a preset separator.

3. The method according to claim 2, characterized in that, Before performing slot recognition on each second word vector through the first slot recognition model to obtain at least one slot result corresponding to each intent result, the method further includes: Mark the positions of each word in the third text sequence in the second text sequence to obtain first position information; The performing of slot recognition on each second word vector through the first slot recognition model to obtain at least one slot result corresponding to each intent result includes: Perform slot recognition on each second word vector through the first slot recognition model to obtain respective slot results corresponding to each intent result in the second text sequence; Based on the first position information, filter out at least one slot result corresponding to each intent result in the third text sequence from the respective slot results corresponding to each intent result.

4. The method according to claim 1, characterized in that, After performing average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence, the method further includes: Perform domain recognition on the target sentence vector through a domain recognition model to recognize at least one domain result, and each domain result corresponds to an intent result.

5. The method according to claim 4, characterized in that, The concatenating of each intent result with the first text sequence respectively to obtain at least one second text sequence includes: Concatenate each intent result and the corresponding domain result with the first text sequence respectively to obtain the at least one second text sequence.

6. The method according to claim 1, characterized in that, Before the concatenating of each intent result with the first text sequence respectively to obtain at least one second text sequence, the method further includes: Determine whether the number of the at least one intended result is greater than 1; The splicing each intended result with the first text sequence respectively to obtain at least one second text sequence includes: When the number of the at least one intended result is greater than 1, splicing each intended result with the first text sequence respectively to obtain at least one second text sequence; The method further includes: When the number of the at least one intended result is equal to 1, performing slot recognition on the first word vector through a second slot recognition model to obtain at least one slot result corresponding to each intended result.

7. The method according to any one of claims 1 to 6, characterized in that, After performing slot recognition on each second word vector through the first slot recognition model to obtain at least one slot result corresponding to each intended result, the method further includes: Determining the arrangement order of each intention in the first text sequence based on the positional relationship of a specific slot in the at least one slot result corresponding to each intended result in the first text sequence.

8. An apparatus for identifying intent slots, characterized in that, Including: An acquisition module, configured to acquire a first text sequence corresponding to a to-be-recognized statement; An encoding module, configured to perform encoding processing on the first text sequence through a first encoding model to obtain a first word vector corresponding to the first text sequence; A pooling module, configured to perform average pooling processing on the first word vector to obtain a target sentence vector corresponding to the first text sequence; An intention recognition module, configured to perform intention recognition on the target sentence vector through an intention recognition model to recognize at least one intended result; A splicing module, configured to splice each intended result with the first text sequence respectively to obtain at least one second text sequence; The encoding module is further configured to perform encoding processing on each second text sequence respectively through a second encoding model to obtain at least one second word vector, and each second word vector corresponds to a second text sequence; A slot recognition module, configured to perform slot recognition on each second word vector through a first slot recognition model to obtain at least one slot result corresponding to each intended result.

9. The device according to claim 8, wherein, The acquisition module is configured to receive the to-be-recognized statement; Performing speech recognition on the to-be-recognized statement to obtain a third text sequence in the to-be-recognized statement, where the third text sequence includes a plurality of words; Based on an explicit knowledge base, performing knowledge base matching on each of the plurality of words respectively to obtain at least one classification knowledge, and each classification knowledge corresponds to a word; Inserting the at least one classification knowledge into the third text sequence respectively to obtain the first text sequence, in the first text sequence, each classification identifier is located behind the corresponding word and is separated by a preset separator.

10. The device according to claim 9, wherein, The device further includes: A marking module, configured to mark the positions of each word in the third text sequence in the second text sequence before performing slot recognition on each second word vector through the first slot recognition model to obtain at least one slot result corresponding to each intended result, so as to obtain first position information; The slot recognition module is specifically configured to perform slot recognition on each second word vector through the first slot recognition model to obtain respective slot results corresponding to each intention result in the second text sequence; Based on the first position information, at least one slot result corresponding to each intention result in the third text sequence is filtered out from the respective slot results corresponding to each intention result.

Citation Information

Patent Citations

  • Natural language processing method, natural language processing device and intelligent question-answering system

    CN111026842A

  • Slot prediction method and device for multi-intention text and computer equipment

    CN113887237A