Natural language recognition method, device and equipment in task-based dialogue system

By implementing natural language recognition methods in a task-based dialogue system, and automatically matching and labeling natural language data, the problem of high corpus labeling costs when developing new skills is solved and development efficiency is improved.

CN114357125BActive Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011086835.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-12
Publication Date
2025-06-06
Estimated Expiration
2040-10-12

AI Technical Summary

Technical Problem

When developing new skills in task-based dialogue systems, a large amount of corpus data needs to be written and marked, resulting in high labor and time costs and affecting development efficiency.

Method used

By implementing a natural language recognition method in a task-based dialogue system, matching the semantic information of the natural language to be recognized with the skill description information in the task-based dialogue system, determining the target skills and intentions, and extracting the slot information through the slot extraction model, thereby automatically completing the labeling of natural language and the acquisition of sample data.

Benefits of technology

This method greatly improves the labeling efficiency of natural language and the development efficiency of task-based dialogue systems, reduces labor and time costs, and enables developers to develop new skills more quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114357125B_ABST
    Figure CN114357125B_ABST
Patent Text Reader

Abstract

The present application provides a natural language recognition method, device and equipment in a task-based dialogue system, which relates to the field of computer technology, and in particular to the field of artificial intelligence. It includes: determining the target skill that matches the semantics of the natural language to be recognized according to the matching conditions between the semantic information of the natural language to be recognized and the skill description information of each dialogue-type task in the task-based dialogue system; determining the target intent according to the matching conditions between the semantic information of the natural language to be recognized and the intention description information of each candidate intention corresponding to the target skill; extracting the slot information of each candidate slot from the natural language to be recognized according to the slot description information of each candidate slot corresponding to the target intention and the slot information extraction condition; obtaining the dialogue-type task recognition result according to the skill description information of the target skill, the intention description information of the target intention, the slot description information of each candidate slot and its slot information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device and equipment for natural language recognition in a task-based dialogue system. Background Art

[0002] The open dialogue platform includes multiple task-based dialogue systems with different functions. Task-based dialogue systems are dialogue systems that can help users solve specific tasks. Each task-based dialogue system corresponds to a related skill, such as skills that help users query and book related air tickets in the field of air tickets. The task-based dialogue system uses a natural language understanding model to correctly understand various possible user inputs related to skills. In related technologies, the open dialogue platform allows developers to independently develop new skills. The number of skills supported by the open dialogue platform greatly affects the user experience. Therefore, enabling developers to develop new skills more quickly and easily has become the core capability of the open dialogue platform.

[0003] When developing new skills, in order to implement a robust natural language understanding model, developers need to write a large amount of skill-related corpus and annotate the written corpus to provide sample data for the natural language understanding model. Developers usually need to write and annotate thousands of corpora to train a good natural language understanding model to achieve recognition of conversational tasks. This requires developers to pay huge manpower and time costs, affecting development efficiency. Therefore, how to improve the efficiency of natural language annotation and conversational task recognition is an issue that needs to be considered. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus and device for natural language recognition in a task-based dialogue system, which are used to improve the efficiency of natural language annotation and the development efficiency of the task-based dialogue system.

[0005] In a first aspect, the present application provides a natural language recognition method in a task-based dialogue system, comprising:

[0006] According to a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system, a target skill matching the semantics of the natural language to be recognized is determined from each dialog task;

[0007] Determine a target intent that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intent description information of each candidate intent corresponding to the target skill, wherein each intent description information is used to describe a subtask included in the conversational task;

[0008] According to the slot description information of each candidate slot corresponding to the target intent and the slot information extraction condition, the slot information of each candidate slot is extracted from the natural language to be recognized, and the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent;

[0009] According to the skill description information of the target skill, the intent description information of the target intent, and the slot description information and slot information of each candidate slot, a dialog task recognition result of the natural language to be recognized is obtained.

[0010] In a possible implementation, the method further includes:

[0011] Annotate natural language identified as untrained conversational tasks;

[0012] Obtain sample data for untrained conversational tasks;

[0013] Based on the obtained sample data, the skill classification model, the intent recognition model and the slot extraction model are updated and trained; and

[0014] The untrained conversational task is updated to a trained conversational task.

[0015] In a second aspect of the present application, a natural language annotation method in a task-based dialogue system is provided, comprising:

[0016] According to a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system, determining a target skill that matches the semantics of the natural language to be recognized from each dialog task;

[0017] Determine a target intent that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intent description information of each candidate intent corresponding to the target skill, wherein each intent description information is used to describe a subtask included in the conversational task;

[0018] Extracting slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent;

[0019] According to the skill description information of the target skill, the intent description information of the target intent, and the slot description information and slot information of each candidate slot, the natural language to be recognized is labeled to obtain sample data of the conversational task.

[0020] In a third aspect of the present application, a natural language recognition device in a task-based dialogue system is provided, comprising:

[0021] a first skill determination unit, configured to determine, from each dialog-type task, a target skill that matches the semantics of the natural language to be recognized, based on a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog-type task in the task-based dialog system;

[0022] A first intention determination unit, configured to determine a target intention that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intention description information of each candidate intention corresponding to the target skill, wherein each intention description information is used to describe a subtask included in the dialog task;

[0023] A first slot determination unit is used to extract slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent;

[0024] The recognition unit is used to obtain the dialogue task recognition result of the natural language to be recognized based on the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot.

[0025] In a possible implementation, the target skill is obtained based on a trained skill classification model, the target intent is obtained based on a trained intent recognition model, and the slot information is obtained based on a trained slot extraction model, wherein:

[0026] The dialog tasks include trained dialog tasks and untrained dialog tasks, and the skill classification model, intent recognition model and slot extraction model are trained using sample data of each trained dialog task;

[0027] The untrained dialogue task includes: skill description information, intent description information of each alternative intent, and alternative slot description information of each alternative intent.

[0028] In a possible implementation, the skill classification model includes a first vector representation module and a first classification module, the skill classification corpus sample for training the skill classification model includes at least one sample corresponding to each trained dialog task, each skill classification corpus sample is annotated with a similarity probability value of the corresponding dialog task, and the trained skill classification model is trained in the following manner:

[0029] Obtaining first target vector representations of skill description information of each trained task based on the first vector representation module;

[0030] For each skill classification corpus sample, obtaining a first reference vector representation of the skill classification corpus sample based on a first vector representation module;

[0031] Obtaining first similarities between the first reference vector representation and each first target vector representation based on the first classification module;

[0032] The parameters of the first vector representation module and the first classification module are adjusted based on the obtained first similarities until the first similarities meet the set conditions.

[0033] In a possible implementation manner, the skill determination unit is specifically configured to:

[0034] Based on the first vector representation module, first comparison vector representations of skill description information of each trained conversational task and untrained conversational task are obtained respectively;

[0035] Obtaining a first to-be-recognized vector representation of the to-be-recognized natural language material based on the first vector representation module;

[0036] Based on the first classification module, second similarities between the first comparison vector representation and each first vector representation to be identified are respectively obtained, and a classification result is obtained based on the second similarities.

[0037] In a possible implementation, the intent recognition model includes a second vector representation module and a second classification module, the intent recognition corpus sample for training the intent recognition model includes at least one sample corresponding to each trained dialogue-type task, each intent recognition corpus sample is annotated with a similarity probability value of the candidate intent to which it belongs, and the trained intent recognition model is trained in the following manner:

[0038] Based on the second vector representation module, second target vector representations of the intent description information corresponding to each subtask in each trained task are respectively obtained;

[0039] For each intent recognition corpus sample, obtaining a second reference vector representation of the intent recognition corpus sample based on the second vector representation module;

[0040] Obtaining third similarities between the second reference vector representation and each second target vector representation based on the second classification module;

[0041] The parameters of the second vector representation module and the second classification module are adjusted based on the obtained third similarities until the third similarities meet the set conditions.

[0042] In a possible implementation manner, the intention determination unit is specifically configured to:

[0043] Based on the second vector representation module, second comparison vector representations of the intention description information of each subtask in the trained conversational task and each subtask in the untrained conversational task are respectively obtained;

[0044] Obtaining a second to-be-recognized vector representation of the to-be-recognized natural language material based on the second vector representation module;

[0045] Based on the second classification module, fourth similarities between the second comparison vector representation and each second vector representation to be identified are respectively obtained, and a classification result is obtained based on the fourth similarities.

[0046] In a possible implementation, the slot extraction model includes a third vector representation module and a slot extraction module, the slot extraction corpus sample for training the slot extraction model includes at least one sample corresponding to each trained dialogue task, each slot extraction corpus sample is annotated with reference slot information of the candidate slot to which it belongs, and the trained slot extraction model is trained in the following manner:

[0047] For each slot extraction corpus sample, obtaining a third reference vector representation of the slot extraction corpus sample based on a third vector representation module;

[0048] Based on the slot extraction module, first target slot information of slot description information corresponding to each candidate slot in the third reference vector representation is respectively obtained;

[0049] The parameters of the third vector representation module and the slot extraction module are adjusted based on the obtained first target slot information until the obtained first target slot information meets the set conditions.

[0050] In a possible implementation manner, the slot determination unit is specifically configured to:

[0051] Obtaining a third to-be-recognized vector representation of the to-be-recognized natural language material based on the third vector representation module;

[0052] Based on the slot extraction module, second target slot information represented by the third vector to be identified is obtained respectively, and a classification result is obtained based on the second target slot information.

[0053] In a possible implementation, the method further includes a first marking unit, which is specifically configured to:

[0054] Annotate natural language identified as untrained conversational tasks;

[0055] Obtain sample data for untrained conversational tasks;

[0056] Based on the obtained sample data, the skill classification model, the intent recognition model and the slot extraction model are updated and trained; and

[0057] The untrained conversational task is updated to a trained conversational task.

[0058] In a fourth aspect of the present application, a natural language annotation device in a task-based dialogue system is provided, comprising:

[0059] a second skill determination unit, configured to determine, from each dialog task, a target skill that matches the semantics of the natural language to be recognized, based on a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system;

[0060] A second intention determination unit, configured to determine a target intention that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intention description information of each candidate intention corresponding to the target skill, wherein each intention description information is used to describe a subtask included in the dialog task;

[0061] A second slot determination unit is used to extract slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent;

[0062] The second labeling unit is used to label the natural language to be recognized according to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot to obtain sample data of the dialogue task.

[0063] In a possible implementation, the target skill is obtained based on a trained skill classification model, the target intent is obtained based on a trained intent recognition model, and the slot information is obtained based on a trained slot extraction model, wherein:

[0064] The conversational tasks include trained conversational tasks and untrained conversational tasks. The skill classification model, intent recognition model and slot extraction model are trained using sample data of each trained conversational task.

[0065] In a possible implementation, a training unit is further included, which is specifically used to:

[0066] Based on the obtained sample data, the skill classification model, intent recognition model and slot extraction model are further updated and trained, wherein when the obtained sample data includes natural language identified as an untrained conversational task, the untrained conversational task is updated to a trained conversational task after the updated training.

[0067] In a fifth aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method of the first aspect is implemented when the processor executes the program.

[0068] In a sixth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions. When the computer instructions are executed on a computer, the computer executes the method of the first aspect.

[0069] Since the embodiment of the present application adopts the above technical solution, it has at least the following technical effects:

[0070] The present application directly matches the semantic information of the natural language to be recognized with the skill description information of the task-based dialogue system, so that the skill to which the natural language to be recognized belongs can be determined among the multiple skills corresponding to the task-based dialogue system, and the semantic information of the recognized natural language can be directly matched with the intent description information corresponding to the skill to which it belongs, so that the intent to which the natural language to be recognized belongs can be determined, and the slot information in the natural language to be recognized can be determined through the slot description information of the alternative slots and the slot extraction conditions, so as to recognize the natural language to be recognized. When developing new skills, the skill description information of the new skills can be used to extract the slot information. The natural language matching the new skill is obtained from the natural language library, and the natural language is automatically labeled to obtain sample data for the new skill. There is no need to manually write and label a large amount of sample data. You only need to define the skill description information, intent description information and alternative slot description information of the new skill you want to develop, so that the task-based dialogue system can obtain sample data for the new skill. The obtained sample data can then be used to train the conversational task system, thereby simplifying the process of developing new skills, saving time and labor costs, and allowing developers to develop new skills more quickly and conveniently. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0072] Figure 1 This is an example diagram of skill-related information, intention-related information, and slot-related information in the field of air tickets in an embodiment of the present application;

[0073] Figure 2This is an example diagram of annotating natural language input by a user in an embodiment of the present application;

[0074] Figure 3 A schematic diagram of an application scenario applicable to an embodiment of the present application;

[0075] Figure 4 This is a flowchart of a natural language recognition method in a task-based dialogue system in an embodiment of the present application;

[0076] Figure 5 This is an example diagram of multiple natural language understanding models in the related technology of the embodiments of the present application;

[0077] Figure 6 This is an example diagram of a general natural language understanding model in an embodiment of the present application;

[0078] Figure 7 This is a structural example diagram of a general natural language understanding model in an embodiment of the present application;

[0079] Figure 8 This is a structural example diagram of a skill classification model in an embodiment of the present application;

[0080] Fig. 9 This is a structural example diagram of a BERT-based classification model in an embodiment of the present application;

[0081] Fig.10 This is a structural example diagram of an intent recognition model in an embodiment of the present application;

[0082] Fig.11 This is a structural example diagram of a slot extraction model in an embodiment of the present application;

[0083] Fig.12 This is a structural example diagram of a BERT-based QA model in an embodiment of the present application;

[0084] Fig.13 A sample data example diagram of skill description information and some examples of a weather conversation task provided in an embodiment of the present application;

[0085] Fig.14 An example diagram of sample data of skill description information and some examples of a flight dialogue task provided in an embodiment of the present application;

[0086] Fig.15 A diagram showing skill description information for a hotel conversational task and an example of natural language to be recognized provided in an embodiment of the present application;

[0087] Fig.16 An example diagram of a natural language recognition result for a hotel conversation task provided in an embodiment of the present application;

[0088] Fig.17 This is a flowchart of a natural language annotation method in a task-based dialogue system in an embodiment of the present application;

[0089] Fig.18 Schematic diagram of the structure of a natural language recognition device in a task-based dialogue system in an embodiment of the present application;

[0090] Fig.19 Schematic diagram of the structure of a natural language annotation device in a task-based dialogue system in an embodiment of the present application;

[0091] Fig. 20 It is a schematic diagram of the structure of the computing device in the embodiment of the present application. DETAILED DESCRIPTION

[0092] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of the application for protection. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.

[0093] The terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in the present application can mean at least two, for example, two, three or more, and the embodiments of the present application are not limited.

[0094] The embodiments of the present application relate to artificial intelligence (AI) and machine learning technology, and are designed based on natural language processing technology and machine learning (ML) in artificial intelligence.

[0095] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology mainly includes computer vision technology, natural language processing technology, and machine learning / deep learning.

[0096] With the research and advancement of artificial intelligence technology, artificial intelligence has been studied and applied in many fields, such as common smart homes, smart customer service, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, robots, smart medical care, etc. I believe that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.

[0097] Machine learning is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning and other technologies. In the process of semantic feature extraction of video titles, the embodiment of the present application adopts a semantic feature extraction model based on machine learning or deep learning to learn video title samples with category labels, so that the feature vector of the semantic features of the input video title can be extracted.

[0098] Natural language processing technology is an important direction in the fields of computer science and artificial intelligence. Its research can realize various theories and methods for effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language used by people in daily life, so it is closely related to the study of linguistics. Natural language processing technology generally includes technologies such as text generation, text processing, semantic understanding, machine translation, robot question and answer, knowledge graph, etc. The embodiment of the present application adopts the semantic understanding technology in natural language processing technology to perform semantic understanding on the video titles of each video, and based on the obtained feature vector that can characterize the semantic features of the video title, cluster processing is performed on each video in the video data set to obtain multiple video sets.

[0099] In order to help those skilled in the art better understand the technical solution of the present application, the technical terms involved in the present application are explained below.

[0100] 1) Task-based dialogue system: refers to a dialogue system that helps users solve specific tasks and provides information or services under specific conditions. Usually, it is to meet the needs of users with clear purposes, such as checking traffic, checking phone bills, ordering meals, booking tickets, consulting and other task-based scenarios, which is different from chat-based dialogue systems. Because user needs are more complex, they usually require multiple rounds of interaction. Users may also continuously modify and improve their needs during the dialogue process. Task-based dialogue systems need to help users clarify their purposes through questioning, clarification and confirmation. The core modules of task-based dialogue systems mainly include natural language understanding module (Natural Language Understanding), dialogue management module (Dialog Management) and natural language generation module (Natural Language Generation).

[0101] 2) Query: The natural language input by the user to the task-based dialogue system. For example, if the user's need is to book a flight ticket, then a request such as "I want to book a high-speed train to Beijing tonight" can be input to the task-based dialogue system.

[0102] 3) Natural Language Understanding (NLU), domain, intent, and slot;

[0103] Natural language understanding: natural language understanding technology in task-based dialogue systems. When the natural language input by the user passes through the natural language understanding module, it needs to go through three sub-modules: domain identification, user intent identification, and slot extraction;

[0104] Domain: The domain in which the user needs to complete a specific task, such as the air ticket domain, restaurant domain, music domain, etc. In a task-based dialogue system, a domain is usually defined by a set of intents and slots.

[0105] Intent: In a specific field, there are usually some segmented user intents. For example, in the air ticket field, there may be intent to search for air tickets, intent to book air tickets, intent to cancel air tickets, etc.

[0106] Slot: represents the key information that needs to be collected to complete a task in a specific field. For example, in the field of air tickets, there are slots such as departure city, destination city, departure date, etc. Each slot needs to identify the corresponding key information, such as departure city: Beijing, destination city: Shenzhen, departure date: *month*day, etc.

[0107] 4) Applications, or application programs, are computer programs that can complete one or more services. They generally have a visual display interface and can interact with users, such as electronic maps and calendars. Some applications require users to install them on the terminal device they use before they can be used, while others do not require application installation, such as various small programs and web pages in some social applications. Small programs do not need to be downloaded and installed to be used. Users can open the application by scanning or searching.

[0108] 5) Client, or user end, refers to a program that corresponds to a server and provides local services to clients. Except for some applications that only run locally, they are generally installed on ordinary client computers and need to work in conjunction with the server. After the development of the Internet, the more commonly used user ends include web browsers used for the World Wide Web, email clients for sending and receiving emails, and client software for instant messaging. For this type of application, there needs to be a corresponding server and service program in the network to provide the corresponding services, such as database services, email services, etc. In this way, a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application.

[0109] The design concept of this application is described below.

[0110] In related technologies, different task-based dialogue systems are developed separately. It usually takes three main steps to configure the natural language understanding capability of a new skill in an open dialogue platform:

[0111] 1) Define the intent and slots related to a skill: A skill usually contains one or more intents. After defining the intent and slots of a skill, the natural language understanding module corresponding to the skill needs to identify the intent corresponding to the skill in the natural language input by the user and extract the slot information in the natural language input by the user. It should be noted that in the relevant technology, the natural language understanding model and the skill are in a one-to-one correspondence, such as Figure 1 Shown are the skill-related information, intention-related information and slot-related information in the field of air tickets proposed in the embodiment of the present application.

[0112] 2) Write and annotate skill-related corpus: The natural language understanding module of the current open dialogue platform is usually based on machine learning methods. In order to train a natural language understanding model with better performance, a large amount of annotated data must be provided for the dialogue tasks that need to be developed. Therefore, developers must first write the user's natural language that may trigger the skill, and then annotate the intent and involved slots in the user's natural language. Due to the diversity and richness of natural language, in order to train a robust natural language understanding model, developers are usually required to write a large amount of rich and diverse natural language data and annotate this data, such as Figure 2 Shown is a schematic diagram for labeling natural language input by a user.

[0113] 3) Use the labeled data for model training. After obtaining the labeled data, you can start training the natural language understanding model. Intent recognition is a classification task, which can be, but not limited to, using classification models such as TextCNN and LSTM (Long Short-Term Memory); slot extraction is a sequence labeling task, which can be, but not limited to, using a combination of LSTM model and CRF (Conditional Random Field) model.

[0114] In the above three steps, the preparation and labeling of sample data for each conversational task usually consumes a lot of manpower and time costs, creating obstacles to the development of new skills.

[0115] In order to develop skills corresponding to task-based dialogue systems more quickly and conveniently, the embodiment of the present application takes into account the similarity of related dialogue tasks and proposes a general task-based dialogue system. The sample data of the general task-based dialogue system includes sample data used when each dialogue task is trained respectively. The trained general task-based dialogue system can learn the common features of each dialogue task. Then, when used, the semantic information of the natural language to be recognized is directly matched with the description information of each dialogue task, and the recognition result of the natural language to be recognized is determined according to the similarity. In this way, a set of systems can be used to realize the functions of multiple task-based dialogue systems. When a new task-based dialogue system needs to be developed, the description information of the new task-based dialogue system to be developed and the natural language to be recognized can be similarly matched, and the matched natural language can be used as candidate corpus for training the new task-based dialogue system. After the obtained candidate corpus is sorted and labeled, it can be used as sample data for training the new task-based dialogue system. Therefore, the general task-based dialogue system provided in the embodiment of the present application can not only identify dialogue tasks for the natural language to be recognized, but also find sample data for the training of new task-based dialogue systems, thereby greatly improving the development efficiency of task-based dialogue systems.

[0116] Specifically, the semantic information of the natural language to be recognized is matched with the skill description information of different dialogue tasks, so that the skill to which the natural language to be recognized belongs can be determined among the multiple skills corresponding to the task-based dialogue system, and by directly matching the semantic information of the recognized natural language with the intent description information corresponding to the skill to which it belongs, the intent to which the natural language to be recognized belongs can be determined, and the slot information in the natural language to be recognized is determined through the slot description information and slot extraction conditions of the alternative slots, and then the natural language to be recognized is recognized. Compared with the task-based dialogue system in the related art in which each natural language understanding model corresponds to a skill, the task-based dialogue system of the present application corresponds to a natural language understanding model, which corresponds to multiple different skills. Therefore, when developing new skills, there is no need to write and annotate a large amount of sample data. It is only necessary to define the skill description information, intent description information and alternative slot description information of the new skill to be developed, so that the task-based dialogue system can add a new skill, which simplifies the process of developing new skills, saves time and labor costs, and enables developers to develop new skills more quickly and conveniently.

[0117] It should be noted that the task-based dialogue system of the present application can also identify the skills, intentions and slot information corresponding to the natural language input by the user, and mark the identified skills, intentions and slot information on the natural language input by the user, and use the marked natural language as sample data for the corresponding skills, or pre-mark the natural language through the task-based dialogue system, and then manually proofread or mark it, so as to further reduce the manpower cost required for marking the data.

[0118] It is particularly important to note that when a new conversational task needs to be developed, the natural language identified as an untrained conversational task can be labeled to obtain sample data of the untrained conversational task, and the skill classification model, intent recognition model, and slot extraction model can be updated and trained based on the obtained sample data; and the untrained conversational task can be updated to a trained conversational task, wherein the natural language identified as an untrained conversational task can be labeled, but is not limited to, after the natural language of the untrained conversational task is identified, manually labeled, or directly labeled after the target skill, target intent, and slot information of the natural language of the untrained conversational task are determined through the skill classification model, intent recognition model, and slot extraction model, or manually corrected and labeled; updating the skill classification model, intent recognition model, and slot extraction model with the obtained sample data can improve the recognition accuracy of the skill classification model, intent recognition model, and slot extraction model.

[0119] In order to better understand the technical solutions provided in the embodiments of the present application, the following briefly introduces the application scenarios to which the technical solutions provided in the embodiments of the present application are applicable. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited thereto. In specific implementation, the technical solutions provided in the embodiments of the present application can be flexibly applied according to actual needs.

[0120] refer to Figure 3 , which is a schematic diagram of an application scenario of a natural language recognition method in a task-based dialogue system provided in an embodiment of the present application. The application scenario includes multiple terminal devices 301 and a server 302. The terminal device 301 and the server 302 are connected via a wireless or wired communication network, and the terminal device 301 includes but is not limited to a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a smart TV, a vehicle-mounted device, a personal digital assistant (PDA), and other electronic devices. The server 302 can be a single server, a server cluster consisting of several servers, or a cloud computing center.

[0121] A client is installed in the terminal device 301, and the client can be provided with natural language recognition services, dialogue management system services, and natural language generation services in the task-based dialogue system by the server 302, or the server 302 provides annotation services in natural language. The client in the terminal device 301 includes a general task-based dialogue system, which can recognize the natural languages ​​corresponding to multiple skills. It should be noted that the skill information, domain information and slot information in the natural language obtained in the natural language recognition service are used as input to the dialogue management system. The dialogue management system consists of two parts: state tracking and dialogue strategy. The state tracking module includes various information of the ongoing dialogue. It updates the current dialogue state based on the old state, the annotation information of the natural language recognition service and the system state (i.e., through the query with the database). For example, for the natural language "I want to book a flight to Beijing", the state tracking module can directly query the flight tickets to Beijing and obtain relevant information. The dialogue strategy is closely related to the task scenario and is usually used as the output of the dialogue management module, such as the counter-question strategy for the missing slots in the scenario. For example, in the above natural language, in the related services for booking flights, there is a lack of slot information for travel time. In this case, the task-based dialogue system can consult or remind the user.

[0122] In an embodiment of the present application, after receiving the natural language of the user, the client of the terminal device 301 determines the skills, intentions and slot information of the natural language input by the user through the natural language recognition service, and then provides corresponding services to the user. As an optional implementation method, after receiving the natural language of the user, the client of the terminal device 301 marks the natural language through the identified skills, intentions and slot information to obtain new sample data. Of course, the client of the terminal device 301 can also provide relevant services to the user after receiving the natural language of the user, and mark the identified skills, intentions and slot information as sample data.

[0123] Of course, the method provided in the embodiment of the present application is not limited to Figure 3 The application scenarios shown can also be used in other possible application scenarios, and the embodiments of the present application are not limited thereto. Figure 3 The functions that can be implemented by each device in the application scenario shown will be described in the subsequent method embodiments, and will not be described in detail here.

[0124] To further illustrate the technical solution provided by the embodiment of the present application, this is described in detail below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiment of the present application provides the method operation steps shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or no creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiment of the present application. The method may be executed in the order of the method shown in the embodiment or drawings or in parallel during the actual processing process or when the device is executed.

[0125] The present application provides a natural language recognition method in a task-based dialogue system, which can be executed by a task-based dialogue server, for example Figure 3 The natural language recognition method in the task-based dialogue system provided in the embodiment of the present application is as follows Figure 4 As shown, Figure 4 The flowchart shown is described as follows.

[0126] Step S401, determining a target skill that matches the semantics of the natural language to be recognized from each dialog task according to a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system;

[0127] The natural language to be recognized in the embodiment of the present application is the natural language input by the user, that is, the specific task that the user needs to complete. Compared with the related technology, the current task-based dialogue system only supports pre-trained dialogue tasks. In the embodiment of the present application, as long as the developer defines the skill description information of the dialogue task, the intent description information of each alternative intent, and the alternative slot description information of each alternative intent in advance, relevant services can be provided to the user without the need for one-to-one training in advance.

[0128] like Figure 5 As shown, each domain in the task-based dialogue system in the related art corresponds to a natural language understanding model, and the domain of each natural language understanding model has been determined. Therefore, the natural language understanding model corresponding to each domain needs to be trained with sample data labeled in the domain, which greatly increases the manpower and time costs of writing and labeling sample data. The task-based dialogue system of the embodiment of the present application is as follows Figure 6 As shown, the task-based dialogue system of the embodiment of the present application corresponds to a universal natural language understanding model, and the universal natural language understanding model corresponds to multiple different skills. It should be noted that each dialogue task corresponds to a corresponding skill. In the embodiment of the present application, the skill description information of the dialogue task is defined in advance, and the semantic information of the natural language to be identified is matched with the skill description information of each dialogue task in the task-based dialogue system. According to the matching conditions between the semantic information of the natural language to be identified and the skill description information of each dialogue task in the task-based dialogue system, the target skill matching the semantics of the natural language to be identified is determined from each dialogue task, wherein the matching condition may be, but is not limited to, determining the similarity between the semantic information of the natural language to be identified and the skill description information of each dialogue task in the task-based dialogue system, determining the dialogue task whose similarity reaches the target threshold as the target skill matching the natural language to be identified, or determining the dialogue task with the greatest similarity as the target skill matching the natural language to be identified.

[0129] It should be noted that, in the embodiment of the present application, the semantic information of the natural language to be recognized can be but is not limited to the vector representation of the natural language to be recognized. Optionally, the vector representation of the skill description information of each conversational task in the task-based dialogue system can also be determined. By comparing the similarity between the vector representation of the natural language to be recognized and the vector representation of the skill description information of each conversational task in the task-based dialogue system, the target skill that matches the semantics of the natural language to be recognized can be determined from each conversational task.

[0130] The fields in which the specific tasks that users need to complete may include air tickets, hotels, etc., but it should be noted that due to different developers, the same field may include multiple similar skills, but the services provided by multiple similar skills in the same field may also be the same, or, multiple skills in the same field have only minor differences in intent and slot extraction. Therefore, in the implementation of this application, each conversational task directly corresponds to a skill, and the corresponding skill can be selected for the user based on the user's historical operation data.

[0131] Step S402, determining a target intent that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intent description information of each candidate intent corresponding to the target skill, wherein each intent description information is used to describe a subtask included in the conversational task;

[0132] In the embodiment of the present application, each target skill corresponds to multiple alternative intentions. After determining the target skill corresponding to the natural language to be identified, the target intention that matches the natural language to be identified is determined from the multiple alternative intentions corresponding to the target skill. In the embodiment of the present application, the target intention that matches the semantics of the natural language to be identified is determined based on the matching condition between the semantic information of the natural language to be identified and the intention description information of each alternative intention corresponding to the target skill, wherein the matching condition may be, but is not limited to, determining the similarity between the semantic information of the natural language to be identified and the intention description information of each alternative intention corresponding to the target skill, determining the alternative intention whose similarity reaches the target threshold as the target intention that matches the natural language to be identified, or determining the alternative intention with the greatest similarity as the target intention that matches the natural language to be identified.

[0133] It should be noted that in the embodiment of the present application, the semantic information of the natural language to be identified can be but is not limited to the vector representation of the natural language to be identified. Optionally, the vector representation of the intention description information of each alternative intention corresponding to the target skill can also be determined. By comparing the vector representation of the natural language to be identified and the similarity between the vector representation of the intention description information of each alternative intention corresponding to the target skill, the target intention that matches the semantics of the natural language to be identified is determined. In the embodiment of the present application, multiple alternative intentions under the same skill are used as multiple subtasks in the task-based dialogue system. For example, the air ticket task includes a subtask of booking an air ticket and a subtask of querying an air ticket.

[0134] Step S403, extracting slot information of each candidate slot from the natural language to be recognized according to the slot description information of each candidate slot corresponding to the target intent and the slot information extraction condition, wherein the slot information of each candidate slot is used to define each key information of the subtask corresponding to the intent;

[0135] Each intent in the embodiment of the present application corresponds to at least one alternative slot. For example, in the intent of booking an air ticket, it is necessary to know the user identity information of the booked air ticket, the travel time of the air ticket, the departure city and the destination, etc. In general, the natural language to be identified usually contains all the slot information under the corresponding intent. In the embodiment of the present application, the slot extraction condition is to find the slot information in the natural language to be identified that corresponds to the slot description information of the alternative slot. For example, if the slot description information indicates that the corresponding alternative slot is to find the departure city, then according to the slot extraction condition, the information corresponding to the place is searched from the natural language to be identified, and according to the semantic information of the natural language to be identified, it is determined to be the location information of the departure place. During the implementation process, the slot information of each alternative slot is extracted from the natural language to be identified according to the slot description information of the alternative slot and the slot extraction condition. In the embodiment of the present application, the slot information of each alternative slot is used to limit the key information of the subtask corresponding to the intent.

[0136] Step S404, obtaining a dialog task recognition result of a natural language to be recognized according to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot.

[0137] In an embodiment of the present application, based on the skill description information of the target skill, the intent description information of the target intent, and the slot description information and slot information of each candidate slot, a task recognition result of the natural language to be recognized can be obtained, and then services can be provided to the user based on the task recognition result. As an optional implementation method, the natural language to be recognized can also be labeled based on the skill description information of the target skill, the intent description information of the target intent, and the slot description information and slot information of each candidate slot to obtain sample data under the corresponding skill.

[0138] like Figure 7 As shown, the natural language understanding model in the task-based dialogue system in the embodiment of the present application includes three sub-models, namely, a skill classification model, an intent recognition model, and a slot extraction model. Among them, the target skill in the embodiment of the present application is obtained based on the trained skill classification model, that is, the natural language to be recognized is input into the trained skill classification model to obtain the target skill that matches the natural language to be recognized, the target intent is obtained based on the trained intent recognition model, that is, the natural language to be recognized is input into the trained intent recognition model to obtain the target skill that matches the natural language to be recognized, and the slot information is obtained based on the trained slot extraction model, that is, the natural language to be recognized is input into the trained slot extraction model to extract the slot information of each candidate slot from the natural language to be recognized.

[0139] Conversational tasks include trained conversational tasks and untrained conversational tasks. Skill classification models, intent recognition models, and slot extraction models are trained using sample data from each trained conversational task. The skill classification model is trained using labeled sample data from multiple skills, the intent recognition model is trained using labeled sample data from multiple intents, and the slot extraction model is trained using labeled sample data from multiple slots. Tasks that have been trained using sample data are trained conversational tasks, and untrained conversational tasks include: skill description information, intent description information for each candidate intent, and candidate slot description information for each candidate intent. The skill classification model obtained after training in the embodiment of the present application is a general skill classification model. When a new skill needs to be developed, it is only necessary to define the skill description information of the new skill. Similarly, the intent recognition model obtained after training in the embodiment of the present application is a general intent recognition model. When a new intent recognition needs to be developed, it is only necessary to define the intent description information of the new intent. The slot extraction model obtained after training in the embodiment of the present application is a general slot extraction model. When a new slot needs to be developed, it is only necessary to define the slot description information of the new slot.

[0140] The following describes the training and use of the skill classification model, intent recognition model, and slot extraction model:

[0141] 1. Skill Classification Model

[0142] like Figure 8 As shown, the skill classification model includes a first vector representation module and a first classification module, wherein the first vector representation module is used to perform vector representation on the input data, and the first classification model is used to determine the similarity between at least two obtained vectors;

[0143] The skill classification corpus samples for training the skill classification model include at least one sample corresponding to each trained conversational task, and each sample includes the skill description information of the conversational task and a similarity probability value annotated with the conversational task to which it belongs. In an embodiment of the present application, the similarity probability value of the conversational task to which each skill classification corpus sample belongs can be, but is not limited to, 0 and 1. When the similarity probability value between the skill classification corpus sample and the conversational task is 1, the skill classification corpus sample matches the corresponding conversational task, otherwise it does not match.

[0144] Specifically, if the existing sample data contains N skills, namely D 1 ,D 2 ,…D N , Skill D n Contains I n Intention, S n slots and M nThe skill classification corpus sample with strip annotations, the set of all natural language data in all fields is M, then for any natural language data q and any skill D in M i , can generate a skill classification corpus sample (l,D i ,q), if q∈Di, then l=1, otherwise l=0.

[0145] In the embodiment of the present application, the trained skill classification model is trained in the following manner:

[0146] 1) obtaining first target vector representations of skill description information of each trained task based on the first vector representation module;

[0147] Among them, the vector representation obtained based on the first vector representation module can be but is not limited to a sentence vector representation. Those skilled in the art can set it according to actual needs, which will not be repeated here.

[0148] 2) for each skill classification corpus sample, obtaining a first reference vector representation of the skill classification corpus sample based on the first vector representation module;

[0149] 3) obtaining first similarities between the first reference vector representation and each first target vector representation based on the first classification module;

[0150] The first similarity between the first reference vector representation and each first target vector representation may be, but is not limited to, a value between 0 and 1. The greater the first similarity, the more similar the first reference vector representation and the corresponding first target vector representation are.

[0151] 4) adjusting the parameters of the first vector representation module and the first classification module based on the obtained first similarities until the first similarities meet the set conditions;

[0152] Among them, the similarity satisfies the set condition, and a similarity threshold can be set but is not limited to. When the first similarity is greater than the similarity threshold, it is considered that the first similarity satisfies the set condition. As an optional implementation method, the highest value of the first similarity can be used as satisfying the set condition. Those skilled in the art can set it according to actual needs, which will not be elaborated here.

[0153] The following describes the process of obtaining target skills based on the skill classification model:

[0154] 1) obtaining first comparison vector representations of skill description information of each trained conversational task and untrained conversational task based on the first vector representation module;

[0155] 2) obtaining a first vector representation of the natural language material to be recognized based on the first vector representation module;

[0156] 3) Based on the first classification module, second similarities between the first comparison vector representation and each first vector representation to be identified are respectively obtained, and a classification result is obtained based on the second similarities.

[0157] As an optional implementation, the skill classification model may be, but is not limited to, a BERT pre-trained model. Specifically, the structure of a BERT-based classification model is used. A schematic diagram of the structure of a BERT-based classification model is shown in FIG. Fig. 9 shown.

[0158] 2. Intent Recognition Model

[0159] like Fig.10 As shown, the intention recognition model includes a second vector representation module and a second classification module, wherein the second vector representation module is used to perform vector representation on the input data, and the second classification model is used to determine the similarity between at least two obtained vectors;

[0160] The intention recognition corpus samples for training the intention recognition model include at least one sample corresponding to each trained conversational task, and each intention recognition corpus sample is annotated with a similarity probability value of the alternative intent to which it belongs. In an embodiment of the present application, the similarity probability value of the alternative intent to which each intention recognition corpus sample belongs can be, but is not limited to, 0 and 1. When the similarity probability value between the intention recognition corpus sample and the alternative intent is 1, the intention recognition corpus sample matches the corresponding alternative intent, otherwise it does not match.

[0161] Specifically, if the existing sample data contains N skills, namely D 1 ,D 2 ,…D N , Skill D n Contains I n Intention, S n slots and M n The skill classification corpus samples with strip annotations, the set of all natural language data in all fields is M, and the intent recognition corpus samples include: for a certain field D i Sample data M i Any data q in the field and any intention Ii in the field can generate a labeled data (l, Ii, q) of the intention classification model. If the intention of q is Ii, then l = 1, otherwise l = 0.

[0162] In the embodiment of the present application, the trained intent recognition model is trained in the following manner:

[0163] 1) Based on the second vector representation module, a second target vector representation of the intent description information corresponding to each subtask in each trained task is obtained respectively;

[0164] Among them, the vector representation obtained based on the second vector representation module can be but is not limited to a sentence vector representation. Those skilled in the art can set it according to actual needs, which will not be elaborated here.

[0165] 2) for each intent recognition corpus sample, obtaining a second reference vector representation of the intent recognition corpus sample based on the second vector representation module;

[0166] 3) obtaining third similarities between the second reference vector representation and each second target vector representation based on the second classification module;

[0167] The second similarity between the second reference vector representation and each second target vector representation may be, but is not limited to, a value between 0 and 1. The greater the second similarity, the more similar the second reference vector representation and the corresponding second target vector representation are.

[0168] 4) adjusting the parameters of the second vector representation module and the second classification module based on each third similarity obtained until the third similarity meets the set condition;

[0169] Among them, the similarity satisfies the set condition, and a similarity threshold can be set but is not limited to. When the third similarity is greater than the similarity threshold, it is considered that the third similarity satisfies the set condition. As an optional implementation method, the highest value of the third similarity can be used as satisfying the set condition. Those skilled in the art can set it according to actual needs, which will not be elaborated here.

[0170] The following describes the process of obtaining the target intent based on the intent recognition model:

[0171] 1) obtaining second comparison vector representations of the intention description information of each subtask in the trained conversational task and each subtask in the untrained conversational task based on the second vector representation module;

[0172] 2) obtaining a second vector representation of the natural language material to be recognized based on the second vector representation module;

[0173] 3) Based on the second classification module, fourth similarities between the second comparison vector representation and each second vector representation to be identified are respectively obtained, and a classification result is obtained based on the fourth similarities.

[0174] As an optional implementation, the intent recognition model may, but is not limited to, adopt a BERT pre-trained model, specifically using the structure of a BERT-based classification model.

[0175] 3. Slot extraction model

[0176] like Fig.11As shown, the slot extraction model includes a third vector representation module and a slot extraction module, wherein the third vector representation module is used to perform vector representation on the input data, and the slot extraction module is used to extract slot information in the natural language;

[0177] The slot extraction corpus samples for training the slot extraction model include at least one sample for each trained conversational task. Each slot extraction corpus sample is annotated with reference slot information of the candidate slot to which it belongs. The trained slot extraction model is trained in the following way:

[0178] 1) for each slot extraction corpus sample, obtaining a third reference vector representation of the slot extraction corpus sample based on the third vector representation module;

[0179] Among them, the vector representation obtained based on the third vector representation module can be but is not limited to sentence vector representation, word vector representation or character vector representation. Technical personnel in this field can set it according to actual needs, which will not be elaborated here.

[0180] 2) obtaining first target slot information of slot description information corresponding to each candidate slot in the third reference vector representation based on the slot extraction module;

[0181] The first target slot information is also the slot information of the corresponding candidate slot in the slot extraction corpus sample. For example, if the candidate slot is a specific time representation, the first target slot information is the specific time information in the slot extraction corpus sample.

[0182] 3) adjusting the parameters of the third vector representation module and the slot extraction module based on the obtained first target slot information until the obtained first target slot information meets the set conditions;

[0183] Among them, the first target slot information meets the set conditions, and the slot information corresponding to the candidate slot can be accurately obtained in the slot extraction corpus sample.

[0184] The method for obtaining slot information based on the trained slot extraction model is as follows:

[0185] 1) obtaining a third vector representation of the natural language data to be recognized based on the third vector representation module;

[0186] 2) Based on the slot extraction module, second target slot information represented by the third vector to be identified is obtained respectively, and a classification result is obtained based on the second target slot information.

[0187] As an optional implementation, the slot extraction model may be, but is not limited to, a BERT pre-trained model, and specifically uses a structure of a BERT-based QA model. The schematic diagram of the structure of a BERT-based QA model is as follows: Fig.12 shown.

[0188] The following is a detailed description of a natural language recognition method in a task-based dialogue system provided by the present application in conjunction with a specific implementation method. Assume that the existing open dialogue platform contains two skills, weather dialogue task skills and flight dialogue task skills, as well as sample data of the related natural language recognition model. The skill description information and some sample data of the weather dialogue task are as follows: Fig.13 As shown in Figure 1, the skill description information and some sample data of the flight dialogue task are as follows: Fig.14 shown.

[0189] Through the skill description information of these two skills and the annotated sample data, the three general natural language recognition models proposed by the present invention can be trained, namely: a general skill classification model, a general intent recognition model, and a general slot extraction model. The main function of these three general natural language recognition models is to discover the semantic correlation between the skill description information of the skill and the natural language to be recognized input by the user in terms of skills, intent, and slots.

[0190] The skill classification model will learn the semantic relevance between the skill description information and the natural language input by the user to be recognized. For example, "air ticket" has significant semantic relevance to the keywords "airplane" and "flight" in the natural language input by the user to be recognized; "weather" has significant semantic relevance to the keywords "weather", "heat", "humidity" and other keywords in the natural language input by the user to be recognized.

[0191] The intent recognition model will learn the semantic correlation between the intent description information and the natural language input by the user to be recognized. For example, there is a significant semantic correlation between "search" and the key words "find" and "search" in the natural language input by the user to be recognized; there is a significant semantic correlation between "book" and the key words "book" and "buy" in the natural language input by the user to be recognized.

[0192] The slot extraction model will learn the semantic correlation between the slot description information and the natural language to be recognized input by the user. For example, "city" and city names such as "Beijing" and "Shanghai" have semantic correlation, "date" and "today" and "tomorrow" have semantic correlation; "departure" and "from" have semantic correlation; "arrival" and "go" and "to" have semantic correlation; "number of people" and numerals such as "two" have semantic correlation.

[0193] Assume that a developer has newly developed a skill for a hotel conversation task. The skill description information and the possible natural language input to be recognized by the user are as follows: Fig.15As shown. By utilizing the semantic correlation between the skill description information learned by the general natural language understanding model and the input natural language to be recognized, the general natural language understanding model proposed in the present invention can recognize the skill, intent and slot of the natural language to be recognized input by the user corresponding to the skill without labeled data.

[0194] First, the skill classification model will find that the semantic relevance of skills such as "hotel", "guesthouse", and "room" to the natural language to be identified by the user is low, but the relevance to the skill "hotel" is high, thereby completing the skill classification and obtaining that the skill to which the natural language to be identified by the user belongs is hotel. Through keywords such as "find" and "book" in the natural language to be identified by the user, the intent recognition model can distinguish between the two different intents of "query hotel" and "book hotel". The slot extraction model can also extract relevant keywords in the natural language to be identified by the user based on the keywords defined in the slots such as "city" and "date". The recognition results of the general natural language understanding model are as follows. Fig.16 shown.

[0195] Based on the natural language recognition method in the task-based dialogue system provided by the present application, the present application also provides a natural language annotation method, which can be executed by the task-based dialogue server, for example, Figure 3 The natural language annotation method in the task-based dialogue system provided in the embodiment of the present application is as follows Fig.17 As shown, Fig.17 The flowchart shown is described as follows.

[0196] Step S1701, determining a target skill that matches the semantics of the natural language to be recognized from each dialog task according to a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system;

[0197] Step S1702, determining a target intent that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intent description information of each candidate intent corresponding to the target skill, wherein each intent description information is used to describe a subtask included in the conversational task;

[0198] Step S1703, extracting slot information of each candidate slot from the natural language to be recognized according to the slot description information of each candidate slot corresponding to the target intent and the slot information extraction condition, wherein the slot information of each candidate slot is used to define each key information of the subtask corresponding to the intent;

[0199] Step S1704, annotating the natural language to be recognized according to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot to obtain sample data of the conversational task.

[0200] Among them, the target skill is obtained based on the trained skill classification model, the target intent is obtained based on the trained intent recognition model, the slot information is obtained based on the trained slot extraction model, the conversational tasks include trained conversational tasks and untrained conversational tasks, and the skill classification model, intent recognition model and slot extraction model are trained using sample data from each trained conversational task.

[0201] As mentioned above, after obtaining the general natural language understanding model, the embodiment of the present application can directly identify the natural language through the general natural language understanding model. Similarly, after obtaining the skill, intent and slot information of the natural language according to the general understanding model, the natural language can be labeled and the labeled natural language can be used as the sample data of the conversational task. According to the obtained sample data, the skill classification model, the intent recognition model and the slot extraction model are further updated and trained. It should be noted that the obtained sample data can include the natural language of the trained conversational task and the natural language of the untrained conversational task. The skill classification model, the intent recognition model and the slot extraction model are further updated and trained by the obtained sample data, which can improve the processing accuracy of the skill classification model, the intent recognition model and the slot extraction model for the trained conversational task. Secondly, when the obtained sample data includes the natural language identified as the untrained conversational task, the untrained conversational task is updated to the trained conversational task after the updated training.

[0202] As an optional implementation method, the embodiment of the present application can not only annotate the natural language input by the user through the general natural language understanding model, but also obtain and annotate the skills, intentions and slot information of the natural language through the skill language understanding model, intention recognition model and slot model in the relevant technology. The recognition and annotation of natural language, as well as the training of the model, can be performed regularly according to the update of the natural language library. Since the recognition and annotation of natural language can be completed automatically, the efficiency of model training and the efficiency of new skill development are greatly improved.

[0203] Based on the same inventive concept, the embodiment of the present application provides a natural language recognition device in a task-based dialogue system, and the natural language recognition device in the task-based dialogue system can be a hardware structure, a software module, or a hardware structure plus a software module. The natural language recognition device in the task-based dialogue system is, for example, the aforementioned Figure 3The server 302 itself, or a functional device disposed in the server 302, the natural language recognition device in the task-based dialogue system can be implemented by a chip system, which can be composed of a chip, or can include a chip and other discrete devices. Fig.18 As shown, the natural language recognition device in the task-based dialogue system in the embodiment of the present application includes a first skill determination unit 1801, a first intention determination unit 1802, a first slot determination unit 1803, and a recognition unit 1804, wherein:

[0204] The first skill determination unit 1801 is used to determine, from each dialog task, a target skill that matches the semantics of the natural language to be recognized based on a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system;

[0205] The first intention determination unit 1802 is used to determine a target intention that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intention description information of each candidate intention corresponding to the target skill, where each intention description information is used to describe a subtask included in the dialog task;

[0206] The first slot determination unit 1803 is used to extract slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent;

[0207] The recognition unit 1804 is used to obtain a dialog task recognition result of a natural language to be recognized based on the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot.

[0208] In a possible implementation, the target skill is obtained based on a trained skill classification model, the target intent is obtained based on a trained intent recognition model, and the slot information is obtained based on a trained slot extraction model, wherein:

[0209] Conversational tasks include trained conversational tasks and untrained conversational tasks. The skill classification model, intent recognition model, and slot extraction model are trained using sample data from each trained conversational task.

[0210] The untrained dialogue task includes: skill description information, intent description information of each alternative intent, and alternative slot description information of each alternative intent.

[0211] In a possible implementation, the skill classification model includes a first vector representation module and a first classification module. The skill classification corpus sample for training the skill classification model includes at least one sample corresponding to each trained dialog task. Each skill classification corpus sample is annotated with a similarity probability value of the corresponding dialog task. The trained skill classification model is trained in the following manner:

[0212] Obtaining first target vector representations of skill description information of each trained task based on the first vector representation module;

[0213] For each skill classification corpus sample, obtaining a first reference vector representation of the skill classification corpus sample based on a first vector representation module;

[0214] Obtaining first similarities between the first reference vector representation and each first target vector representation based on the first classification module;

[0215] The parameters of the first vector representation module and the first classification module are adjusted based on the obtained first similarities until the first similarities meet the set conditions.

[0216] In a possible implementation, the first skill determination unit 1801 is specifically configured to:

[0217] Based on the first vector representation module, first comparison vector representations of skill description information of each trained conversational task and untrained conversational task are obtained respectively;

[0218] Obtaining a first to-be-recognized vector representation of the natural language material to be recognized based on the first vector representation module;

[0219] Based on the first classification module, second similarities between the first comparison vector representation and each first vector representation to be identified are respectively obtained, and a classification result is obtained based on the second similarities.

[0220] In a possible implementation, the intent recognition model includes a second vector representation module and a second classification module. The intent recognition corpus sample for training the intent recognition model includes at least one sample corresponding to each trained conversational task. Each intent recognition corpus sample is annotated with a similarity probability value of the candidate intent to which it belongs. The trained intent recognition model is trained in the following manner:

[0221] Based on the second vector representation module, second target vector representations of the intent description information corresponding to each subtask in each trained task are respectively obtained;

[0222] For each intent recognition corpus sample, obtaining a second reference vector representation of the intent recognition corpus sample based on the second vector representation module;

[0223] Obtaining third similarities between the second reference vector representation and each second target vector representation based on the second classification module;

[0224] The parameters of the second vector representation module and the second classification module are adjusted based on the obtained third similarities until the third similarities meet the set conditions.

[0225] In a possible implementation, the first intention determination unit 1802 is specifically configured to:

[0226] Based on the second vector representation module, second comparison vector representations of the intention description information of each subtask in the trained conversational task and each subtask in the untrained conversational task are respectively obtained;

[0227] Obtaining a second to-be-recognized vector representation of the natural language material to be recognized based on the second vector representation module;

[0228] Based on the second classification module, fourth similarities between the second comparison vector representation and each second vector representation to be identified are respectively obtained, and a classification result is obtained based on the fourth similarities.

[0229] In a possible implementation, the slot extraction model includes a third vector representation module and a slot extraction module. The slot extraction corpus sample for training the slot extraction model includes at least one sample corresponding to each trained dialogue task. Each slot extraction corpus sample is annotated with reference slot information of the candidate slot to which it belongs. The trained slot extraction model is trained in the following manner:

[0230] For each slot extraction corpus sample, obtaining a third reference vector representation of the slot extraction corpus sample based on a third vector representation module;

[0231] Based on the slot extraction module, first target slot information of slot description information corresponding to each candidate slot in the third reference vector representation is respectively obtained;

[0232] The parameters of the third vector representation module and the slot extraction module are adjusted based on the obtained first target slot information until the obtained first target slot information meets the set conditions.

[0233] In a possible implementation, the first slot determining unit 1803 is specifically configured to:

[0234] Obtaining a third to-be-recognized vector representation of the natural language material to be recognized based on the third vector representation module;

[0235] Based on the slot extraction module, second target slot information represented by the third vector to be identified is obtained respectively, and a classification result is obtained based on the second target slot information.

[0236] In a possible implementation, the first labeling unit 1805 is further included, which is specifically configured to:

[0237] Annotate natural language identified as untrained conversational tasks;

[0238] Obtain sample data for untrained conversational tasks;

[0239] Based on the obtained sample data, the skill classification model, intent recognition model and slot extraction model are updated and trained; and

[0240] Update untrained conversational tasks to trained conversational tasks.

[0241] All relevant contents of each step involved in the embodiment of the natural language recognition method in the aforementioned task-based dialogue system can be referred to the functional description of the functional module corresponding to the natural language recognition device in the task-based dialogue system in the embodiment of the present application, and will not be repeated here.

[0242] Based on the same inventive concept, the embodiment of the present application provides a natural language annotation device in a task-based dialogue system, and the natural language annotation device in the task-based dialogue system can be a hardware structure, a software module, or a hardware structure plus a software module. The natural language recognition device in the task-based dialogue system is, for example, the aforementioned Figure 3 The server 302 itself, or a functional device disposed in the server 302, the natural language annotation device in the task-based dialogue system can be implemented by a chip system, which can be composed of a chip, or can include a chip and other discrete devices. Fig.19 As shown, the natural language annotation device in the task-based dialogue system in the embodiment of the present application includes a second skill determination unit 1901, a second intention determination unit 1902, a second slot determination unit 1903, and a second annotation unit 1904, wherein:

[0243] The second skill determination unit 1901 is used to determine, from each dialog task, a target skill that matches the semantics of the natural language to be recognized based on a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system;

[0244] The second intention determination unit 1902 is used to determine a target intention that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intention description information of each candidate intention corresponding to the target skill, wherein each intention description information is used to describe a subtask included in the dialog task;

[0245] The second slot determination unit 1903 is used to extract slot information of each candidate slot from the natural language to be recognized according to the slot description information of each candidate slot corresponding to the target intent and the slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent;

[0246] The second labeling unit 1904 is used to label the natural language to be recognized according to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot to obtain sample data of the dialogue task.

[0247] In a possible implementation, the target skill is obtained based on a trained skill classification model, the target intent is obtained based on a trained intent recognition model, and the slot information is obtained based on a trained slot extraction model, wherein:

[0248] The conversational tasks include trained conversational tasks and untrained conversational tasks. The skill classification model, intent recognition model and slot extraction model are trained using sample data of each trained conversational task.

[0249] In a possible implementation, a training unit 1905 is further included, which is specifically configured to:

[0250] Based on the obtained sample data, the skill classification model, intent recognition model and slot extraction model are further updated and trained, wherein when the obtained sample data includes natural language identified as an untrained conversational task, the untrained conversational task is updated to a trained conversational task after the updated training.

[0251] All relevant contents of each step involved in the embodiment of the natural language annotation method in the aforementioned task-based dialogue system can be referred to the functional description of the functional module corresponding to the natural language annotation device in the task-based dialogue system in the embodiment of the present application, and will not be repeated here.

[0252] The division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application may be integrated into a processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0253] Based on the same inventive concept, the present application embodiment provides a computing device, which is, for example, the aforementioned Figure 3The server 302 in the embodiment of the present application can execute the natural language recognition method in the task-based dialogue system provided by the embodiment of the present application, such as Fig. 20 As shown, the computing device in the embodiment of the present application includes at least one processor 2001, and a memory 2002 and a communication interface 2003 connected to the at least one processor 2001. The specific connection medium between the processor 2001 and the memory 2002 is not limited in the embodiment of the present application. Fig. 20 In the example, the processor 2001 and the memory 2002 are connected via a bus 2000. Fig. 20 The connections between other components are shown in bold lines, and are not intended to be limiting. The bus 2000 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 20 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0254] In an embodiment of the present application, the memory 2002 stores a computer program that can be executed by at least one processor 2001. By executing the computer program stored in the memory 2002, the at least one processor 2001 can execute the steps included in the natural language recognition method in the aforementioned task-based dialogue system.

[0255] Among them, the processor 2001 is the control center of the computing device, and can use various interfaces and lines to connect various parts of the entire computing device, and through running or executing instructions stored in the memory 2002 and calling data stored in the memory 2002, various functions of the computing device and processing data, so as to monitor the computing device as a whole. Optionally, the processor 2001 may include one or more processing modules, and the processor 2001 may integrate an application processor and a modem processor, wherein the processor 2001 mainly processes the operating system, user interface and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 2001. In some embodiments, the processor 2001 and the memory 2002 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on independent chips.

[0256] Processor 2001 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware processor to be executed, or can be executed by a combination of hardware and software modules in the processor.

[0257] The memory 2002 is a non-volatile computer-readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 2002 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 2002 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 2002 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0258] The communication interface 2003 is a transmission interface that can be used for communication. Data can be received or sent through the communication interface 2003. For example, data can be exchanged with other devices through the communication interface 2003 to achieve the purpose of communication.

[0259] Furthermore, the computing device also includes a basic input / output system (I / O system) 2004 for facilitating information transmission between various components within the computing device, and a large-capacity storage device 2008 for storing an operating system 2005, application programs 2006 and other program modules 2007.

[0260] The basic input / output system 2004 includes a display 2009 for displaying information and an input device 2010 such as a mouse and a keyboard for user inputting information. The display 2009 and the input device 2010 are connected to the processor 2001 via the basic input / output system 2004 connected to the system bus 2000. The basic input / output system 2004 may also include an input / output controller for receiving and processing inputs from a plurality of other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller also provides output to a display screen, a printer, or other types of output devices.

[0261] The mass storage device 2008 is connected to the processor 2001 through a mass storage controller (not shown) connected to the system bus 2000. The mass storage device 2008 and its associated computer readable media provide non-volatile storage for the server package. That is, the mass storage device 2008 may include a computer readable medium (not shown) such as a hard disk or a CD-ROM drive.

[0262] According to various embodiments of the present application, the computing device package can also be connected to a remote computer on the network through a network such as the Internet. That is, the computing device can be connected to the network 2011 through the communication interface 2003 connected to the system bus 2000, or the communication interface 2003 can be used to connect to other types of networks or remote computer systems (not shown).

[0263] Based on the same inventive concept, an embodiment of the present application also provides a storage medium, which may be a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed on a computer, the computer executes the steps of the natural language recognition method in the aforementioned task-based dialogue system.

[0264] Based on the same inventive concept, the embodiment of the present application also provides a chip system, which includes a processor and may also include a memory, for implementing the steps of the natural language recognition method in the aforementioned task-based dialogue system. The chip system may be composed of a chip, or may include a chip and other discrete devices.

[0265] In some possible implementations, various aspects of the method for recommending content provided in the embodiments of the present application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer, the program code is used to enable the computer to execute the steps of the control method for updating content in the content recommendation pool according to various exemplary embodiments of the present application described above.

[0266] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0267] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A natural language recognition method in a task-based dialogue system, It is characterized in that The method includes: According to a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system, determining a target skill that matches the semantics of the natural language to be recognized from each dialog task; Determine a target intent that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intent description information of each candidate intent corresponding to the target skill, wherein each intent description information is used to describe a subtask included in the conversational task; Extracting slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent; According to the skill description information of the target skill, the intention description information of the target intent, and the slot description information and slot information of each candidate slot, the dialog task recognition result of the natural language to be recognized is obtained; the dialog task includes a trained dialog task and an untrained dialog task; the untrained dialog task is used to label the natural language identified as the untrained dialog task according to the skill description information, the intention description information and the slot description information of each candidate slot corresponding to the untrained dialog task to obtain sample data of the untrained dialog task, and update the general natural language understanding model based on the sample data of the untrained dialog task to update the untrained dialog task to a trained dialog task; the target skill, target intent and slot information are obtained based on the general natural language understanding model.

2. The method according to claim 1, It is characterized in that The universal natural language understanding model includes a skill classification model, an intent recognition model, and a slot extraction model; the target skill is obtained based on the skill classification model in the trained universal natural language understanding model, the target intent is obtained based on the intent recognition model in the trained universal natural language understanding model, and the slot information is obtained based on the slot extraction model in the trained universal natural language understanding model, wherein: The skill classification model, intent recognition model and slot extraction model are trained using sample data of each trained conversational task; The untrained dialogue task includes: skill description information, intent description information of each alternative intent, and alternative slot description information of each alternative intent.

3. The method according to claim 2, It is characterized in that The skill classification model includes a first vector representation module and a first classification module. The skill classification corpus sample for training the skill classification model includes at least one sample corresponding to each trained dialog task. Each skill classification corpus sample is annotated with a similarity probability value of the corresponding dialog task. The trained skill classification model is trained in the following manner: Based on the first vector representation module, first target vector representations of skill description information of each trained task are respectively obtained; For each skill classification corpus sample, obtaining a first reference vector representation of the skill classification corpus sample based on a first vector representation module; Obtaining first similarities between the first reference vector representation and each first target vector representation based on the first classification module; The parameters of the first vector representation module and the first classification module are adjusted based on the obtained first similarities until the first similarities meet the set conditions.

4. The method according to claim 3, It is characterized in that The target skills are obtained based on the trained skill classification model, and specifically include: Based on the first vector representation module, first comparison vector representations of skill description information of each trained conversational task and untrained conversational task are obtained respectively; Obtaining a first to-be-recognized vector representation of the to-be-recognized natural language material based on the first vector representation module; Based on the first classification module, second similarities between the first comparison vector representation and each first vector representation to be identified are respectively obtained, and a classification result is obtained based on the second similarities.

5. The method according to claim 2, It is characterized in that The intention recognition model includes a second vector representation module and a second classification module. The intention recognition corpus sample for training the intention recognition model includes at least one sample corresponding to each trained dialogue-type task. Each intention recognition corpus sample is annotated with a similarity probability value of the candidate intention. The trained intention recognition model is trained in the following manner: Based on the second vector representation module, second target vector representations of the intent description information corresponding to each subtask in each trained task are respectively obtained; For each intent recognition corpus sample, obtaining a second reference vector representation of the intent recognition corpus sample based on the second vector representation module; Obtaining third similarities between the second reference vector representation and each second target vector representation based on the second classification module; The parameters of the second vector representation module and the second classification module are adjusted based on the obtained third similarities until the third similarities meet the set conditions.

6. The method according to claim 5, It is characterized in that The target intent is obtained based on the trained intent recognition model, specifically including: Based on the second vector representation module, second comparison vector representations of the intention description information of each subtask in the trained conversational task and each subtask in the untrained conversational task are respectively obtained; Obtaining a second to-be-recognized vector representation of the to-be-recognized natural language material based on the second vector representation module; Based on the second classification module, fourth similarities between the second comparison vector representation and each second vector representation to be identified are respectively obtained, and a classification result is obtained based on the fourth similarities.

7. The method according to claim 2, It is characterized in that The slot extraction model includes a third vector representation module and a slot extraction module. The slot extraction corpus sample for training the slot extraction model includes at least one sample corresponding to each trained dialogue task. Each slot extraction corpus sample is annotated with reference slot information of the candidate slot to which it belongs. The trained slot extraction model is trained in the following manner: For each slot extraction corpus sample, obtaining a third reference vector representation of the slot extraction corpus sample based on a third vector representation module; Based on the slot extraction module, first target slot information of slot description information corresponding to each candidate slot in the third reference vector representation is respectively obtained; The parameters of the third vector representation module and the slot extraction module are adjusted based on the obtained first target slot information until the obtained first target slot information meets the set conditions.

8. The method according to claim 7, It is characterized in that The slot information is obtained based on the trained slot extraction model, specifically including: Obtaining a third to-be-recognized vector representation of the to-be-recognized natural language material based on the third vector representation module; Based on the slot extraction module, second target slot information represented by the third vector to be identified is obtained respectively, and a classification result is obtained based on the second target slot information.

9. A natural language annotation method in a task-based dialogue system. It is characterized in that include: According to a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system, determining a target skill that matches the semantics of the natural language to be recognized from each dialog task; Determine a target intent that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intent description information of each candidate intent corresponding to the target skill, wherein each intent description information is used to describe a subtask included in the conversational task; Extracting slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent; According to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot, the natural language to be recognized is labeled to obtain sample data of a dialog task; the dialog task includes a trained dialog task and an untrained dialog task; The untrained conversational task is used to annotate the natural language identified as the untrained conversational task according to the skill description information, the intention description information, and the slot description information of each candidate slot corresponding to the untrained conversational task to obtain sample data of the untrained conversational task, and update the general natural language understanding model based on the sample data of the untrained conversational task to update the untrained conversational task to a trained conversational task; The target skills, target intentions, and slot information are obtained based on the general natural language understanding model.

10. The method according to claim 9, It is characterized in that The universal natural language understanding model includes a skill classification model, an intent recognition model, and a slot extraction model; the target skill is obtained based on the skill classification model in the trained universal natural language understanding model, the target intent is obtained based on the intent recognition model in the trained universal natural language understanding model, and the slot information is obtained based on the slot extraction model in the trained universal natural language understanding model, wherein: The skill classification model, intent recognition model and slot extraction model are trained using sample data of each trained dialogue task.

11. A natural language recognition device in a task-based dialogue system, It is characterized in that include: a first skill determination unit, configured to determine, from each dialog-type task, a target skill that matches the semantics of the natural language to be recognized, based on a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog-type task in the task-based dialog system; A first intention determination unit, configured to determine a target intention that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intention description information of each candidate intention corresponding to the target skill, wherein each intention description information is used to describe a subtask included in the dialog task; A first slot determination unit is used to extract slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent; an identification unit, configured to obtain a recognition result of a dialog task of the natural language to be identified according to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot; the dialog task includes a trained dialog task and an untrained dialog task; The untrained conversational task is used to annotate the natural language identified as the untrained conversational task according to the skill description information, the intention description information, and the slot description information of each candidate slot corresponding to the untrained conversational task to obtain sample data of the untrained conversational task, and update the general natural language understanding model based on the sample data of the untrained conversational task to update the untrained conversational task to a trained conversational task; The target skills, target intentions, and slot information are obtained based on a general natural language understanding model.

12. A natural language annotation device in a task-based dialogue system, It is characterized in that include: a second skill determination unit, configured to determine, from each dialog task, a target skill that matches the semantics of the natural language to be recognized, based on a matching condition between the semantic information of the natural language to be recognized and the skill description information of each dialog task in the task-based dialog system; A second intention determination unit, configured to determine a target intention that matches the semantics of the natural language to be identified based on a matching condition between the semantic information of the natural language to be identified and the intention description information of each candidate intention corresponding to the target skill, wherein each intention description information is used to describe a subtask included in the dialog task; A second slot determination unit is used to extract slot information of each candidate slot from the natural language to be recognized according to slot description information of each candidate slot corresponding to the target intent and a slot information extraction condition, wherein the slot information of each candidate slot is used to limit each key information of the subtask corresponding to the intent; a second labeling unit, configured to label the natural language to be recognized according to the skill description information of the target skill, the intention description information of the target intention, and the slot description information and slot information of each candidate slot, so as to obtain sample data of a dialog task; the dialog task includes a trained dialog task and an untrained dialog task; The untrained conversational task is used to annotate the natural language identified as the untrained conversational task according to the skill description information, the intention description information, and the slot description information of each candidate slot corresponding to the untrained conversational task to obtain sample data of the untrained conversational task, and update the general natural language understanding model based on the sample data of the untrained conversational task to update the untrained conversational task to a trained conversational task; The target skills, target intentions, and slot information are obtained based on the general natural language understanding model.

13. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the method described in any one of claims 1-8, or 9-10 are implemented.

14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 8 or 9 to 10.

Citation Information

Patent Citations

  • Model training method and device based on dialogue template

    CN110008319A

  • Multi-round dialogue intelligent customer service system based on proprietary word correction and cold start

    CN110825865A

  • Conversation state tracking method and device and calculation equipment

    CN111090728A