Reply Information Generation Method, Device, Computer Device, and Storage Medium

Through binary classification of dialogue information and semantic analysis of large language model, the intent type is determined and the target dialogue model is used to generate reply information, which solves the problem of poor accuracy of reply information in the existing technology and achieves higher accuracy of reply information.

CN117520504BActive Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311499361.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-07-22
Estimated Expiration
2043-11-10

AI Technical Summary

Technical Problem

In the prior art, the accuracy of the reply information generation in artificial intelligence dialogue scenarios is poor.

Method used

The dialogue information is classified using a binary classification method, the large language model is used for semantic analysis, the intent type is determined, and the reply information is generated through the corresponding target dialogue model.

Benefits of technology

It improves the accuracy of reply information, ensures that the dialogue model is more consistent with the dialogue needs, and improves the accuracy of reply information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117520504B_ABST
    Figure CN117520504B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method, device, computer device and storage medium for generating reply information, belonging to the field of computer technology. The method includes: classifying the dialogue information to obtain the category of the dialogue information, and in the case where the category of the dialogue information is the first category, performing semantic analysis on the dialogue information through a large language model to obtain the semantic information of the dialogue information; based on the semantic information, determining a first intention type from multiple intention types, and the first intention type matches the semantic information; processing the dialogue information through a target dialogue model under the first intention type to obtain a first reply information. The present application adopts a binary classification method for rough screening to determine which type of dialogue model to use for reply. In the case of determining to use a dialogue model for reply under multiple intention types, the large language model is used for accurate matching to ensure that the dialogue model used for reply is more matched with the dialogue requirement, thereby ensuring the accuracy of the reply information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and particularly to a method, apparatus, computer device, and storage medium for generating reply information. Background Art

[0002] With the development of computer technology, the application of artificial intelligence conversations is becoming more and more widespread. In the scenario of artificial intelligence conversations, a conversation model is usually trained for this scenario. When a user inputs conversation information, the conversation model is used to generate a reply message for the conversation information and output the reply message to achieve artificial intelligence conversations. However, this way of generating reply messages results in poor accuracy of the reply messages. Summary of the Invention

[0003] The embodiments of the present application provide a method, apparatus, computer device, and storage medium for generating reply information, which can ensure the accuracy of the reply information. The technical solutions are as follows:

[0004] On the one hand, a method for generating reply information is provided. The method includes:

[0005] Classify the conversation information to obtain the category of the conversation information. The category includes a first category or a second category. The first category indicates multiple target conversation models, and each target conversation model is used to reply to a type of conversation information with a specific intention. The second category indicates a general conversation model;

[0006] When the category of the conversation information is the first category, perform semantic analysis on the conversation information through a large language model to obtain the semantic information of the conversation information;

[0007] Based on the semantic information, determine a first intention type from the multiple intention types, where the first intention type matches the semantic information;

[0008] Process the conversation information through the target conversation model under the first intention type to obtain a first reply message.

[0009] On the other hand, a device for generating reply information is provided. The device includes:

[0010] A classification module, configured to classify the conversation information to obtain the category of the conversation information. The category includes a first category or a second category. The first category indicates multiple target conversation models, and each target conversation model is used to reply to a type of conversation information with a specific intention. The second category indicates a general conversation model;

[0011] An analysis module, configured to, when the category of the conversation information is the first category, perform semantic analysis on the conversation information through a large language model to obtain semantic information of the conversation information;

[0012] A determination module, configured to determine a first intent type from the multiple intent types based on the semantic information, where the first intent type matches the semantic information;

[0013] A processing module, configured to process the conversation information through a target conversation model under the first intent type to obtain a first reply message.

[0014] In a possible implementation, the conversation information is question information; the analysis module is configured to, when the category of the conversation information is the first category, identify the question type of the question information through the large language model; classify the question information through the large language model to obtain a second intent type, where the second intent type is the intent type in the multiple intent types that matches the question type; and form the semantic information by using the second intent type and the question type through the large language model.

[0015] In another possible implementation, the apparatus further includes:

[0016] An identification module, configured to identify at least one of the topic of the question information, the entity words in the question information, or the word types of the entity words through the large language model;

[0017] The analysis module is configured to form the semantic information by using at least one of the topic, the entity words, or the word types, the second intent type, and the question type through the large language model.

[0018] In another possible implementation, the determination module is configured to query a type mapping table based on the semantic information, where the type mapping table includes the question types corresponding to each intent type in the multiple intent types; and when it is queried that the question type corresponding to the second intent type in the type mapping table is the same as the question type in the semantic information, determine the second intent type as the first intent type.

[0019] In another possible implementation, the apparatus further includes:

[0020] An acquisition module, configured to acquire first indication information, where the first indication information instructs the large language model to perform semantic analysis on input information according to an example of semantic analysis, and the example includes an input information example and a semantic information example of the input information example;

[0021] The analysis module is used to, when the category of the conversation information is the first category, perform semantic analysis on the conversation information based on the first indication information through the large language model to obtain the semantic information of the conversation information.

[0022] In another possible implementation, the processing module is further used to, when the category of the conversation information is the second category, process the conversation information through the general conversation model to obtain a second reply message.

[0023] In another possible implementation, the processing module is further used to, when none of the multiple intent types match the semantic information, process the conversation information through the general conversation model to obtain a second reply message.

[0024] In another possible implementation, the device further includes:

[0025] An acquisition module, configured to acquire sample conversation information and second indication information, where the second indication information instructs a semantic analysis model to perform semantic analysis on input information according to an example of semantic analysis;

[0026] The analysis module is further used to perform semantic analysis on the sample conversation information based on the second indication information through the semantic analysis model to obtain sample semantic information;

[0027] The processing module is further used to process the sample conversation information through the large language model to obtain predicted semantic information;

[0028] A training module, configured to train the large language model based on the predicted semantic information and the sample semantic information.

[0029] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the reply message generation method as described in the above aspect.

[0030] On the other hand, a computer-readable storage medium is provided, where at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the reply message generation method as described in the above aspect.

[0031] On yet another hand, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the operations performed by the reply message generation method as described in the above aspect.

[0032] The solution provided by the embodiments of the present application pre-sets a general dialogue model and multiple target dialogue models. The multiple target dialogue models and the general dialogue model belong to different categories. After obtaining the dialogue information, the dialogue information is first classified to identify which type of dialogue model to use for reply. In the case where the category of the dialogue information is determined to be the first category, the semantic information of the dialogue information is analyzed through a large language model, so as to use the semantic information to determine the first intention type that matches the semantic information from multiple intention types, and then use the dialogue model under the first intention type to generate the corresponding reply information. In this way, a simple binary classification method is adopted to roughly screen the dialogue information to determine which type of dialogue model to use for reply. In the case where it is determined to use the dialogue models under multiple intention types for reply, the large language model will be used for precise matching to determine which intention type the dialogue requirement of the dialogue information matches, and then use the dialogue model under the determined intention type for reply. This can ensure that the used dialogue model better matches the dialogue requirement of the dialogue information, and thus ensure the accuracy of the reply information. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 is a schematic structural diagram of an implementation environment provided by the embodiments of the present application;

[0035] Figure 2 is a flowchart of a method for generating reply information provided by the embodiments of the present application;

[0036] Figure 3 is a flowchart of another method for generating reply information provided by the embodiments of the present application;

[0037] Figure 4 is a schematic diagram for classifying dialogue information provided by the embodiments of the present application;

[0038] Figure 5 is a flowchart of still another method for generating reply information provided by the embodiments of the present application;

[0039] Figure 6 is a schematic diagram of the proportion of dialogue information provided by the embodiments of the present application;

[0040] Figure 7 is a schematic structural diagram of a device for generating reply information provided by the embodiments of the present application;

[0041] Figure 8 It is a schematic structural diagram of another reply information generation device provided by an embodiment of the present application;

[0042] Figure 9 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;

[0043] Figure 10 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0045] The terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first category may be referred to as the second category, and similarly, the second category may be referred to as the first category.

[0046] The terms "at least one", "multiple", "each", "any one" used in the present application, at least one includes one, two or more than two, multiple includes two or more than two, and each refers to each one in the corresponding multiple, and any one refers to any one in the multiple. For example, multiple intention types include 3 intention types, and each refers to each of these 3 intention types, and any one refers to any one of these 3 intention types, which can be the first intention type, or the second intention type, or the third intention type.

[0047] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the conversation information involved in the present application is obtained under full authorization.

[0048] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0049] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0050] Computer Vision (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition and measurement, and further performing graphic processing to make the computer process the image into a more suitable one for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build an artificial intelligence system that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision such as Swin Transformer (a deep learning model), ViT (Vision Transformer, a deep learning model), V-MOE (Vision MoE, a vision model), MAE (Masked Auto Encoders, auto encoder) can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3 Dimensions) technology, virtual reality, and augmented reality.

[0051] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves important technologies for model training in artificial intelligence fields such as computer science and mathematics. The pre-trained model (Pre-trained Models, PTM) is developed from the large language model (Large Language Model, LLM) in the NLP field. The pre-trained model, also known as the foundation model or large model, refers to a deep neural network (Deep Neural Network, DNN) with large parameters. It is trained on a large amount of unlabeled data, and uses the function approximation ability of the large-parameter DNN to enable the PTM to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and Prompt-Tuning (prompt tuning), it is applicable to downstream tasks. Therefore, the pre-trained model can achieve ideal results in few-shot or zero-shot scenarios. PTMs are divided into language models, visual models, speech models, multi-modal models, etc. according to the data modalities they process. For example, language models include ELMO (Embeddings from Language Model, a language model), BERT (Bidirectional Encoder Representations from Transformers, a bidirectional pre-trained language model), GPT (Generative Pre-trained Transformer, a pre-trained generative model), etc. Among them, the multi-modal model refers to a model that establishes feature representations of two or more data modalities. The pre-trained model is an important tool for outputting Artificial Intelligence Generated Content (AIGC), and can also be used as a general interface connecting multiple specific task models. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0052] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0053] The solution provided by the embodiments of this application, based on the machine learning technology of artificial intelligence, can train a large language model, and then use the trained large language model to implement a method for generating reply information.

[0054] The method for generating reply information provided by the embodiments of this application can be executed by a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers. Optionally, it is a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, intelligent voice interaction device, smart home appliance, in-vehicle terminal, etc., but is not limited thereto.

[0055] In some embodiments, the computer program involved in the embodiments of this application can be deployed to be executed on a single computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network can form a blockchain system.

[0056] In some embodiments, the computer device is provided as a server. Figure 1 It is a schematic diagram of an implementation environment provided by the embodiments of this application. Refer to Figure 1 In this, the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected through a wireless or wired network.

[0057] The terminal 101 is used to obtain the conversation information input by the user and send the conversation information to the server 102. The server 102 is used to receive the conversation information sent by the terminal 101, generate a reply message for the conversation information based on the reply message generation method provided in the embodiments of the present application, send the reply message to the terminal 101, so that the terminal 101 displays the reply message, and realize human-machine intelligent conversation.

[0058] In some embodiments, an application provided by the server 102 is installed on the terminal 101, and the terminal 101 can implement functions such as human-machine intelligent conversation through this application. Optionally, the application is an application in the operating system of the terminal 101, or an application provided by a third party. For example, the application is a conversation application, and this conversation application has the function of human-machine intelligent conversation. Of course, this conversation application can also have other functions, such as a review function, a shopping function, a navigation function, a game function, etc.

[0059] The terminal 101 is used to log in to the application based on the user identification. Through the application, a conversation interface can be displayed. The user can input conversation information in the conversation interface through the terminal 101. After the terminal 101 obtains the conversation information, it sends the conversation information to the server 102 through the application. The server 102 is used to receive the conversation information, generate a reply message for the conversation information based on the reply message generation method provided in the embodiments of the present application, send the reply message to the terminal 101, and the terminal 101 receives the reply message and displays the reply message in the conversation interface.

[0060] Figure 2 It is a flowchart of a reply message generation method provided in the embodiments of the present application. This method is executed by a computer device, such as Figure 2 shown, and this method includes:

[0061] 201. The computer device classifies the conversation information to obtain the category of the conversation information. The category includes a first category or a second category. The first category indicates multiple target conversation models, and each target conversation model is used to reply to a conversation information of a certain intent type. The second category indicates a general conversation model.

[0062] In the embodiments of the present application, a general conversation model and multiple target conversation models are preset. The multiple target conversation models correspond one-to-one with multiple intent types, and the intent types corresponding to different target conversation models are different, so that subsequently, multiple target conversation models can be used to reply to conversation information of multiple intent types, and for conversation information that does not belong to multiple intent types, it is replied through the general conversation model.

[0063] In the embodiments of the present application, multiple target dialogue models are equivalent to dialogue models in multiple special scenarios. When the dialogue information applies to one of these multiple special scenarios, a reply message is subsequently generated by one of the multiple target dialogue models. When the dialogue information does not apply to the multiple special scenarios, a reply message is subsequently generated by the general dialogue model. The multiple target dialogue models and the general dialogue model belong to different categories. When the dialogue information is obtained, the dialogue information is classified to determine which type of dialogue model will reply to the dialogue information subsequently.

[0064] Among them, the dialogue information can be any type of information. For example, the dialogue information is text, image, video, etc. The intent type is any type. For example, the intent type includes weather type, code type, text-to-image type, stock type, etc. The target dialogue model can be any network model, and the general dialogue model can be any network model. For example, both the target dialogue model and the general dialogue model can be large language models.

[0065] 202. When the category of the dialogue information is the first category, the computer device performs semantic analysis on the dialogue information through a large language model to obtain the semantic information of the dialogue information.

[0066] In the embodiments of the present application, the category of the dialogue information being the first category means that the dialogue information may match any one of multiple intent types. Therefore, semantic analysis is performed on the dialogue information through a large language model to determine the semantics of the dialogue information, so as to subsequently determine which intent type the dialogue information matches based on the semantics of the dialogue information.

[0067] Among them, the semantic information can be any type of information. For example, the semantic information is text information.

[0068] 203. The computer device determines a first intent type from multiple intent types based on the semantic information, and the first intent type matches the semantic information.

[0069] In the embodiments of the present application, when the semantic information of the dialogue information is determined, the semantic information can indicate the semantics of the dialogue information. Then, based on the semantics indicated by the semantic information, an intent type that matches the semantic information can be determined from multiple intent types, that is, the intent type that matches the dialogue information is determined.

[0070] 204. The computer device processes the dialogue information through the target dialogue model under the first intent type to obtain a first reply message.

[0071] In the embodiment of the present application, if the first intent type matches the dialogue information, the target dialogue model under the first intent type is used to process the dialogue information so that the obtained first response information matches the dialogue information, ensuring that the first response information is more accurate.

[0072] The solution provided by the embodiment of the present application presets a general dialogue model and multiple target dialogue models. The multiple target dialogue models and the general dialogue model belong to different categories. After obtaining the dialogue information, the dialogue information is first classified to identify which type of dialogue model to use for the response. In the case where the category of the dialogue information is determined to be the first category, the semantic information of the dialogue information is analyzed through a large language model, so as to use the semantic information to determine the first intent type that matches the semantic information from multiple intent types, and then use the dialogue model under the first intent type to generate the corresponding response information. In this way, a simple binary classification method is adopted to roughly screen the dialogue information to determine which type of dialogue model to use for the response. In the case where it is determined to use the dialogue models under multiple intent types for the response, the large language model will be used for precise matching to determine which intent type the dialogue requirement of the dialogue information matches, and then use the dialogue model under the determined intent type for the response. This can ensure that the used dialogue model better matches the dialogue requirement of the dialogue information, and thus ensure the accuracy of the response information.

[0073] In Figure 2 Based on the embodiment shown, the embodiment of the present application takes the dialogue information as the question information as an example. The semantic information obtained through the large language model includes the intent type and the question type, and then in combination with the type mapping table, the first intent type is determined. The specific process is detailed in the following embodiments.

[0074] Figure 3 is a flowchart of a method for generating response information provided by an embodiment of the present application. This method is executed by a computer device, as Figure 3 shown, and this method includes:

[0075] 301. The computer device classifies the dialogue information to obtain the category of the dialogue information. The category includes the first category or the second category. The first category indicates multiple target dialogue models, and each target dialogue model is used to respond to the dialogue information of one intent type. The second category indicates the general dialogue model.

[0076] In the embodiment of the present application, the dialogue information is question information. For example, the question information is "What's the weather like today" or "Write a piece of code to traverse the folder".

[0077] In a possible implementation manner, step 301 includes: classifying the dialogue information through a classification model to obtain the category of the dialogue information.

[0078] Among them, the classification model can be any network model. For example, the classification model is a lightweight model, such as BERT (Bidirectional Encoder Representations from Transformers), as Figure 4 shown, the dialogue information includes n characters. The dialogue information is input into BERT, and BERT classifies the dialogue information and outputs the category of the dialogue information.

[0079] In the embodiment of the present application, the classification model is a binary classification model, which is used to classify the dialogue information to determine whether to reply through one of multiple target dialogue models or through a general dialogue model.

[0080] 302. When the category of the dialogue information is the first category, the computer device identifies the problem type of the problem information through a large language model.

[0081] In the embodiment of the present application, the dialogue information is problem information, and the problem types include query type, programming type, creation type, etc. The large language model can be any type of model. For example, the large language model is GPT (Generative Pre-Trained Transformer), and this large language model is obtained by performing SFT (Supervised Fine-Tuning) on the pre-trained GPT. Among them, GPT is constructed using a Transformer decoder module.

[0082] 303. The computer device classifies the problem information through the large language model to obtain a second intention type, and the second intention type is the intention type that matches the problem type among multiple intention types.

[0083] In the embodiment of the present application, the problem information is classified through the large language model to determine the intention type that matches the problem type of the problem information from multiple intention types.

[0084] For example, if the problem information is "draw a moon", the second intention type is the text-to-image type; if the problem information is "write a quicksort code", the second intention type is the code type; if the problem information is "what's the weather like today", the second intention type is the weather type; if the problem information is "what's 1 + 1 equal to", the second intention type is the calculation type.

[0085] 304. The computer device forms semantic information by using the second intention type and the problem type through the large language model.

[0086] In the embodiments of the present application, the semantic information includes a second intention type and a question type. When the conversation information is question information, the large language model is used to process the conversation information to identify the question type of the question information and the second intention type matching the question type, and the second intention type and the question type are used to form the semantic information, so that the semantic information can indicate the type of the question information and the related intention type, enriching the content of the semantic information and ensuring the accuracy of the semantics indicated by the semantic information for the question information, and further ensuring the accuracy of the semantic information, so as to be able to determine the accurate intention type based on the semantic information subsequently.

[0087] In a possible implementation manner, step 304 includes: identifying, by the large language model, at least one of the theme of the question information, the entity words in the question information, or the word type of the entity words; and forming, by the large language model, the semantic information with at least one of the theme, the entity words, or the word type, the second intention type, and the question type.

[0088] In the embodiments of the present application, the semantic information includes a second intention type and question information, and also includes at least one of a theme, entity words, or a word type. The large language model is used to perform feature characterization on the conversation information in different dimensions to output an intention type, a theme, a question type, entity words, or an entity word type to form the semantic information, enriching the content included in the semantic information, so that the semantic information can more detailedly point out the semantics of the conversation information and ensuring the accuracy of the semantic information.

[0089] For example, if the question information is "What's the weather like today", in the semantic information of this question information, the second intention type is "weather type", the question type of the question information is "query type", the theme of the question information is "weather query", the entity words in the question information are "today" and "weather", the word type of "today" is "time", and the word type of "weather" is "concept".

[0090] For another example, if the question information is "Draw an apple", in the semantic information of this question information, the second intention type is "image generation from text type", the question type of the question information is "creation type", the theme of the question information is "drawing", the entity word in the question information is "apple", and the word type of "apple" is "fruit".

[0091] It should be noted that in the embodiments of the present application, taking the identification of the second intention type and the question type by the large language model as an example for illustration, in another embodiment, the large language model can also adopt the method of feature extraction and decoding to output semantic information, and the output semantic information includes the second intention type, the question type, etc.

[0092] In a possible implementation, the semantic information includes multiple characters. The process of obtaining the semantic information of the dialogue information includes: through a large language model, extracting features for each character in the dialogue information to obtain dialogue features, where the dialogue features include the features of each character in the dialogue information; for the first character in the dialogue information, updating the features of the first character based on the features of the first character and the features of the characters before the first character in the dialogue information to obtain the updated features of the first character, where the first character is any character in the dialogue information; forming the updated dialogue features with the updated features of the multiple characters in the dialogue information; through the large language model, performing the first decoding on the updated dialogue features to obtain the first character; through the large language model, based on the currently obtained character, performing the i-th decoding on the updated dialogue features to obtain the i-th character, where i is an integer greater than 0. When the number of decoding times reaches the threshold number of times, or when the currently decoded termination character is obtained, decoding is no longer performed, and the obtained characters are formed into semantic information.

[0093] In the embodiments of the present application, by updating the features of each character based on the features of multiple characters in the semantic information and adopting a decoding method based on the updated semantic features, multiple characters are gradually decoded, and then the semantic information is obtained, ensuring the accuracy of the obtained semantic information.

[0094] Optionally, multiple candidate characters are configured in the large language model. The decoding process includes: through the large language model, performing the first decoding on the updated dialogue features to obtain the first probabilities of the multiple candidate characters, and determining the candidate character with the largest first probability among the multiple candidate characters as the first character; through the large language model, based on the currently obtained character, performing the i-th decoding on the updated dialogue features to obtain the i-th probabilities of the multiple candidate characters, and determining the candidate character with the largest i-th probability among the multiple candidate characters as the i-th character, where i is an integer greater than 0; through the large language model, based on the currently obtained character, performing the j-th decoding on the updated dialogue features to obtain the j-th probabilities of the multiple candidate characters. When the largest probability among the j-th probabilities of the multiple candidate characters is less than the probability threshold, the decoding process is no longer executed, and the currently obtained j - 1 characters are formed into semantic information, where j is an integer greater than i; or, when the candidate character with the largest probability among the j-th probabilities of the multiple candidate characters is a termination character, the decoding process is no longer executed, and the currently obtained j - 1 characters are formed into semantic information; or, when the number of decoding times j is equal to the threshold number of times, the decoding process is no longer executed, and the currently obtained j characters are formed into semantic information.

[0095] Among them, the first probability of a candidate character refers to the possibility of the candidate character being the first character in the semantic information, and the i-th probability of a candidate character refers to the possibility of the candidate character being the i-th character in the semantic information.

[0096] In an embodiment of the present application, the large language model includes multiple candidate characters, and the multiple candidate characters include all characters as much as possible. In the process of generating semantic information of the dialogue information, the large language model selects characters from the multiple candidate characters to form semantic information. In the decoding process, one character is selected from the multiple candidate characters each time. By adopting this step-by-step decoding method, multiple characters can be decoded, and then the decoded multiple characters form semantic information, which can ensure the accuracy of the obtained semantic information. In the process of multiple decodings, in the current decoding process, if the maximum probability of the j-th probability of the multiple candidate characters obtained by decoding is less than the probability threshold, it means that the semantic information has been decoded, and there is no need to select characters from the multiple candidate characters as characters in the semantic information, thereby ensuring the accuracy of the semantic information.

[0097] 305. The computer device queries a type mapping table based on the semantic information, where the type mapping table includes a question type corresponding to each intent type in a plurality of intent types.

[0098] In the embodiment of the present application, a type mapping table stores multiple intent types and the question type corresponding to each intent type. The question type corresponding to the intent type indicates that the target dialogue model under the intent type can reply to the dialogue information belonging to the question type. When the semantic information of the dialogue information is obtained through the large language model, the type mapping table is queried in combination with the intent type and the question type in the semantic information to determine the real intent type of the dialogue information, so as to ensure the accuracy of the intent type finally identified.

[0099] In a possible implementation, the type mapping table stores at least one of the topic, entity word or word type and question type corresponding to each intent type. For example, the type mapping table includes the topic, entity word or word type and question type corresponding to each intent type.

[0100] 306. When the computer device finds that the question type corresponding to the second intent type in the type mapping table is the same as the question type in the semantic information, the computer device determines the second intent type as the first intent type.

[0101] In an embodiment of the present application, the semantic information includes a second intent type and a question type, and the type mapping table stores the question type corresponding to each intent type. Therefore, the question type corresponding to the second intent type in the query type table is the same as the semantic information, and then it is determined whether the target semantic model under the second intent type can reply to the dialogue message. When it is found that the question type corresponding to the second intent type in the type mapping table is the same as the question type in the semantic information, the second intent type is determined as the first intent type, so as to determine whether the dialogue message can be replied through the target semantic model under the first intent type.

[0102] In an embodiment of the present application, when obtaining the semantic information of the conversation information through a large language model, the type mapping table is queried in combination with the intent type and question type in the semantic information to determine the true intent type of the conversation information. Further verification is performed on the intent type matched by the conversation information to ensure the accuracy of the finally recognized intent type. Furthermore, it is ensured that the subsequent reply is made using a conversation model that matches the conversation requirements of the conversation information, ensuring the accuracy of the subsequent reply information.

[0103] In a possible implementation manner, the type mapping table further includes at least one of the topic, entity words, or word types corresponding to each intent type. Taking the type mapping table further including the topic, entity words, and word types corresponding to each intent type as an example, step 306 includes: when it is queried that the topic, entity words, or word types corresponding to the second intent type in the type mapping table and the question type are the same as the topic, entity words, or word types and the question type in the semantic information, the second intent type is determined as the first intent type.

[0104] For example, in the type mapping table, the topic corresponding to the "weather type" is "weather query". If the question information is "What's the weather like today", in the semantic information of this question information, the second intent type of the question information is "weather type" and the topic is "weather query", then it is determined that the "weather type" is the true intent type of the question information; if the question information is "Encyclopedic introduction of weather", in the semantic information of this question information, the second intent type of the question information is "weather type" and the topic is "meteorological knowledge", then it is determined that the "weather type" is not the true intent type of the question information.

[0105] It should be noted that in the embodiment of the present application, the first intent type is determined through the type mapping table. In another embodiment, the above steps 305-306 are not required to be executed, but other methods are adopted to determine the first intent type from multiple intent types based on the semantic information of the conversation information, and the first intent type matches the semantic information.

[0106] 307. The computer device processes the conversation information through the target conversation model under the first intent type to obtain the first reply information.

[0107] In the embodiment of the present application, since the first intent type is the true intent type to which the conversation information belongs, the conversation information is processed through the target conversation model under the first intent type to obtain the first reply information, so as to ensure that the obtained reply information matches the conversation information and ensure the accuracy of the obtained first reply information.

[0108] For example, the conversation information is "What's the weather like today". By classifying the conversation information, when it is determined that the category of the conversation information belongs to the first category, semantic analysis is performed on the conversation information through a large language model, and then posterior correction is performed using the obtained semantic information to determine that the true intention type of the conversation information is the "weather type". Through the target conversation model under the "weather type", the conversation information is processed, and the first reply information obtained is "The temperature today is 26 degrees and it is sunny."

[0109] 308. When the computer device does not match multiple intention types and semantic information, the conversation information is processed through a general conversation model to obtain the second reply information.

[0110] In the embodiment of the present application, considering that the category obtained by classifying the conversation information may be inaccurate, resulting in the semantic information not matching multiple intention types, in order to ensure that the conversation information can be replied, the conversation information is processed through a general conversation model, so that the second reply information can be timely based on for subsequent reply, avoiding the user waiting for too long, ensuring the real-time nature of the reply, and also avoiding inaccurate reply caused by replying through one of multiple target conversation models, thereby ensuring the accuracy of the reply information.

[0111] 309. When the category of the conversation information belongs to the second category, the conversation information is processed through a general conversation model to obtain the second reply information.

[0112] In the embodiment of the present application, the category of the conversation information being the second category means that the conversation information does not match multiple intention types, so it is no longer possible to reply through multiple target conversation models. Instead, the conversation information is replied through a general conversation model to ensure the accuracy of the generated second reply information.

[0113] The solution provided by the embodiments of the present application pre-sets a general dialogue model and multiple target dialogue models. The multiple target dialogue models and the general dialogue model belong to different categories. After obtaining the dialogue information, the dialogue information is first classified to identify which type of dialogue model to use for the reply. When it is determined that the category of the dialogue information is the first category, the semantic information of the dialogue information is analyzed through a large language model, so as to use the semantic information to determine the first intention type that matches the semantic information from multiple intention types. Then, the dialogue model under the first intention type is used to generate the corresponding reply information. In this way, a simple binary classification method is adopted to roughly screen the dialogue information to determine which type of dialogue model to use for the reply. When it is determined to use the dialogue models under multiple intention types for the reply, the large language model will be used for precise matching to determine which intention type the dialogue requirement of the dialogue information matches. Then, the dialogue model under the determined intention type is used for the reply. This can ensure that the used dialogue model better matches the dialogue requirement of the dialogue information, and thus ensure the accuracy of the reply information.

[0114] The embodiments of the present application propose an intention recognition method based on text general understanding posterior. In the large language model question and answer scenario, first, the dialogue information is roughly screened based on a binary classification model. For the dialogue information that does not belong to multiple intention types, the reply information is output through the general dialogue model. For the dialogue information that belongs to multiple intention types, the semantic information of the dialogue information is analyzed through the large language model, and then the true intention type of the dialogue information can be determined. Through the target dialogue model under the true intention type, the reply information of the dialogue information is output, thereby improving the efficiency and accuracy of the entire link of intention recognition and ensuring the accuracy of the reply information.

[0115] Based on the above Figure 3 On the basis of the shown embodiments, through the classification model, the large language model, multiple target dialogue models and the general dialogue model, human-machine intelligent dialogue can be realized. As Figure 5 shown, for any dialogue information input by any user, the dialogue information is classified through the classification model to obtain the category of the dialogue information. When the category of the dialogue information is the first category, the semantic analysis of the dialogue information is carried out through the large language model to obtain the semantic information of the dialogue information. Based on the semantic information, among multiple intention types such as the text-to-image intention type, the code intention type, the weather intention type, and the calculation intention type, it is determined that the dialogue information matches the code intention type; through the target dialogue model under the code intention type, the dialogue information is processed to obtain the first reply information. When the category of the dialogue information is the second category, the dialogue information is processed through the general dialogue model to obtain the second reply information.

[0116] It should be noted that the above Figure 3The illustrated embodiment takes the semantic information including the second intention type and the question type as an example for illustration. In another embodiment, instead of performing the above steps 302-304, other methods are adopted. When the category of the conversation information is the first category, the large language model is used to perform semantic analysis on the conversation information to obtain the semantic information of the conversation information.

[0117] In a possible implementation manner, the large language model combines the first indication information to generate semantic information. That is, the process of generating semantic information includes: obtaining the first indication information, where the first indication information instructs the large language model to perform semantic analysis on the input information according to the examples of semantic analysis, and the examples include input information examples and semantic information examples of the input information examples; when the category of the conversation information is the first category, the large language model is used to perform semantic analysis on the conversation information based on the first indication information to obtain the semantic information of the conversation information.

[0118] In the embodiment of the present application, the first indication information indicates the examples of semantic information and instructs the large language model to perform semantic analysis on the input information according to the examples of semantic information. Since the large language model has powerful reasoning ability, through the large language model, according to the first indication information, it can learn the examples of semantic analysis in the first indication information to perform semantic analysis on the conversation information according to the examples of semantic analysis, ensuring the accuracy of the semantic information.

[0119] Among them, the first indication information can be represented in any form. For example, the first indication information is text. Both the input information example and the semantic information example can be represented in any form. For example, both the input information example and the semantic information example are text.

[0120] Optionally, the examples of semantic analysis in the first indication information include positive examples and negative examples. The positive examples include input information examples and semantic information examples, and the negative examples include input information examples and semantic information examples. The intention type included in the semantic information example in the positive examples indicates the general dialogue model.

[0121] For example, taking the input information of a large language model as text, the first indication information is: based on the instruction, complete the following text understanding tasks: identify the intention type of the question, identify the theme of the question, identify the question type of the question, identify the entity words and word types of the question, where the intention types include text-to-image type, code type, calculation type, weather type, calendar type, acrostic type, map type, website type, picture description type, translation type, etc. Example of semantic analysis: "Input: What's the weather like today and what's the temperature? Output: Intention type: [weather type], theme: [weather query], question type: [query type], entity words and word types: [today: time | weather: concept]"; "Input: Write a piece of code to traverse a folder. Output: Intention type: [code type], theme: [programming], type: [programming type], entity: []".

[0122] Based on the above Figures 2 to 3 On the basis of the embodiments shown above, before classifying the dialogue information through a classification model, the classification model is also trained. The process of training the classification model includes: obtaining a plurality of sample dialogue information and the sample category corresponding to each sample dialogue information; classifying each sample dialogue information through the classification model to obtain the predicted category of each sample dialogue information; training the classification model based on the sample category and predicted category of the plurality of sample dialogue information.

[0123] In the embodiments of the present application, the classification model is a binary classification model for classifying input information to determine whether the input information belongs to the first category or the second category. The sample category corresponding to the sample dialogue information is the first category or the second category. By classifying the sample dialogue information through the classification model, the predicted category of the sample dialogue information is obtained. The difference between the predicted category and the sample category can reflect the accuracy of the classification model. The classification model is trained through the predicted type and the sample category to improve the accuracy of the classification model.

[0124] Among them, the sample dialogue information can be any dialogue information. For example, the sample dialogue information is "What's the weather like today and what's the temperature?", or "Draw an apple", etc.

[0125] Optionally, the sample dialogue information includes positive sample dialogue information or negative sample dialogue information. The sample category corresponding to the positive sample dialogue information is the first category, and the sample category corresponding to the negative sample dialogue information is the second category.

[0126] In the embodiments of the present application, the plurality of sample dialogue information includes positive sample dialogue information and negative sample dialogue information. The ratio of positive sample dialogue information to negative sample dialogue information in the plurality of sample dialogue information can be any ratio. For example, the ratio of positive sample dialogue information to negative sample dialogue information is 6:4, such as Figure 6As shown, among multiple sample conversation messages, 60% of the sample conversation messages are positive sample conversation messages, that is, the category of 60% of the sample conversation messages is the first category, and 40% of the sample conversation messages are negative sample conversation messages, that is, the category of 40% of the sample conversation messages is the second category. Since the data of the sample conversation messages and the sample categories are simple, large-scale data can be generated quickly and efficiently. The classification model only needs to identify whether the input message belongs to the first category or the second category, without identifying which intent type among the multiple intent types corresponding to the first category the input message belongs to, making the structure of the classification model simple and enabling the classification model to be trained quickly.

[0127] Optionally, based on the predicted category and the sample category of each sample conversation message, a loss value is determined, and based on the loss value, the classification model is trained, where the loss value represents the difference between the predicted category and the sample category of the sample conversation message. In the embodiments of the present application, cross-entropy loss is adopted to train the classification model to improve the accuracy of the classification model.

[0128] In the above Figures 2 to 3 Based on the embodiment shown above, before generating the semantic information of the conversation message through the large language model, the large language model is also trained, and the training process is executed by the computer device. The training process includes:

[0129] Step 1: The computer device obtains the sample conversation message and the second indication information, and the second indication information instructs the semantic analysis model to perform semantic analysis on the input message according to the example of semantic analysis.

[0130] In the embodiments of the present application, the second indication information instructs the semantic analysis model to be able to complete the semantic analysis task according to the instruction and indicates the example of semantic analysis, so that the subsequent semantic analysis model can complete the semantic analysis task according to the indication information.

[0131] Among them, the sample conversation message is any conversation message. For example, the sample conversation message is "What's the weather like today and what's the temperature", or "Draw an apple", etc. The semantic analysis model is a large language model that has been trained and completed.

[0132] Step 2: The computer device performs semantic analysis on the sample conversation message through the semantic analysis model based on the second indication information to obtain the sample semantic information.

[0133] In the embodiment of the present application, the semantic analysis model has a powerful reasoning function and can process the input information based on the input indication information to achieve the task indicated by the indication information. If the second indication information indicates that the semantic analysis model completes the semantic analysis task, the semantic analysis model can perform semantic analysis on the sample dialogue information according to the second indication information, so as to obtain the sample semantic information of the sample dialogue information according to the example of semantic analysis in the second indication information.

[0134] In the embodiment of the present application, the semantic analysis model can perform tasks based on the indication information, and the semantic analysis model is a trained model, so the sample semantic information obtained through the semantic analysis model is accurate enough.

[0135] Step 3: The computer device processes the sample dialogue information through the large language model to obtain the predicted semantic information.

[0136] This step 3 is the same as the above step 202 and will not be elaborated here.

[0137] Step 4: The computer device trains the large language model based on the predicted semantic information and the sample semantic information.

[0138] In the embodiment of the present application, the difference between the predicted semantic information and the sample semantic information can reflect the accuracy of the large language model. The smaller the difference between the predicted semantic information and the sample semantic information, the more accurate the large language model is. The larger the difference between the predicted semantic information and the sample semantic information, the less accurate the large language model is. Therefore, training the large language model based on the predicted semantic information and the sample semantic information can improve the accuracy of the large language model.

[0139] In the solution provided by the embodiment of the present application, the semantic analysis model is a trained large language model with a powerful reasoning function. When the sample dialogue information is obtained, the sample semantic information of the sample dialogue information can be obtained by using the semantic analysis model, which ensures the accuracy of the sample semantic information. Then, based on the sample dialogue information and the sample semantic information, the predicted semantic information of the sample dialogue information is predicted through the large language model. Training the large language model based on the difference between the predicted semantic information and the sample semantic information can ensure the training effect of the large language model and improve the accuracy of the large language model.

[0140] It should be noted that in the above embodiment, the large language model is trained based on the predicted semantic information and the sample semantic information output by the large language model. In another embodiment, the sample dialogue information is processed through the large language model to obtain the predicted semantic information that is the same as the sample semantic information and the probability of each character in the predicted semantic information, and the large language model is trained based on the probability of each character in the predicted semantic information.

[0141] In the embodiments of the present application, when the large language model processes the sample dialogue information, it will output multiple characters in a character-by-character decoding manner and form the predicted semantic information with the output characters. Since the sample semantic information is the true semantic information of the sample dialogue information, when the large language model processes the sample dialogue information, according to the sample semantic information, the large language model is controlled to output the predicted semantic information identical to the sample semantic information, and the probability of each character in the predicted semantic information is determined. Then, the probability of each character in the predicted semantic information can reflect the accuracy of the large language model. The greater the probability of each character in the predicted semantic information, the greater the possibility that the large language model outputs the predicted semantic information identical to the sample semantic information. The smaller the probability of each character in the predicted semantic information, the smaller the possibility that the large language model outputs the predicted semantic information identical to the sample semantic information. Therefore, based on the probability of each character in the predicted semantic information, the large language model is trained to improve the accuracy of the large language model.

[0142] Optionally, the process of obtaining the predicted semantic information identical to the sample semantic information and the probability of each character in the predicted semantic information through the large language model includes: extracting features of each character in the sample dialogue information through the large language model to obtain sample dialogue features, where the sample dialogue features include the features of each character in the sample dialogue information; for the second character in the sample dialogue information, updating the features of the second character based on the features of the second character and the features of the characters before the second character in the dialogue information to obtain the updated features of the second character, where the second character is any character in the sample dialogue information; forming the updated sample dialogue features with the updated features of multiple characters in the sample dialogue information; decoding the updated sample dialogue features for the first time through the large language model to obtain the first probability of multiple alternative characters in the large language model and determining the first probability of the first character in the sample semantic information; decoding the updated sample dialogue features for the second time through the large language model based on the first character in the sample semantic information to obtain the second probability of multiple alternative characters and determining the second probability of the second character in the sample semantic information; decoding the updated sample dialogue features for the (k + 1)-th time through the large language model based on the first k characters in the sample semantic information to obtain the (k + 1)-th probability of multiple alternative characters and determining the (k + 1)-th probability of the (k + 1)-th character in the sample semantic information; repeating the above process until the probability of the last character in the sample semantic information is determined. At this time, the predicted semantic information identical to the sample semantic information and the probability of each character in the predicted semantic information are obtained. Wherein, k is an integer greater than 1.

[0143] In a possible implementation, if the predicted semantic information is the same as the sample semantic information, the process of training the large language model includes: determining a loss value based on the probability of each character in the predicted semantic information, and training the large language model based on the loss value.

[0144] Optionally, the large language model minimizes the maximum likelihood function to determine the loss value, and the loss value satisfies the following relationship:

[0145]

[0146] Where L1(u) is used to represent the loss value, Θ is used to represent the sample semantic information, i is used to represent the serial number of the character in the sample semantic information, i is an integer greater than 1, u i-1 represents the (i - 1)-th character in the sample semantic information, u i-k represents the (i - k)-th character in the sample semantic information, u i represents the i-th character in the sample semantic information, k is an integer greater than 0, k represents the window size, P(u i |u i-k , …, u i-1 ; Θ) is used to represent the probability of obtaining the i-th character in the sample semantic information when the first k characters of the predicted semantic information are the same as those of the sample semantic information, that is, the probability of using the first k characters in the sample semantic information to predict the i-th character.

[0147] In the embodiments of the present application, the large language model is obtained by SFT based on the pre-trained GPT in an autoregressive manner, and GPT is constructed using the Transformer decoder module. When the Transformer decoder processes the dialogue information, it outputs multiple characters in a step-by-step decoding manner, and the semantic information is composed of the output multiple characters. During the process of outputting multiple characters, the next character is output based on the dialogue information and the currently obtained characters. During the process of training the Transformer decoder, the sample dialogue information is processed, and the predicted semantic information identical to the sample semantic information is output character by character according to the sample semantic information. During the process of outputting the predicted semantic information, the next character is output based on the currently obtained character, that is, the next Token is predicted by the current Token (representation) and the previous Tokens, and the characters after the currently obtained character in the sample semantic information are masked (hidden).

[0148] The Transformer decoder can more efficiently capture the long-range dependencies in sequence data. The Transformer decoder consists of multiple self-attention (Masked Self-attention) layers and position-wise feed-forward neural networks, and is stacked together through residual connections and layer normalization. Self-Attention captures context-related information in the sequence through the self-attention mechanism. The calculation of self-attention involves three weight matrices (query matrix Q, key matrix K, and value matrix V), and the final attention weights are calculated through dot product, scaling, Softmax activation, and weighted summation. Masked Self-attention uses a mask in the self-attention mechanism to mask the information after the current token, ensuring that the prediction is only based on the information of previous tokens. Layer Normalization is used to accelerate the convergence of the model. After the output of each layer, layer normalization is used to normalize it, alleviating the problem of gradient vanishing / explosion in the network.

[0149] The parameters of the Transformer decoder can be arbitrary parameters. For example, the parameters of the Transformer decoder are set as follows: MODEL_SIZE (model scale) is 7B (7 Billion, 7 billion), NUM_LAYERS (number of network layers) is 32, HIDDEN_SIZE (hidden layer size) is 4096, NUM_ATTN_HEADS (number of multi-head self-attention heads) is 32, FFN_HIDDEN_SIZE (hidden layer size of the feed-forward neural network) is 16384, and ATTN_HEAD_SIZE (self-attention layer size) is 128. It should be noted that large language models can also adopt the GPT model with larger-scale parameters for SFT.

[0150] In the above embodiments, the training data of the large language model includes sample dialogue information and sample semantic information. When constructing the training data of the large language model, based on the ICL (In Context Learning, analogical learning) method, sample dialogue information is constructed, and the sample semantic information of the sample dialogue information is obtained through a semantic analysis model. Furthermore, the sample dialogue information and the sample semantic information can form the training data of the large language model. The training data constructed in this way can cover a variety of NLP tasks, including intent types, topics, question types, entity recognition, etc.

[0151] In the process of constructing the training data for the large language model, the second indication information of the semantic analysis model is first constructed, and the second indication information indicates the tasks that the semantic analysis model needs to perform and examples of semantic analysis.

[0152] The tasks that the semantic analysis model needs to perform indicate that the semantic analysis model completes the following text understanding tasks based on the instructions: identifying the intent type to which the question belongs, identifying the theme to which the question belongs, identifying the question type to which the question belongs, and identifying the entity words and word types of the question. The intent types include text-to-image type, code type, calculation type, weather type, calendar type, acrostic type, map type, website type, describe-image type, translation type, etc. Examples of semantic analysis: "Input: What's the weather like today and what's the temperature? Output: Intent type: [Weather type], Theme: [Weather query], Question type: [Query type], Entity words and word types: [Today: Time | Weather: Concept]"; "Input: Write a piece of code to traverse a folder. Output: Intent type: [Code type], Theme: [Programming], Type: [Programming type], Entity: []".

[0153] The examples of semantic analysis in the second indication information indicate the input and output of the semantic analysis model. For example, the input of the semantic analysis model is "What's the weather like today and what's the temperature?", and the output of the semantic analysis model is ""Plugin" belongs to "Weather plugin", "Theme" belongs to "Weather query", "Type" belongs to "Query class"; "Entity" is stored in the form of "Entity fragment: Entity type", and multiple entities are separated by "|", including two, namely "Today: Time" and "Weather: Concept". The examples of semantic analysis in the second indication information include positive and negative examples of multiple intent types. For example, taking the "weather intent type" as an example, the positive example is ("What's the weather like today and what's the temperature?") and an easily confused negative example ("Can I go to the camping concert on Friday?"), and the same applies to other intent types. The semantic information output by the semantic analysis model includes intent type, theme, question type, entity words and word types. This application embodiment is only illustrated by taking the semantic information including intent type, theme, question type, entity words and word types as an example. In various NLP understanding tasks, the content of the semantic information can be flexibly increased or decreased.

[0154] For example, the examples of semantic analysis in the second indication information are as follows:

[0155] "Input: What's the weather like today and what's the temperature? Output: Intent type: [Weather type], Topic: [Weather query], Question type: [Query type], Entities: [Today: Time|Weather: Concept]"; "Input: Can I go to the camping concert on Friday?; Output: Intent type: [General dialogue model], Topic: [Event invitation], Question type: [Consultation type], Entities: [Friday: Time|Camping concert: Event]"; "Input: Write a code to traverse a folder; Output: Intent type: [Code type], Topic: [Programming], Question type: [Programming type], Entities: []"; "Input: What's the difference between a linear regression function and a coding function? Output: Intent type: [General dialogue model], Topic: [Programming], Question type: [Introduction type], Entities: []"; "Input: Draw an apple; Output: Intent type: [Text-to-image type], Topic: [Painting], Question type: [Creation type], Entities: [Apple: Fruit]"; "Input: Describe a picture: An animal is sleeping in a wardrobe and blowing a fan; Output: Intent type: [General dialogue model], Topic: [Picture description], Question type: [Description type], Entities: [Animal: Anime image|Wardrobe: Item|Fan: Item]".

[0156] Based on the above second indication information, taking the sample dialogue information "Describe a picture: Students are playing basketball" as an example, through the semantic analysis model, based on the second indication information, the sample dialogue information is processed, and the obtained sample semantic information is "Intent type: [General dialogue model], Topic: [Picture description], Question type: [Description type], Entities: [Students: People|Playing basketball: Activity]".

[0157] After obtaining the training data through the semantic analysis model above, the training data can be screened, and the large language model can be trained using the screened training data. In the embodiment of the present application, through the sample information template provided by the embodiment of the present application, the sample dialogue information and the corresponding sample semantic information are formed into a piece of training data, and the sample information template and the training data are as follows: "Input: {Dialogue information} Output: {Reply information}". For example, the training data is "Input: What is 1.1 to the power of 7? Output: Intent type: [Calculation type], Topic: [Mathematical problem], Question type: [Calculation type], Entities: [Power: Mathematical concept]".

[0158] In a possible implementation manner, a template for the large language model is also constructed. The template includes the input and output of the large language model. Taking the input information of the large language model as the text, the large language model processes the input information to obtain the reply information. The dialogue information and the reply information can form the following template: "Input: {Dialogue information} Output: {Reply information}".

[0159] Through the method provided by the embodiments of the present application, the accuracy of the reply information can be guaranteed. In the solution provided by the embodiments of the present application, the dialogue model includes a general dialogue model and multiple target dialogue models. The training data of the general dialogue model and multiple target dialogue models is simply constructed, and the training data of each dialogue model can be quickly generated. Then, based on the training data of each dialogue model, each dialogue model is trained respectively, and the large language model is used to process the intention type to which the dialogue information belongs. Only a small amount of high-quality training data is required to train the large language model. Since the classification model has low requirements for configuration and video memory, and the large language model has high requirements for configuration and video memory, the binary classification model is used to identify the category of the dialogue information, and only the dialogue information belonging to the first category will be distributed to the large language model for intention recognition, without processing each dialogue information through the large language model, which can save the resources of the device.

[0160] Figure 7 is a schematic structural diagram of a reply information generation device provided by an embodiment of the present application, as Figure 7 shown, the device includes:

[0161] A classification module 701, configured to classify dialogue information to obtain the category of the dialogue information, where the category includes a first category or a second category. The first category indicates multiple target dialogue models, and each target dialogue model is used to reply to dialogue information of one intention type. The second category indicates a general dialogue model;

[0162] An analysis module 702, configured to, when the category of the dialogue information is the first category, perform semantic analysis on the dialogue information through a large language model to obtain the semantic information of the dialogue information;

[0163] A determination module 703, configured to determine a first intention type from multiple intention types based on the semantic information, where the first intention type matches the semantic information;

[0164] A processing module 704, configured to process the dialogue information through the target dialogue model under the first intention type to obtain a first reply information.

[0165] In a possible implementation manner, the dialogue information is question information; the analysis module 702 is configured to, when the category of the dialogue information is the first category, identify the question type of the question information through a large language model; classify the question information through a large language model to obtain a second intention type, where the second intention type is the intention type that matches the question type among multiple intention types; and form semantic information by using the second intention type and the question type through a large language model.

[0166] In another possible implementation manner, as Figure 8 shown, the device further includes:

[0167] An identification module 705, configured to identify at least one of the topic of the question information, the entity words in the question information, or the word types of the entity words through a large language model;

[0168] An analysis module 702, configured to form semantic information through a large language model with at least one of the topic, entity words, or word types, as well as the second intention type and the question type.

[0169] In another possible implementation manner, a determination module 703 is configured to query a type mapping table based on the semantic information. The type mapping table includes the question types corresponding to each intention type among multiple intention types. When it is queried that the question type corresponding to the second intention type in the type mapping table is the same as the question type in the semantic information, the second intention type is determined as the first intention type.

[0170] In another possible implementation manner, as Figure 8 shown, the apparatus further includes:

[0171] An acquisition module 706, configured to acquire first indication information, where the first indication information instructs the large language model to perform semantic analysis on the input information according to an example of semantic analysis. The example includes an input information example and a semantic information example of the input information example;

[0172] The analysis module 702 is configured to, when the category of the conversation information is the first category, perform semantic analysis on the conversation information through the large language model based on the first indication information to obtain the semantic information of the conversation information.

[0173] In another possible implementation manner, a processing module 704 is further configured to, when the category of the conversation information is the second category, process the conversation information through a general conversation model to obtain a second reply message.

[0174] In another possible implementation manner, the processing module 704 is further configured to, when none of the multiple intention types match the semantic information, process the conversation information through a general conversation model to obtain a second reply message.

[0175] In another possible implementation manner, as Figure 8 shown, the apparatus further includes:

[0176] An acquisition module 706, configured to acquire sample conversation information and second indication information, where the second indication information instructs the semantic analysis model to perform semantic analysis on the input information according to an example of semantic analysis;

[0177] The analysis module 702 is further configured to perform semantic analysis on the sample conversation information through the semantic analysis model based on the second indication information to obtain sample semantic information;

[0178] The processing module 704 is further configured to process the sample dialogue information through a large language model to obtain predicted semantic information;

[0179] The training module 707 is configured to train the large language model based on the predicted semantic information and the sample semantic information.

[0180] It should be noted that: for the reply information generation device provided in the above embodiment, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the reply information generation device provided in the above embodiment and the embodiment of the reply information generation method belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.

[0181] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the reply information generation method in the above embodiment.

[0182] Optionally, the computer device is provided as a terminal. Figure 9 The structural block diagram of a terminal 900 provided by an exemplary embodiment of the present application is shown. The terminal 900 includes a processor 901 and a memory 902.

[0183] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0184] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 901 to implement the reply information generation method provided in the method embodiments of the present application.

[0185] In some embodiments, the terminal 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, and a power supply 908.

[0186] The peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0187] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0188] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905. The touch signal can be input to the processor 901 as a control signal for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, which is disposed on the front panel of the terminal 900; in other embodiments, there may be at least two display screens 905, which are respectively disposed on different surfaces of the terminal 900 or are in a foldable design; in other embodiments, the display screen 905 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the terminal 900. Even, the display screen 905 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 905 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0189] The camera module 906 is used to collect images or videos. Optionally, the camera module 906 includes a front camera and a rear camera. The front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 906 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0190] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 901 for processing, or input to the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 907 may further include a headphone jack.

[0191] The power supply 908 is used to supply power to each component in the terminal 900. The power supply 908 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 908 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0192] Those skilled in the art can understand that Figure 9 the structure shown in

[0193] does not limit the terminal 900, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout. Figure 10

[0194] Optionally, the computer device is provided as a server.

[0195] The embodiment of the present application also provides a computer program product, including a computer program, and the operations performed by the computer program when executed by a processor implement the method for generating reply information in the above embodiment.

[0196] Those of ordinary skill in the art can understand that all or part of the steps in the above embodiment can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0197] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.

Claims

1. A method for generating reply information, characterized in that, The method includes: Classify the problem information to obtain the category of the problem information. The category includes a first category or a second category. The first category indicates multiple target dialogue models, and each target dialogue model is used to reply to problem information of a certain intention type. The second category indicates a general dialogue model; When the category of the problem information is the first category, identify the problem type of the problem information through a large language model; Classify the problem information through the large language model to obtain a second intention type, where the second intention type is the intention type that matches the problem type among multiple intention types; Identify at least one of the theme of the problem information, the entity words in the problem information, or the word type of the entity words through the large language model; Through the large language model, construct the semantic information of the problem information from at least one of the theme, the entity words, or the word type, the second intention type, and the problem type; Based on the semantic information, query a type mapping table. The type mapping table includes at least one of the theme, entity words, or word type corresponding to each intention type among the multiple intention types and the problem type corresponding to each intention type. The problem type corresponding to the intention type indicates that the target dialogue model under the intention type can reply to problem information belonging to the problem type; When it is queried that the theme, entity words, or word type and problem type corresponding to the second intention type in the type mapping table are the same as the theme, entity words, or word type and problem type in the semantic information, determine the second intention type as the first intention type; Process the problem information through the target dialogue model under the first intention type to obtain a first reply message.

2. The method according to claim 1, wherein The method further includes: Obtain first indication information, where the first indication information instructs the large language model to perform semantic analysis on the input information according to an example of semantic analysis. The example includes an input information example and a semantic information example of the input information example; When the category of the problem information is the first category, perform semantic analysis on the problem information through the large language model based on the first indication information to obtain the semantic information of the problem information.

3. The method according to claim 1, characterized in that, After classifying the problem information to obtain the category of the problem information, the method further includes: When the category of the problem information is the second category, process the problem information through the general dialogue model to obtain a second reply message.

4. The method according to claim 1, wherein After querying the type mapping table based on the semantic information, the method further includes: When none of the multiple intention types match the semantic information, process the problem information through the general dialogue model to obtain a second reply message.

5. The method according to claim 1, wherein The method further includes: Obtain sample dialogue information and second indication information, where the second indication information instructs a semantic analysis model to perform semantic analysis on the input information according to an example of semantic analysis; Based on the second indication information, perform semantic analysis on the sample dialogue information through the semantic analysis model to obtain sample semantic information; Process the sample dialogue information through the large language model to obtain predicted semantic information; Train the large language model based on the predicted semantic information and the sample semantic information.

6. A reply information generation device, characterized in that The device includes: A classification module, configured to classify the question information to obtain the category of the question information. The category includes a first category or a second category. The first category indicates multiple target dialogue models, and each target dialogue model is used to reply to question information of a specific intention type. The second category indicates a general dialogue model; An analysis module, configured to, when the category of the question information is the first category, identify the question type of the question information through the large language model; classify the question information through the large language model to obtain a second intention type, and the second intention type is the intention type that matches the question type among multiple intention types; An identification module, configured to identify at least one of the topic of the question information, the entity words in the question information, or the word type of the entity words through the large language model; The analysis module is further configured to, through the large language model, form the semantic information of the question information by using at least one of the topic, the entity words, or the word type, as well as the second intention type and the question type; A determination module, configured to query a type mapping table based on the semantic information. The type mapping table includes at least one of the topic, entity words, or word type corresponding to each intention type among the multiple intention types, and the question type corresponding to each intention type. The question type corresponding to the intention type indicates that the target dialogue model under the intention type can reply to question information belonging to the question type. When it is queried that the topic, entity words, or word type and the question type corresponding to the second intention type in the type mapping table are the same as those in the semantic information, determine the second intention type as the first intention type; A processing module, configured to process the question information through the target dialogue model under the first intention type to obtain a first reply information.

7. The device according to claim 6, characterized in that, The device further includes: An acquisition module, configured to acquire first indication information, where the first indication information instructs the large language model to perform semantic analysis on the input information according to an example of semantic analysis. The example includes an input information example and a semantic information example of the input information example; The analysis module is further configured to, when the category of the question information is the first category, perform semantic analysis on the question information through the large language model based on the first indication information to obtain the semantic information of the question information.

8. The device according to claim 6, characterized in that, The processing module is further configured to, when the category of the question information is the second category, process the question information through the general dialogue model to obtain a second reply information.

9. The device according to claim 6, wherein The processing module is further configured to, when the multiple intent types do not match the semantic information, process the question information through the general dialogue model to obtain a second reply message.

10. The device according to claim 6, wherein The device further includes: an acquisition module, configured to acquire sample dialogue information and a second indication information, where the second indication information instructs the semantic analysis model to perform semantic analysis on the input information according to an example of semantic analysis; the analysis module is further configured to perform semantic analysis on the sample dialogue information through the semantic analysis model based on the second indication information to obtain sample semantic information; the processing module is further configured to process the sample dialogue information through the large language model to obtain predicted semantic information; a training module, configured to train the large language model based on the predicted semantic information and the sample semantic information.

11. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the reply message generation method according to any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the reply message generation method according to any one of claims 1 to 5.

13. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the operations performed by the reply message generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Visual question and answer processing method and device, computer readable medium and program product

    CN113722458A

  • Intention recognition method, man-machine interaction method, electronic equipment and storage medium

    CN115658724A