Message processing method and device, computer equipment and storage medium

Through the big model, process multiple rounds of conversation messages, generate statements and identify user intentions, the problem that traditional intelligent customer service cannot accurately identify user intentions is solved, and a more accurate reply and a more natural conversation experience is achieved.

CN120296113APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410028268.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When traditional intelligent customer service handles multiple user descriptions and supplements, it cannot accurately identify the user's true intentions, resulting in inaccurate reply.

Method used

By obtaining multiple rounds of session messages, using a big model for statement generation and intent recognition, combining historical and current session messages, locking user intent and providing accurate replies.

Benefits of technology

Improve the accuracy of intention recognition, make the interactive process closer to the real-person dialogue experience, and achieve a smooth dialogue process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296113A_ABST
    Figure CN120296113A_ABST
Patent Text Reader

Abstract

The invention relates to a message processing method and device, computer equipment, a storage medium and a computer program product. The invention relates to man-machine conversation, and the method comprises the steps: obtaining multiple rounds of conversation messages from a user terminal in an interaction process, the multiple rounds of conversation messages comprising a current conversation message and historical conversation messages before the current conversation message; jointly inputting the current session message and the historical session message into a large model, and outputting a generated statement through the large model; performing intention recognition based on the generated statement to obtain a first intention corresponding to the current session message; for any historical session message, acquiring a second intention corresponding to the historical session message; and under the condition that the first intention is successfully matched with any second intention, determining reply content based on the first intention and the current session message, and feeding back the reply content. By adopting the method, the accurate reply content can be fed back to the user terminal by accurately positioning the intention of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a message processing method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] An intelligent customer service is a robotic customer service that processes user problems. During the interaction and conversation between a user and the intelligent customer service, due to the user's understanding of the problem or communication habits, etc., the user often describes and supplements the same problem multiple times. During the interaction with the user, traditional intelligent customer services usually identify the user's problem through NLP (Natural Language Processing) technology to determine the user's intention, and then reply according to the determined user intention.

[0003] However, NLP technology generally identifies based on complete single sentences, and problems such as fragmented descriptions or lack of key information often occur in the user's problem descriptions, making it impossible to focus on the user's true intention, and thus impossible to provide an accurate problem solution for the user. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a message processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can accurately locate the user's intention and thus feedback accurate reply content to the user terminal.

[0005] In a first aspect, the present application provides a message processing method. The method includes:

[0006] Obtain multi-round session messages from a user terminal during an interaction, where the multi-round session messages include the current session message and historical session messages before the current session message;

[0007] Input the current session message and the historical session messages into a large model together, and generate a sentence through the large model;

[0008] Perform intention recognition based on the generated sentence to obtain a first intention corresponding to the current session message;

[0009] For any historical session message, obtain a second intention corresponding to the targeted historical session message;

[0010] In the case where the first intention matches any of the second intentions successfully, determine the reply content based on the first intention and the current session message, and feedback the reply content.

[0011] In a second aspect, the present application further provides a message processing apparatus. The apparatus includes:

[0012] An acquisition module, configured to acquire multi-round session messages from a user terminal during an interaction process, where the multi-round session messages include the current session message and historical session messages before the current session message;

[0013] An output module, configured to input the current session message and the historical session messages into a large model together, and generate a sentence through the large model;

[0014] An identification module, configured to perform intent identification based on the generated sentence to obtain a first intent corresponding to the current session message;

[0015] The identification module is further configured to, for any historical session message, obtain a second intent corresponding to the targeted historical session message;

[0016] A feedback module, configured to, when the first intent matches any of the second intents successfully, determine a reply content based on the first intent and the current session message, and feedback the reply content.

[0017] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above message processing method are implemented.

[0018] In a fourth aspect, the present application further provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the steps of the above message processing method are implemented.

[0019] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the above message processing method are implemented.

[0020] The above message processing method, apparatus, computer device, storage medium, and computer program product obtain multi-round session messages from a user terminal during an interaction, where the multi-round session messages include the current session message and historical session messages before the current session message. The current session message and the historical session messages are jointly input into a large model, and then a standardized generated statement can be output based on the large model. Further, intent recognition is performed on the standardized generated statement, which can reduce the interference caused by similar, redundant, and colloquial content in the interaction to intent recognition and improve the accuracy of intent recognition. Furthermore, for any historical session message, a second intent corresponding to the targeted historical session message is obtained, and when the first intent matches any second intent successfully, the user intent is locked. Based on this, the response content is determined based on the first intent and the current session message, and an accurate response content is fed back to the user terminal, making the interaction process closer to a real-person conversation experience and achieving a smooth conversation flow. Description of the Drawings

[0021] Figure 1 It is an application environment diagram of the message processing method in an embodiment;

[0022] Figure 2 It is a flowchart of the message processing method in an embodiment;

[0023] Figure 3 It is a flowchart block diagram of the message processing method in an embodiment;

[0024] Figure 4 It is a flowchart block diagram of determining a response vector in an embodiment;

[0025] Figure 5 It is a flowchart block diagram of the message processing method in another embodiment;

[0026] Figure 6 It is a flowchart of the message processing method in another embodiment;

[0027] Figure 7 It is a diagram of rewriting conditions in an embodiment;

[0028] Figure 8 It is a flowchart block diagram of the message processing method in another embodiment;

[0029] Figure 9 It is a structural block diagram of the message processing apparatus in an embodiment;

[0030] Figure 10 It is an internal structure diagram of a computer device in an embodiment. Detailed Description of the Invention

[0031] To make the objectives, technical solutions and advantages of this application more clear, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.

[0032] The message processing method provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other network servers. The terminal 102 and the server 104 can be used alone to execute the message processing method in this application, and the terminal 102 and the server 104 can be used in cooperation to execute the message processing method in this application.

[0033] Taking the cooperation between the terminal 102 and the server 104 to execute this application as an example, when specifically processing messages, the user can carry out business consultations in combination with their own business processing needs, such as shopping needs, payment needs, game guidance needs, etc. For example, the user can describe the problem through the terminal 102 in ways such as text or voice. Then, the terminal 102 interacts with the server 104. The server 104 can obtain multiple rounds of session messages from the terminal 102 during the interaction. The multiple rounds of session messages include the current session message and the historical session messages before the current session message. The server 104 inputs the current session message and the historical session messages into the large model and generates sentences through the large model. The server 104 performs intent recognition based on the generated sentences to obtain the first intent corresponding to the current session message. For any historical session message, the server 104 obtains the second intent corresponding to the targeted historical session message. When the first intent matches any of the second intents successfully, the server 104 determines the reply content based on the first intent and the current session message, and feeds back the reply content to the terminal 102, so that the user can perform relevant operations based on the reply content on the terminal 102 to solve the relevant business.

[0034] Among them, an application program runs on the terminal 102, and the application program refers to any computer program that can provide an interaction platform between the terminal 102 and the server 104. For example, various applications with business consultation services such as social applications, educational applications, shopping applications, and video applications. The terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, portable wearable devices, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, vehicle-mounted intelligent devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods.

[0035] It should be noted that the embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0036] The artificial intelligence (AI) involved herein is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0037] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The message processing method of the present application will be introduced in detail below:

[0038] In one embodiment, as Figure 2 shown, a message processing method is provided. This method is applied to a computer device (the specific computer device can be Figure 1Taking the terminal 102 or the server 104 in it as an example for illustration, it includes the following steps:

[0039] Step 202, obtain multi-round session messages from the user terminal during the interaction process. The multi-round session messages include the current session message and the historical session messages before the current session message.

[0040] Among them, the interaction process is a process in which the user conducts conversations and exchanges with the computer device through the user terminal to obtain the required business information.

[0041] The session message is the conversation information generated during the interaction process between the user terminal and the computer device.

[0042] The multi-round session messages include the current session message and the historical session messages. Current and historical are two relative concepts. The current session message is the session message generated at the current moment or during the current time period, and the historical session message is the session message generated before the current moment or before the current time period.

[0043] In some embodiments, the current session message and the historical session message may be consultation messages for the same business. For the same business, the user may make multiple descriptions about the business. At this time, the current session message and the historical session message are related content. The current session message can specifically be the follow-up message of the historical session message, that is, a description that supplements, extends, and correlates with the historical session message. For example, the historical session message can be "Why hasn't the cash withdrawal arrived yet", and the current session message can be "100 yuan yesterday". At this time, both the historical session message and the current session message are consultations regarding the cash withdrawal problem, and the current session message is a further supplement to the specific time, amount, etc. of the cash withdrawal.

[0044] In some other embodiments, the current session message and the historical session message may be consultation messages for different businesses. For different businesses, the current session message and the historical session message have no or weak correlation. For example, the historical session message can be "Why hasn't the cash withdrawal arrived yet", and the current session message can be "How to retrieve the payment password". The historical session message is a consultation regarding cash withdrawal, while the current session message is a consultation regarding the payment password.

[0045] In some other embodiments, the current session message may correspond to multiple historical session messages. The current session message and some of the historical session messages may be consultation messages for the same business. Of course, the current session message and multiple historical session messages may all be consultation messages for the same business, or all belong to consultation messages for different businesses.

[0046] In some embodiments, every time the computer device receives a new session message from the user terminal, the new session message is the current session message, and the session messages generated before the current session message are the historical session messages. That is, every time a new session message is received, the message processing method provided by the embodiments of the present application is executed, so that for each session message in the interaction process, there can be a corresponding reply content adapted thereto, and a smooth conversation experience can be achieved. In a specific application, during the shopping process, the user will consult about the product model and the usage method of the product. The user can conduct shopping business consultations through text or voice based on the shopping program on the user terminal. The computer device can obtain multiple rounds of session messages related to the shopping business from the user terminal during the interaction process and process the obtained multiple rounds of session messages. For another example, during the payment process, there may be a situation where the payment fails. The user can conduct payment business consultations based on the relevant program on the user terminal to clarify the reason for the payment failure.

[0047] Step 204, input the current session message and the historical session message into the large model together, and generate a sentence through the large model.

[0048] Among them, the large model is a machine learning model with a huge parameter scale and complexity, and specifically can refer to a neural network model with tens to hundreds of billions of parameters. The large model itself has good text understanding and generation capabilities. The large model is pre-trained and can be used to rewrite the current session message and the historical session message.

[0049] The generated sentence is a standardized sentence obtained by the large model summarizing and rewriting the current session message and the historical session message. The standardized sentence has the characteristics of prominent key information, no typos or grammar errors, conciseness, and summarization.

[0050] Specifically, the computer device can input the current session message and the historical session message into the large model together. The large model can summarize and rewrite the current session message and the historical session message based on the message characteristics of the current session message and the historical session message. For example, the large model can determine the correlation degree between the current session message and the historical session message, the key information in the current session message and the historical session message, etc., and adaptively summarize and rewrite the current session message and the historical session message to output a generated sentence.

[0051] In some embodiments, the large model can be fine-tuned based on an existing base model, so as to utilize the text understanding and generation capabilities of the base model to build the summarization and rewriting capabilities. By fine-tuning the base model, a model related to the actual task can be specifically constructed, improving the training efficiency and model performance. For example, the base model can be Baichuan-13B-Chat (Baichuan large language model), or other large-scale language models, such as the GPT-4 model (Generative Pre-trained Transformer 4, a natural language technology based on artificial neural networks), ChatGPT-3 (Generative Pre-trained Transformer 3, a natural language technology based on artificial neural networks), and PaLM 2 (Pretraining-augmented Language Model, a pre-trained large language model), etc.

[0052] In some embodiments, when the computer device fine-tunes the base model, a certain adjustment period can be set, so that the adjustment can be carried out according to the adjustment period, while ensuring the business capabilities of the large model, and saving computing resources, operation costs, etc.

[0053] In some embodiments, when the computer device fine-tunes the base model, methods such as Prompt-Tuning (prompt fine-tuning), instruction-Tuning (instruction fine-tuning), prefix-Tuning (prefix tuning), fine-Tuning (pre-trained model fine-tuning), adapter-Tuning (adapter fine-tuning) can be adopted, and the embodiments of the present application are not limited thereto.

[0054] In a specific application, when the computer device adopts prompt fine-tuning, specific prompts or instructions can be added to the input of the model to guide the model to generate a specific type of response. The prompt can be a question, a description, or a complete sentence, which is used to guide the model to generate relevant content.

[0055] Step 206, perform intent recognition based on the generated statement to obtain the first intent corresponding to the current session message.

[0056] Among them, intent recognition is a processing process for determining the user intent represented by the generated statement.

[0057] Intent is used to represent the business processing requirements expressed by the user during the business consultation process, and specifically can refer to the corresponding solution desired for a certain problem point. The first intent is the user's business processing requirement corresponding to the current session message, that is, the user's business processing requirement at the current moment or during the current time period.

[0058] Specifically, the server can perform intent recognition on the generated statement. For example, it can use a pre-set intent recognition model or other intent recognition algorithms other than the intent recognition model to perform intent recognition based on the generated statement, and obtain the first intent corresponding to the current session message. For instance, when using other intent recognition algorithms for intent recognition, it can be based on entity detection, knowledge graph, or rule template parsing, etc. For entity detection, entities in the generated statement, such as person names, place names, organization names, proper nouns, etc., can be recognized, and then intent recognition can be carried out based on the recognized entities. For rule template parsing, the domain to which the generated statement belongs can be determined, and after determining the domain, intent verbs and intent interrogative words can be used to judge the intent.

[0059] In some embodiments, the intent recognition model can be a network model constructed based on artificial intelligence algorithms. When constructing the intent recognition model, it can be constructed based on various algorithms such as supervised learning algorithms and unsupervised learning algorithms, which are not limited herein. When constructing the intent recognition model, it can be constructed using, for example, recurrent neural networks, long short-term memory networks, recursive neural networks, etc.

[0060] In some embodiments, performing intent recognition based on the generated statement to obtain the first intent corresponding to the current session message includes: performing word segmentation processing on the generated statement to obtain the keywords in the generated statement; performing vectorization representation processing on the keywords to obtain the feature representation of the keywords; and performing intent classification based on the feature representation to obtain the first intent corresponding to the current session message.

[0061] Specifically, the computer device can use a word segmentation tool to perform word segmentation processing on the generated statement to determine various types of words in the generated statement, such as nouns, entity words, non-nouns, stop words, etc. Further, the computer device can remove stop words and non-noun words. Stop words refer to words without actual meaning, such as pronouns, auxiliary words, adjectives, adverbs, etc., and retain entity words and nouns, such as words with actual meaning like time and amount, and use the retained entity words and nouns as keywords. Then, the computer device can perform vectorization representation processing based on the determined keywords to obtain the feature representation of the keywords. The computer device can classify based on the feature representation of the keywords using an intent recognition model, such as an NLP model (Natural Language Processing model), to determine the intent category represented by the vectorization representation processing.

[0062] In the above embodiments, the computer device can determine the keywords of the generated statement and obtain the feature representation of the keywords, so that by performing intent classification based on the feature representation of the keywords, the first intent can be accurately obtained.

[0063] Step 208, for any historical session message, obtain the second intent corresponding to the targeted historical session message.

[0064] Among them, the second intention is the user service processing requirement corresponding to the historical conversation message, that is, the user service processing requirement at the current moment or before the current time period. In multiple rounds of conversation messages, the second intention can specifically refer to the historical intention. The types of historical intentions can be one or multiple. The actual types of historical intentions are related to the actual interaction scenarios, user requirements, business scenarios, user's understanding of the problem, or communication habits, etc.

[0065] Specifically, for any historical conversation message, the computer device can obtain the second intention corresponding to the targeted historical conversation message in the preset intention database.

[0066] In some embodiments, the second intention obtained by the computer device can be the same as the problem point pointed to by the first intention; the second intention obtained by the computer device can also be different from the problem point pointed to by the first intention; the second intention obtained by the computer device can also be partially the same as the problem point pointed to by the first intention and partially different from the problem point pointed to by the first intention.

[0067] In some embodiments, each historical conversation message corresponds to a second intention. When a historical conversation message was once the current conversation message, the first intention corresponding to the current conversation message was obtained by summarizing the current conversation message and its historical conversation messages through a large model to generate a statement, and then performing intention recognition on the generated statement. When a new conversation message arrives, the current conversation message becomes a historical conversation message. Correspondingly, the first intention corresponding to the current conversation message also changes. For the sake of convenience of description, the changed first intention is used as the second intention, and the second intention is the intention corresponding to each historical conversation message.

[0068] It can be understood that the terms "first", "second", etc. used in this application can be used in this text to describe various components, but these components are not limited by these terms. These terms are only used to distinguish the first component from another component. For example, without departing from the scope of this application, the first intention can be called the second intention.

[0069] In a specific application, the second intention can include multiple intentions, such as it can include Figure 1 intention Figure 2 intention Figure 1 intention Figure 2 and the categories of intention Figure 1 are different. The first intention can be the same as the problem point pointed to by intention Figure 2 or can be the same as the problem point pointed to by intention Figure 1 intention Figure 2 or can be different from the problem points pointed to by intention Figure 1 intention Figure 2 and intention Figure 2 are all different.

[0070] Step 210, when the first intent matches any second intent successfully, determine the reply content based on the first intent and the current session message, and feedback the reply content.

[0071] Among them, when the first intent matches any second intent successfully, it means that there is a second intent that is relevant or associated with the first intent. Specifically, it can be that the problem points pointed to by the first intent and any second intent are the same.

[0072] The reply content is the solution corresponding to the current session message among the problem points pointed to by the first intent.

[0073] Specifically, the computer device can match the first intent with each second intent respectively to determine whether there is a second intent that matches the first intent. When there is a second intent that matches the first intent, it indicates a successful match, and then the user's intent can be locked, that is, the user's problem point can be locked. When the computer device locks the user's problem point, it can determine the reply content based on the current session message under the locked problem point, and feedback the reply content to the user terminal to assist the user in processing the business.

[0074] In some embodiments, when the computer device matches the first intent and each second intent, it can be carried out through an algorithm that can achieve category matching. The algorithms used can be various algorithms such as similarity analysis algorithm, decision tree algorithm, and support vector machine algorithm.

[0075] In other embodiments, when the computer device matches the first intent and each second intent, it can obtain the first identifier of the first intent and the respective second identifiers corresponding to each second intent, and compare the first identifier with each second identifier respectively to determine whether there is a second identifier that is the same as the first identifier among the second identifiers. When there is a second identifier that is the same as the second identifier, the second intent corresponding to the second identifier that is the same as the first identifier is determined as the intent that matches the first intent. Among them, the identifier can be presented as text, or various identifiers such as letters, numbers, or feature codes, which are not limited herein.

[0076] In one embodiment, as Figure 3 shown, it is a flowchart of a message processing method in a specific application:

[0077] Figure 3 It involves a user terminal and a server. An application program that supports the intelligent customer service function runs on the user terminal, and the server is the backend of the intelligent customer service, which can process the business consultation messages of the user to feedback the reply content to the user.

[0078] The server can obtain multiple rounds of conversation messages generated by users during business consultation. For the obtained multiple rounds of conversation messages, the server can bring the multiple rounds of conversation messages into the big model, and the big model can summarize and rewrite the multiple rounds of conversation messages to generate a problem summary that conforms to the user's description. Through the rewriting of the big model, while eliminating meaningless content in multiple user interactions, the content that needs to be emphasized in the user's original words can be retained to generate a smooth and logical problem description.

[0079] Furthermore, the server may input the problem summary obtained after the summary is rewritten into the recognition module, and the recognition module may identify the business intent of the problem summary through the NLP natural language processing model to determine the current intent of the user.

[0080] After determining the user's current intention, the computer device can obtain the historical intention, that is, obtain the intention corresponding to the historical conversation message, and carry out intention matching to determine whether the current intention matches any historical intention. If it matches, the computer device can determine the reply content. If it does not match, the computer device can feedback the preset reply words, wherein the reply words can be words used to guide the user to the next round of conversation, or can be a basic solution related to the current intention.

[0081] In the above message processing method, multiple rounds of conversation messages from the user terminal during the interaction process are obtained, wherein the multiple rounds of conversation messages include the current conversation message and the historical conversation message before the current conversation message. The current conversation message and the historical conversation message are input into the big model together, and then a standardized generated sentence can be output based on the big model, and then the intention recognition is performed on the standardized generated sentence, which can reduce the interference of similar, redundant, and colloquial content in the interaction process on the intention recognition, and improve the accuracy of the intention recognition. Furthermore, for any historical conversation message, the second intention corresponding to the historical conversation message is obtained, and when the first intention successfully matches any second intention, the user intention is locked, so that on the basis of locking the user intention, the reply content is determined based on the first intention and the current conversation message, and the accurate reply content is fed back to the user terminal, so that the interaction process is closer to the real person dialogue experience and a smooth dialogue process is achieved.

[0082] In some embodiments, the message processing method further includes: when the first intent does not match any second intent, determining suggestive content based on the first intent, and feeding back the suggestive content; the suggestive content is used to guide the next round of conversation.

[0083] Among them, the prompting content is some understandings and suggestions provided to guide users to ask more focused questions, and the prompting content can be used to guide the next round of conversation. The prompting content may include associated questions, derivative questions, or basic question answers related to the first intention, etc.

[0084] Specifically, when the computer device determines that each second intention does not match the first intention, that is, among the second intentions, there is no intention that points to the same problem point as the first intention. The computer device can obtain the prompting content that matches the first intention by querying the intention prompt library, and feedback the prompting content to the user terminal, thereby guiding the user to start the next round of conversation based on the prompting content.

[0085] In some embodiments, the intention prompt library can be preset. For various types of intentions, in the intention prompt library, there can be established corresponding prompting content for each intention respectively. For example, the intentions can include cash withdrawal intention, payment password intention, limit intention, etc., and thus the prompting content can be respectively guiding content related to cash withdrawal, guiding content related to payment password, and guiding content related to limit.

[0086] In the above embodiments, when the computer device determines that the first intention does not match any of the second intentions, it can determine the prompting content based on the first intention to guide the next round of conversation, can guide the user to ask targeted questions, and thus can quickly locate the user's true intention, improving the efficiency and accuracy of message processing.

[0087] In some embodiments, the message processing method further includes: storing the first intention corresponding to the current session message; when the next round of session message arrives, the current session message becomes a historical session message, and the first intention corresponding to the stored current session message becomes the second intention corresponding to the historical session message when processing the new session message.

[0088] Among them, the historical session message and the current session message are relative concepts, which are specifically related to the rounds of the conversation in the interaction process. When a new round of session message arrives, the current session message that already exists before the arrival of the new round of session message will become a historical session message. Correspondingly, when processing the new round of session message, the stored first intention becomes the second intention.

[0089] In some embodiments, when the computer device recognizes the first intention corresponding to the current session message, it can store the recognized first intention corresponding to the current session message. For example, the first intention can be stored in a preset storage unit, and the storage unit can include a database, a message queue, a linked list, etc.

[0090] In some embodiments, when there are duplicate second intents among the second intents corresponding to the stored historical session messages, the computer device can clear the duplicate second intents, thereby releasing storage resources and improving the efficiency of the computer device without affecting the matching effect, and to a certain extent improving the message processing efficiency.

[0091] In other embodiments, the computer device can also clear a preset number of second intents according to the storage time of the second intents in combination with the chronological order when a preset time period is reached or the number of stored second intents reaches a set number threshold. For example, the second intents with earlier time can be preferentially cleared. Considering the actual conversation process, the user may have changed the intent during a long conversation, and the second intents with longer storage time are no longer of reference significance, so the second intents with longer storage time can be cleared to release the cache resources.

[0092] In the above embodiments, the computer device can store the first intent corresponding to the current session message, and when processing a new session message, the first intent corresponding to the current session message will become the second intent corresponding to the historical session message. Subsequently, before performing intent matching, the stored second intent can be directly obtained, which improves the message processing efficiency.

[0093] In some embodiments, the generated statement includes a question statement. Inputting the current session message and the historical session message into the large model and outputting the generated statement through the large model includes: inputting the current session message and the historical session message into the large model according to a preset input format; through the large model, according to a preset statement rewriting instruction, summarizing the questions for the current session message and the historical session message to obtain the summarized question statement and output it.

[0094] Among them, the preset input format is the format for inputting the set session message into the large model. The preset input format can be determined according to the round corresponding to the session message, the generation time of the session message, the type of the session message, etc.

[0095] The statement rewriting instruction can be used to indicate the specific way for the large model to rewrite the current session message and the historical session message.

[0096] Specifically, the computer device can input the current session message and the historical session message into the large model according to a preset input format. For example, the current session message and the historical session message can be sorted in the order of the time of receiving the messages, with the session message received earlier ranked first and the message received later ranked second, to obtain the sorted historical session message and the current session message, and then input the sorted historical session message and the current session message into the large model. The large model summarizes the questions from the sorted current session message and the historical session message, obtains the summarized question statement and outputs it. By setting the input format of the current session message and the historical session message, the large model can process the obtained session messages according to the input format, improving the rewriting efficiency and accuracy of the large model.

[0097] In some embodiments, the statement rewriting instruction may include an indication to extract key information from the historical session message and an indication to combine the current session information with the extracted key information. For example, when extracting the key message from the historical session message, various types of words in the historical session message can be determined, such as nouns, entity words, non-nouns, stop words, etc. Further, the computer device can remove the stop words and non-noun words. Stop words refer to words without actual meaning, such as pronouns, auxiliary words, adjectives, adverbs, etc., and retain the entity words and nouns, such as words with actual meaning like time, amount, person name, etc., and use the retained entity words and nouns as key information.

[0098] In some embodiments, the computer device can also determine the preset input format based on the type of the message. The type of the message can be specifically divided into historical messages and current messages. The computer device can sort the current session message and the historical session message in the order that the historical message is after and the current message is before, obtain the sorted historical session message and the current session message, and input the sorted historical session message and the current session message into the large model. The large model summarizes the questions from the sorted current session message and the historical session message, obtains the summarized question statement and outputs it.

[0099] In a specific application, the statement rewriting instruction can be "Please extract the key information from the historical session message to complete the current session message". The large model can extract the key information from at least one historical session message respectively, combine it with the current session message to summarize the question, and obtain the summarized question statement. For example, the current session message can be "I didn't use much money either", and the historical session messages can include historical session message 1 "My spare change" and historical session message 2 "Why was there a sudden limit on its use". Based on the statement rewriting instruction, the question statement obtained by the large model can be "I didn't use much of my spare change either. Why was there a sudden limit?".

[0100] In the above embodiments, the computer device may input the current session message and the historical session message into the large model according to a preset input format. Then, the large model may summarize the current session message and the historical session message according to a preset sentence rewriting instruction to obtain a standardized question sentence. The obtained question sentence is more accurate in positioning the user's question, which can improve the accuracy of intent recognition and further improve the message processing accuracy.

[0101] In some embodiments, determining a reply content based on a first intent and a current session message includes: obtaining a question-and-answer database, where the question-and-answer database includes multiple intents and question-and-answer vector pairs under each intent; determining a target intent matching the first intent in the question-and-answer database; and determining the reply content based on the current session message and the question-and-answer vector pairs under the target intent.

[0102] Among them, the question-and-answer database is a database for storing various intents and the corresponding question-and-answer contents. The question-and-answer vector pair is the vector form of the question-and-answer content, that is, the question-and-answer contents in the question-and-answer database are stored in vector form.

[0103] In some embodiments, the computer device may determine a target intent matching the first intent in the question-and-answer database based on the first intent. For example, the computer device may perform similarity matching between the first intent and various intents in the question-and-answer database respectively, and determine the target intent according to the matching result. Or the computer device may also determine the intent identifier of the first intent, determine a target identifier matching the intent identifier of the first intent from the question-and-answer database, and determine the intent corresponding to the target identifier as the target intent.

[0104] In some embodiments, after the computer device determines the target intent, it may determine the reply content based on the current session message and the question-and-answer vector pairs under the target intent.

[0105] In the above embodiments, the computer device first searches for a target intent matching the first intent in the question-and-answer database, so that it can lock the user's intent without special downstream operation for the user's intent, reduce the operation cost, and further determine the reply content accurately according to the current session message and the question-and-answer vector pairs under the target intent.

[0106] In some embodiments, the question-and-answer vector pair includes a question vector. Determining the reply content based on the current session message and the question-and-answer vector pairs under the target intent includes: encoding the current session message to obtain an encoded vector; calculating the similarity between the encoded vector and each question vector under the target intent; and determining the reply content based on the reply vector associated with the question vector corresponding to the similarity meeting the preset similarity condition.

[0107] Among them, the question vector is the vector form corresponding to the question in the Q&A content. For each question, there is a corresponding answer. Therefore, the Q&A vector pair also includes an answer vector.

[0108] The preset similarity condition is the condition set for determining the answer vector from the Q&A database according to the similarity. When setting the similarity condition, it can be adaptively set in combination with the actual business scenario, message processing accuracy requirements, etc.

[0109] Similarity is a parameter used to characterize the degree of similarity between the encoded vector and each question vector under the target intention.

[0110] Specifically, the computer device can encode the current session message to obtain an encoded vector for vector characterization of the current session message, and then calculate the similarity between the encoded vector and each question vector under the target intention respectively to obtain the similarity between the encoded vector and each question vector. The computer device can determine the answer content according to each similarity and the preset similarity condition.

[0111] In some embodiments, it can be the similarity with the largest value among each similarity that meets the preset similarity condition. The computer device can determine the answer vector associated with the question vector corresponding to the similarity with the largest value from each similarity, and perform decoding processing based on the answer vector to obtain the answer content, and feedback the answer content to the user terminal. Thus, the solution most matching the current session message can be fed back to the user.

[0112] In other embodiments, the preset similarity condition can also be determined based on a similarity threshold, that is, when any similarity reaches the set similarity threshold, it is determined that the preset similarity condition is met. The computer device determines the answer vector associated with the question vector corresponding to the similarity that reaches the similarity threshold, and performs decoding processing based on the answer vector to obtain the answer content, and feedbacks the answer content to the user terminal. Thus, at least one solution matching the current session message can be fed back to the user.

[0113] In some embodiments, refer to Figure 4 as shown, it is a flowchart for determining the answer vector:

[0114] Among them, the semantic model is a model used to encode text, combined with Figure 5For example, the semantic model can be a model that encodes user questions and Q&A content. The semantic model can be trained based on BERT (Bidirectional Encoder Representations from Transformers), GPT3 (Generative Pre-trained Transformer 3), GPT4 (Generative Pre-trained Transformer 4), Clip (Contrastive Language-Image Pre-Training), etc. The embodiments of this application do not limit this.

[0115] The semantic model can encode the Q&A content to obtain a Q&A vector pair and store the Q&A vector pair in the Q&A database.

[0116] In the actual process of determining the reply content, the computer device can encode the obtained user question, that is, the current user message, based on the semantic model to obtain an encoded vector, and then perform a similarity analysis on the encoded vector and the question vector of the Q&A vector pair in the Q&A database to determine a reply vector, and determine the reply content based on the reply vector.

[0117] In the above embodiments, the computer device can accurately determine the reply content based on the encoded vector and the question vector and calculate the similarity between the encoded vector and the question vector, ensuring the accuracy of message processing. In the actual business processing process, since the reply content is determined under a locked intention, the operation cost of annotating the fragmented questions of the user's follow-up text or the operation of the task session can be reduced.

[0118] In some embodiments, refer to Figure 5 As shown, it is a flowchart of a message processing method in an embodiment:

[0119] Figure 5 It may involve a user terminal and a computer device. An application program can run on the user terminal, and the application program can support the user to conduct business consultations with the intelligent customer service. The computer device is the processing backend of the intelligent customer service, which can obtain the session messages transmitted by the user terminal and conduct processing. When the user has a business processing requirement, the user can consult the intelligent customer service through the user terminal. For actual problems, the user can describe them in text or voice, which is not limited here.

[0120] When a computer device processes the current session message transmitted by a user terminal, it can be carried out in two parts, namely S1: Summary rewriting below, and S2: Intent recognition model. In the summary rewriting part below, the computer device can judge the obtained current session message to determine whether the current session message is the first sentence. If so, the computer device can directly input the current session message into the intent recognition model, and the intent recognition model can perform intent recognition based on the current session message to obtain the intent corresponding to the first sentence.

[0121] When the current session message is not the first sentence, the computer device can substitute the multi-round session messages into the large model. The multi-round session messages include the current session message and the historical session messages of the current session message. The large model rewrites based on the current session message and the historical session messages and outputs the user summary rewriting, that is, the complete text description of the user. Among them, the large model can be a vertical model debugged by applying P-Tuning (fine-tuning method). Combining the text processing ability of the large model itself, it can generate a problem summary that conforms to the user's description. The large model can rewrite the complete text description of the user, that is, the multi-round session messages including the current session message and the historical session messages, to summarize the content described by the user. While eliminating the meaningless content in the interaction process, it retains the content that needs to be emphasized in the user's original words and generates a smooth and logical problem description.

[0122] Furthermore, the computer device can input the user summary rewriting output by the large model into the intent recognition model. The intent recognition model can specifically be an NLP natural language processing model. The NLP natural language processing model performs intent recognition on the user summary rewriting generated by the large model to obtain the intent corresponding to the user summary rewriting. Among them, the intent corresponding to the user summary rewriting can be a new intent, and the intent corresponding to the first sentence becomes the historical intent.

[0123] The computer device can match the new intent with the historical intent to determine whether the new intent is the same as the historical intent, so as to judge whether the user's intent has switched. When the matching is successful, indicating that the new intent is the same as the historical intent, the computer device can lock the user's intent. When the matching fails, indicating that the new intent is different from the historical intent, that is, the user has switched the intent, the computer device can use the set reply words to reply.

[0124] After the computer device locks the user's intent, it can obtain the Q&A database and determine the reply content for the locked user intent according to the Q&A database. Specifically, it can perform short text semantic matching recognition based on the Q&A database to achieve the positioning of multiple progressive words and perform targeted subsequent progressive word replies.

[0125] In some embodiments, the training steps of the large model include the following steps:

[0126] Step 602, obtain multi-turn conversation message samples; the multi-turn conversation message samples include at least one of multi-intent conversation messages, single-intent conversation messages, or third-party conversation messages.

[0127] Among them, multi-intent conversation messages are related to the business and can include conversation messages with at least two intents; single-intent conversation messages are related to the business and only include conversation messages with one intent; third-party conversation messages are conversation data that is not related or has little association with business processing. By using various types of data to train the large model, the generality of the large model can be improved, so that the trained large model can be applied in various businesses.

[0128] In some embodiments, the computer device can obtain multi-intent conversation messages, single-intent conversation messages, and third-party conversation messages, and train the large model based on the multi-intent conversation messages, single-intent conversation messages, and third-party conversation messages. Of course, the computer device can also only obtain one or two of the multi-intent conversation messages, single-intent conversation messages, and third-party conversation messages to train the large model, and the actual acquisition situation can be adaptively adjusted in combination with actual efficiency requirements, computer resources, etc.

[0129] Step 604, obtain the annotated sentences obtained by rewriting the multi-turn conversation message samples; through the large model to be trained, summarize based on the multi-turn conversation message samples to obtain the predicted sentences.

[0130] Among them, the annotated sentences are the sentences obtained by rewriting the multi-turn conversation message samples. The predicted sentences are the sentences obtained by the large model summarizing the multi-turn conversation message samples.

[0131] Specifically, the computer device obtains the multi-turn conversation message samples obtained and rewrites them to obtain the annotated sentences. The large model can summarize the obtained multi-turn conversation message samples to obtain the summarized predicted sentences.

[0132] In some embodiments, the annotated sentences can also be pre-annotated. The pre-annotated sentences can be stored in the storage service. When the computer device has a model training requirement, it can directly obtain the annotated sentences from the storage service.

[0133] Step 606, based on the predicted sentences and the annotated sentences, adjust the parameters of the large model to be trained and continue training until the training ends to obtain the trained large model.

[0134] In some embodiments, the computer device may train the large model to be trained based on the prediction statement and the annotation statement, and end the training when the training stop condition is met, obtaining a trained large model. The training stop condition may specifically be reaching a preset number of training times, the loss being less than the loss threshold, or the training duration reaching a preset duration, etc., which is not limited in the embodiments of the present application.

[0135] In the above embodiments, the computer device may obtain a large number of multi-round conversation message samples, rewrite the obtained multi-round conversation message samples, and then continue to train the large model to be trained by adjusting the parameters based on the prediction statement and the annotation statement, which can improve the generation quality of the large model and also significantly improve the generality of the large model. Since various types of data are used in the training process and the combination of text spoken language and words is processed, it can be applied to various services in the customer service scenario after the model is built.

[0136] In some embodiments, obtaining the annotation statement obtained by rewriting the multi-round conversation message sample includes: rewriting each multi-round conversation message sample according to the preset rewriting conditions to obtain the rewritten annotation statement; the preset rewriting conditions include at least one of keyword retention, only referring to the previous message when rewriting, typo correction, invalid information filtering, or separating non-strongly related questions into sentences.

[0137] Among them, the preset rewriting conditions are the conditions set for rewriting the multi-round conversation message sample. The preset rewriting conditions may include keyword retention, only referring to the previous message when rewriting, typo correction, invalid information filtering, and separating non-strongly related questions into sentences, or may only include any one or several of keyword retention, only referring to the previous message when rewriting, typo correction, invalid information filtering, and separating non-strongly related questions into sentences, which is not limited here.

[0138] Keyword retention means retaining the key elements in the multi-round conversation message sample, such as key elements like time and amount. Typo correction means correcting grammar errors and typos in the multi-round conversation message sample; only referring to the previous message when rewriting means that when rewriting the current conversation message in the multi-round conversation message sample, only the previous message (i.e., the historical conversation message) in the multi-round conversation message sample can be referred to, and the following message of the current conversation message cannot be referred to for rewriting. Invalid information filtering means removing words without actual meaning such as modal particles and auxiliary words. Separating non-strongly related questions into sentences means that for the current conversation message in the multi-round conversation message sample, if it has little or no association with its historical conversation message, the current conversation message is directly output.

[0139] In one embodiment, refer to Figure 7The following is a diagram of annotation requirements in a specific application. The annotation requirements may include annotating the current sentence in conjunction with the previous sentence as a whole; discarding redundant content and summarizing simply and clearly; using the original sentence and words and avoiding making up new words; rewriting the text without in-depth understanding; correcting typos or grammatical errors. Computer equipment can be used according to Figure 7 The annotation shown requires annotation of multiple rounds of conversation message samples. Of course, annotation can also be performed manually.

[0140] In a specific application, referring to Table 1, a rewriting result obtained by rewriting a multi-intent conversation message in an embodiment is shown:

[0141] Table 1

[0142]

[0143] In combination with Table 1, it can be seen that user questions may include "My withdrawal has not arrived", "Yesterday's", "I withdrew 100 yuan, please help me check" and "How do I retrieve my payment password?". The rewritten annotation corresponding to "My withdrawal has not arrived" is "My withdrawal has not arrived", the rewritten annotation corresponding to "Yesterday's" is "My withdrawal yesterday has not arrived", the rewritten annotation corresponding to "I withdrew 100 yuan, please help me check" is "My withdrawal of 100 yuan yesterday has not arrived, please help me check", and the rewritten annotation corresponding to "How do I retrieve my payment password?" is "How do I retrieve my payment password?"

[0144] In some embodiments, a large model to be trained is used to summarize based on multiple rounds of conversation message samples to obtain a predicted statement, including: inputting current conversation messages and historical conversation messages in the multiple rounds of conversation message samples into the large model to be trained according to a preset input format; using the large model to be trained, according to preset sentence rewriting instructions, summarizing problems for the multiple rounds of conversation message samples, obtaining the summarized predicted statement and outputting it.

[0145] Specifically, the computer device can input the current conversation message and the historical conversation message in the multi-round conversation message sample into the large model according to the preset input format, such as sorting the current conversation message and the historical conversation message in the multi-round conversation message sample according to the time sequence of receiving the messages, putting the conversation message with the earlier receiving time in the front and the message with the later receiving time in the back, obtaining the sorted historical conversation message and the current conversation message, and inputting the sorted historical conversation message and the current conversation message into the large model. The large model summarizes the problem of the current conversation message and the historical conversation message in the sorted multi-round conversation message sample based on the preset sentence rewriting instruction, obtains the summarized problem sentence and outputs it.

[0146] In the above embodiments, the computer device inputs the current session message and the historical session message in the multi-round session message samples into the large model to be trained according to the preset input format, so that the large model can summarize the current session message and the historical session message in the multi-round session message samples according to the preset statement rewriting instruction to obtain a predicted statement, and then the training of the large model is carried out.

[0147] In some embodiments, based on the predicted statement and the annotated statement, the parameters of the large model to be trained are adjusted and then the training is continued until the training is completed to obtain a trained large model, including: determining the statement loss between the predicted statement and the annotated statement based on the predicted statement and the annotated statement; adjusting the parameters of the large model to be trained according to the statement loss and then continuing the training until the training is completed to obtain a trained large model.

[0148] Specifically, the computer device can calculate the statement loss between the predicted statement and the annotated statement, and adjust the parameters of the large model based on the calculated statement loss. For example, the parameters of a specific layer of the large model can be adjusted to reduce the adjustment amount, and based on the large model after adjusting the parameters, the training is continued until the training is completed to obtain a trained large model.

[0149] In the above embodiments, the computer device can adjust the parameters of the large model by calculating the statement loss between the predicted statement and the annotated statement, and continue to train the large model after parameter adjustment until the training is completed to obtain a trained large model.

[0150] The present application also provides an application scenario that applies the above message processing method. Specifically, taking an actual application scenario as an example, the message processing method of the present application is described in detail as follows:

[0151] Figure 8 The user terminal and the computer device are involved. An application program is running on the user terminal, and the application program can support the user to conduct business consultations. Figure 8 It can be the business consultation involved in the user's QR code payment process.

[0152] The session management service is the processing backend of a computer device and may include a message access interface, a session management service, and a message storage system. Among them, the message access interface is an interface for realizing information forwarding between the user terminal and the computer device, and can be used to obtain multi-round session messages transmitted by the user terminal and transmit them to the session management service and the message storage system. The message management service can be used to manage the received multi-round session messages, and the message storage system can store the received multi-round session messages. For example, the received multi-round session messages may include "My spare change", "Why is there a sudden limit when using it", and "I didn't use much money either". Among them, "I didn't use much money either" can be the current session message, and "My spare change" and "Why is there a sudden limit when using it" are historical session messages. The session management service can input the received multi-round session messages into the large model in a certain order, and the large model can rewrite them to output the rewritten text. The rewritten text can be "I didn't use much money for my spare change either. Why is there a sudden limit?".

[0153] The computer device can send the rewritten text to the intent recognition module, and the intent recognition model in the intent recognition module can perform intent recognition on the rewritten text to determine the first intent.

[0154] The computer device can obtain the second intent, match the first intent with each second intent respectively to determine whether there is a second intent that matches the first intent. In the case where there is a second intent that matches the first intent, it indicates that the intent has not been modified, and thus the user's intent can be locked, that is, the user's problem point can be locked. When the computer device determines that the intent has not been modified, it can query the Q&A database under the locked problem intent and determine the reply content based on the Q&A database.

[0155] When the computer device determines the reply content based on the Q&A database, it can perform vector library semantic recognition, that is, determine the target intent that matches the first intent in the Q&A database, and determine the reply content based on the Q&A vector pair corresponding to the target intent. For "I didn't use much money either", the computer device can determine that the matching question vector is Q-1, and the corresponding reply content is A-1. In the case where there is no second intent that matches the first intent, it indicates that the intent has been modified, and the computer device can give the user a prompt based on the basic intent library to guide the user to ask more focused questions.

[0156] The message processing method provided in this application can apply the large model text summarization and generation ability. In the intelligent customer service scenario, when interacting with the user's follow-up text, the large model can rewrite the fragmented description of the user's follow-up single sentence, combine all the content of the previous text to summarize it into a brief description, and then combine the intent recognition model to perform intelligent customer service reply for semantic understanding.

[0157] In addition, for large models, a preset adjustment period can be set to fine-tune the large model regularly to improve the generality and accuracy of the large model. When fine-tuning the large model, three types of sample data can be obtained, and the sample data can be labeled to obtain labeled data. The sample data can include: single-intention data, multi-intention data, and third-party corpora. Among them, the number of labeled data can be 110,000, specifically including: 90,000 single-intention data, 10,000 multi-intention data, and 10,000 third-party corpora.

[0158] When labeling various types of data, a unified rewriting standard is formulated, including rules such as retaining business keywords, correcting typos, filtering invalid information, and separating non-strongly related questions into sentences. The historical questions of users in the dataset are labeled sentence by sentence as training data. Specifically, the descriptions of various types of data are as follows:

[0159] (1)Single-intention data

[0160] Single-intention data can roughly account for 81%. Rewrite the current user's question according to business requirements so that after rewriting the session message of the current round, it can include key elements such as the above-mentioned business requirements, time, amount, etc. At the same time, some invalid remarks such as complaints do not need to be rewritten and can be directly output.

[0161] (2)Multi-intention data

[0162] In the actual application process, during the interaction, users often generate new questions. For example, if the historical session message is "Why hasn't the withdrawal arrived yet" and the current session message is "How to retrieve the payment password", in this case, when labeling, the historical session message cannot be retained, and the current session message can be directly output.

[0163] (3)Third-party data

[0164] In order to maintain the generality of the model's ability in the rewriting task, some Chinese multi-round dialogue chat data can be used as a third-party dataset to supplement the training data.

[0165] After obtaining the labeled data, the Baichuan-13B-Chat base model can be selected, and using the text understanding and generation capabilities of the base model, the rewriting ability can be built, and tens of thousands of manually labeled examples can be imported. The base model is fine-tuned through the P-Tuning method to obtain a multi-round dialogue rewriting model.

[0166] Among them, during the training process, the input of the model can include sample data. The large model can rewrite the sample data according to the set sentence rewriting instruction to obtain the predicted sentence. For example, the sentence rewriting instruction "instruction" can be: "Please extract the key information in the historical question to complete the current question". The input "input" can be: "Current question: I don't have much money either\n; Historical question 1: My WeChat change\n; Historical question 2: Why is there a sudden limit?", and the output "output": "I don't have much change either. Why is there a sudden limit?".

[0167] Among them, instruction represents the instruction, instruction + input is the model input, and output is the model output.

[0168] Among them, the large model can summarize multiple descriptions of the user, reducing the interference caused by similar, redundant, and colloquial content in the user's segmented descriptions to the intention recognition and processing; and it can distinguish multiple questions of the user, discard the above text of non-current questions, and generate more standardized text. When it is determined that the user's intention has not changed, for the user's more detailed follow-up questions about this problem based on the Q&A database, that is, the Redis vector library, short text semantic matching recognition is performed to give a targeted short reply, improving the anthropomorphic degree of the conversation. Among them, the Redis vector library is also called a vector database, which is a type of database that stores data in the mathematical representation of vectors or data points.

[0169] It should be understood that although each step in the flowchart involved in the above-mentioned embodiments is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0170] Based on the same inventive concept, the embodiments of the present application also provide a message processing device for implementing the message processing method involved above. The implementation solution provided by this device to solve problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the message processing device provided below can refer to the limitations on the message processing method in the above text, and will not be repeated here.

[0171] In one embodiment, asFigure 9 As shown, a message processing device is provided, including: an acquisition module 902, an output module 904, an identification module 906, and a feedback module 910, where:

[0172] The acquisition module 902 is configured to acquire multi-round session messages from a user terminal during an interaction process. The multi-round session messages include the current session message and historical session messages before the current session message.

[0173] The output module 904 is configured to input the current session message and the historical session messages into a large model together, and generate a sentence through the large model.

[0174] The identification module 906 is configured to perform intent recognition based on the generated sentence to obtain a first intent corresponding to the current session message.

[0175] The identification module 908 is further configured to, for any historical session message, obtain a second intent corresponding to the targeted historical session message.

[0176] The feedback module 910 is configured to, when the first intent matches any of the second intents successfully, determine a reply content based on the first intent and the current session message, and feedback the reply content.

[0177] In some embodiments, the message processing device further includes a prompt content feedback module;

[0178] The prompt content feedback module is configured to, when the first intent does not match any of the second intents, determine a prompt content based on the first intent and feedback the prompt content; the prompt content is used to guide the next round of session.

[0179] In some embodiments, the message processing device further includes a storage module;

[0180] The storage module is configured to store the first intent corresponding to the current session message; when the next round of session message arrives, the current session message becomes a historical session message, and the stored first intent corresponding to the current session message becomes the second intent corresponding to the historical session message when processing a new session message.

[0181] In some embodiments, the generated sentence includes a question sentence; the output module 904 is further configured to input the current session message and the historical session messages into the large model according to a preset input format; through the large model, according to a preset sentence rewriting instruction, summarize the current session message and the historical session messages to obtain a summarized question sentence and output it.

[0182] In some embodiments, the recognition module 906 is further configured to perform word segmentation on the generated statement to obtain keywords in the generated statement; perform vectorization representation processing on the keywords to obtain feature representations of the keywords; and perform intent classification based on the feature representations to obtain a first intent corresponding to the current session message.

[0183] In some embodiments, the feedback module 910 is further configured to obtain a Q&A database, where the Q&A database includes multiple intents and Q&A vector pairs under each intent; determine a target intent matching the first intent in the Q&A database; and determine a reply content based on the current session message and the Q&A vector pairs under the target intent.

[0184] In some embodiments, the feedback module 910 is further configured to encode the current session message to obtain an encoded vector; calculate the similarity between the encoded vector and each question vector under the target intent; and determine a reply content based on the reply vector associated with the question vector corresponding to the similarity that satisfies a preset similarity condition.

[0185] In some embodiments, the message processing device further includes a model training module;

[0186] The model training module is configured to obtain multi-round session message samples; the multi-round session message samples include at least one of multi-intent session messages, single-intent session messages, or third-party session messages; obtain annotated statements obtained by rewriting the multi-round session message samples; summarize the multi-round session message samples through a large model to be trained to obtain predicted statements; and adjust the parameters of the large model to be trained based on the predicted statements and the annotated statements and continue training until the training is completed to obtain a trained large model.

[0187] In some embodiments, the model training module is further configured to rewrite each multi-round session message sample according to a preset rewriting condition to obtain a rewritten annotated statement; the preset rewriting condition includes at least one of keyword retention, only referring to the previous message during rewriting, typo correction, invalid information filtering, or independent sentence formation for non-strongly related questions.

[0188] In some embodiments, the model training module is further configured to input the current session message and the historical session message in the multi-round session message samples into the large model to be trained according to a preset input format; and summarize the questions in the multi-round session message samples through the large model to be trained according to a preset statement rewriting instruction to obtain a summarized predicted statement and output it.

[0189] In some embodiments, the model training module is further configured to determine the statement loss between the predicted statement and the annotated statement; and adjust the parameters of the large model to be trained according to the statement loss and continue training until the training is completed to obtain a trained large model.

[0190] Each module in the above message processing device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0191] In one embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multi-round session data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a message processing method.

[0192] Those skilled in the art can understand that Figure 10 the structure shown in

[0193] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.

[0194] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0195] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0196] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0197] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0198] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0199] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A message processing method, characterized in that, The method includes: Obtaining multi-round conversation messages from a user terminal during an interaction, where the multi-round conversation messages include the current conversation message and historical conversation messages before the current conversation message; Inputting the current conversation message and the historical conversation messages into a large model together, and generating a sentence through the large model; Performing intent recognition based on the generated sentence to obtain a first intent corresponding to the current conversation message; For any historical conversation message, obtaining a second intent corresponding to the targeted historical conversation message; When the first intent matches any of the second intents successfully, determining a reply content based on the first intent and the current conversation message, and feeding back the reply content.

2. The method according to claim 1, wherein The method further includes: When the first intent does not match any of the second intents, determining a prompt content based on the first intent, and feeding back the prompt content; the prompt content is used to guide the next round of conversation.

3. The method according to claim 1, wherein The method further includes: Storing the first intent corresponding to the current conversation message; When the next round of conversation message arrives, the current conversation message becomes a historical conversation message, and the stored first intent corresponding to the current conversation message becomes a second intent corresponding to the historical conversation message when processing the new conversation message.

4. The method according to claim 1, wherein The generated sentence includes a question sentence. The step of inputting the current conversation message and the historical conversation messages into the large model together and generating a sentence through the large model includes: Inputting the current conversation message and the historical conversation messages into the large model according to a preset input format; Through the large model, summarizing the questions for the current conversation message and the historical conversation messages according to a preset sentence rewriting instruction, obtaining a summarized question sentence and outputting it.

5. The method according to claim 1, wherein The step of performing intent recognition based on the generated sentence to obtain a first intent corresponding to the current conversation message includes: Performing word segmentation on the generated sentence to obtain keywords in the generated sentence; Performing vectorization representation processing on the keywords to obtain feature representations of the keywords; Performing intent classification based on the feature representations to obtain a first intent corresponding to the current conversation message.

6. The method according to claim 1, characterized in that, The step of determining a reply content based on the first intent and the current conversation message includes: Obtaining a question-and-answer database, where the question-and-answer database includes multiple intents and question-and-answer vector pairs under each intent; Determining a target intent that matches the first intent in the question-and-answer database; Determining a reply content based on the current conversation message and the question-and-answer vector pair under the target intent.

7. The method according to claim 6, wherein The question-and-answer vector pair includes a question vector. The step of determining a reply content based on the current conversation message and the question-and-answer vector pair under the target intent includes: Encoding the current conversation message to obtain an encoded vector; Calculating the similarity between the encoded vector and each question vector under the target intent; Determining a reply content based on the reply vector associated with the question vector corresponding to the similarity that meets a preset similarity condition.

8. The method according to any one of claims 1 to 7, characterized in that The training steps of the large model include: Obtain multi-round conversation message samples; the multi-round conversation message samples include at least one of multi-intention conversation messages, single-intention conversation messages, or third-party conversation messages; Obtain the annotated sentences obtained by rewriting the multi-round conversation message samples; Based on the multi-round conversation message samples, summarize through the large model to be trained to obtain predicted sentences; Based on the predicted sentences and the annotated sentences, adjust the parameters of the large model to be trained and continue training until the training is completed to obtain a trained large model.

9. The method according to claim 8, wherein The obtaining of the annotated sentences obtained by rewriting the multi-round conversation message samples includes: According to the preset rewriting conditions, rewrite each multi-round conversation message sample to obtain the rewritten annotated sentences; the preset rewriting conditions include at least one of keyword retention, only referring to the previous message during rewriting, typo correction, invalid information filtering, or separating non-strongly related questions into sentences.

10. The method according to claim 8, wherein The summarizing through the large model to be trained based on the multi-round conversation message samples to obtain predicted sentences includes: Input the current conversation message and the historical conversation messages in the multi-round conversation message samples into the large model to be trained according to the preset input format; Through the large model to be trained, summarize the questions for the multi-round conversation message samples according to the preset sentence rewriting instructions, obtain the summarized predicted sentences and output them.

11. A message processing device, characterized in that, The device includes: An acquisition module for acquiring multi-round conversation messages from a user terminal during an interaction, the multi-round conversation messages including a current conversation message and historical conversation messages before the current conversation message; An output module for jointly inputting the current conversation message and the historical conversation messages into the large model and outputting generated sentences through the large model; An identification module for performing intention identification based on the generated sentences to obtain a first intention corresponding to the current conversation message; The identification module is further configured to obtain a second intention corresponding to any historical conversation message for any historical conversation message; A feedback module for, when the first intention matches any second intention successfully, determining a reply content based on the first intention and the current conversation message and feeding back the reply content.

12. The device according to claim 11, wherein The output module is further configured to input the current conversation message and the historical conversation messages into the large model according to the preset input format; through the large model, summarize the questions for the current conversation message and the historical conversation messages according to the preset sentence rewriting instructions, obtain the summarized question sentences and output them.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 10.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Multi-agent-based unmanned aerial vehicle autonomous flight control method and server

    CN120993955A