Multi-round dialogue processing method, model training method and device

By extracting and comparing feature vectors on multiple rounds of dialogue data, the accuracy of identifying intentions in multiple rounds of dialogue is solved, which improves processing efficiency and accuracy and improves user experience.

CN120163168APending Publication Date: 2025-06-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510322352.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In human-computer interaction, there may be unclear pronouns, omitted key information or partial word errors in multiple rounds of dialogue, making it difficult for human-computer interaction products to accurately identify the user's conversation intentions and affect the user experience.

Method used

By processing the multi-round dialogue data, the context feature vectors of the multi-round dialogue are extracted and compared with the output feature vectors related to the downstream tasks until the preset comparison conditions are reached to determine the recognition results of the multi-round dialogue.

Benefits of technology

It improves the accuracy and efficiency of multiple rounds of conversations, can more accurately identify users' conversation intentions and provide accurate question-and-answer matching, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163168A_ABST
    Figure CN120163168A_ABST
Patent Text Reader

Abstract

The invention provides a multi-round dialogue processing method and device and a model training method and device, and relates to the field of artificial intelligence, in particular to the field of natural language processing. The specific implementation scheme is as follows: in response to multi-round dialogue data, performing first processing on the multi-round dialogue data to obtain a first context feature vector of a multi-round dialogue, so that a comparison result between the first context feature vector and a first output feature vector reaches a preset comparison condition, the first output feature vector is a feature vector about a downstream task extracted based on multi-round dialogue data; and determining an identification result of the multi-round dialogue according to the first context feature vector. Through the method, the processing accuracy and efficiency of multiple rounds of conversations can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to natural language processing technology in artificial intelligence, and particularly to a method for processing multi-turn conversations, a method for training a model, and a device therefor. Background Art

[0002] With the rapid development of artificial intelligence technology, the application of artificial intelligence technology in various fields is becoming more and more extensive, and more and more intelligent products based on human-computer interaction technology have emerged, such as intelligent customer service systems, embedded interaction devices, and so on.

[0003] When a user uses a human-computer interaction product for multi-turn conversations, there may be situations such as unclear pronouns (such as it, he, she, etc.), omission of key information, or errors in some words, which will make it difficult for the human-computer interaction product to accurately retrieve the user's dialogue Q&A or accurately identify the user's dialogue intention, thus affecting the user experience. Summary of the Invention

[0004] The present disclosure provides a method for processing multi-turn conversations, a method for training a model, and a device therefor with higher efficiency.

[0005] According to a first aspect of the present disclosure, there is provided a method for processing multi-turn conversations, including:

[0006] In response to multi-turn conversation data, performing a first processing on the multi-turn conversation data to obtain a first context feature vector of the multi-turn conversation, such that a comparison result between the first context feature vector and a first output feature vector reaches a preset comparison condition, wherein the first output feature vector is a feature vector regarding a downstream task extracted based on the multi-turn conversation data;

[0007] Determining an identification result of the multi-turn conversation according to the first context feature vector.

[0008] According to a second aspect of the present disclosure, there is provided a method for training a model, including:

[0009] Inputting historical multi-turn conversation data into an initial model for processing multi-turn conversations to extract a second context feature vector of the historical multi-turn conversation and a second output feature vector of the historical multi-turn conversation regarding a downstream task;

[0010] Training the initial model according to the second context feature vector and the second output feature vector;

[0011] When a comparison result between the second context feature vector and the second output feature vector reaches a preset comparison condition, obtaining a model for processing multi-turn conversations.

[0012] According to a third aspect of the present disclosure, there is provided a multi-turn dialogue analysis device, including:

[0013] A first processing unit, configured to, in response to multi-turn dialogue data, perform first processing on the multi-turn dialogue data to obtain a first context feature vector of the multi-turn dialogue, such that a comparison result between the first context feature vector and a first output feature vector meets a preset comparison condition, where the first output feature vector is a feature vector regarding a downstream task extracted based on the multi-turn dialogue data;

[0014] An identification unit, configured to determine an identification result of the multi-turn dialogue according to the first context feature vector.

[0015] According to a fourth aspect of the present disclosure, there is provided a model training method, including:

[0016] An input unit, configured to input historical multi-turn dialogue data into an initial model for processing multi-turn dialogue to extract a second context feature vector of the historical multi-turn dialogue and a second output feature vector of the historical multi-turn dialogue regarding a downstream task;

[0017] A training unit, configured to train the initial model according to the second context feature vector and the second output feature vector;

[0018] An obtaining unit, configured to obtain a model for processing multi-turn dialogue when a comparison result between the second context feature vector and the second output feature vector meets a preset comparison condition.

[0019] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0020] At least one processor; and

[0021] A memory communicatively connected to the at least one processor; where

[0022] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the methods provided in the first aspect and / or the second aspect above.

[0023] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the methods provided in the first aspect and / or the second aspect above.

[0024] According to a seventh aspect of the present disclosure, there is provided a computer program product, which includes: a computer program stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to enable the electronic device to execute the method described in the first aspect.

[0025] The technology according to the present disclosure can effectively improve the processing accuracy and efficiency of multi-turn conversations.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0028] Figure 1 is a schematic flowchart of a method for processing multi-turn conversations provided by an embodiment of the present disclosure;

[0029] Figure 2 is a schematic flowchart of another method for processing multi-turn conversations provided by an embodiment of the present disclosure;

[0030] Figure 3 is a schematic flowchart of yet another method for processing multi-turn conversations provided by an embodiment of the present disclosure;

[0031] Figure 4 is a schematic flowchart of a model training method provided by an embodiment of the present disclosure;

[0032] Figure 5 is a schematic diagram of a possible scenario provided by an embodiment of the present disclosure;

[0033] Figure 6 is a schematic structural diagram of a device for processing multi-turn conversations provided by an embodiment of the present disclosure;

[0034] Figure 7 is a schematic structural diagram of a model training device provided by an embodiment of the present disclosure;

[0035] Figure 8 is one of the block diagrams of an electronic device for implementing the method for processing multi-turn conversations and / or of the present disclosure;

[0036] Figure 9 is another block diagram of an electronic device for implementing the method for processing multi-turn conversations and / or of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0038] The semantic analysis solution for multi-turn conversations provided by the embodiments of the present disclosure can be applied to the photo album APPs or mini-programs of various types of terminals. Among them, various types of terminals may include, but are not limited to, computers, smart phones, tablet computers, e-book readers, Moving Picture experts group audio layer III (MP3) players, Moving Picture experts group audio layer IV (MP4) players, portable computers, in-vehicle computers, wearable devices, desktop computers, set-top boxes, smart TVs, and so on. Specifically, it can be applied to intelligent customer service systems, such as multi-turn conversation intent recognition and Frequently Asked Questions (FAQ) matching in banking, e-commerce, and government affairs scenarios. Or, it can be applied to embedded interactive devices, such as context-aware instruction understanding of smart speakers and in-vehicle voice assistants.

[0039] In multi-turn conversations and multi-turn retrievals of human-computer interaction products such as intelligent customer service systems or embedded interactive devices, there may be pronouns such as "it", "he", "she", etc. in the user's conversation request (query), or there may be cases of omission of key information, or there may be cases of partial word errors. These problems will result in inaccurate retrieval results during the retrieval process; or inaccurate recognition results during the intent recognition process. In related technologies, mainly through large models (generative large models, such as Ernie series models, etc.) to rewrite and enhance the user's conversation (query) to solve pronoun disambiguation, information completion, context correction, multi-turn conversation clarification, and Query enhancement, so as to improve the accuracy of conversation retrieval results or the accuracy of conversation intent recognition. Since each user query needs to be rewritten multiple times, a large amount of video memory is required when rewriting multi-turn conversations, which requires very high computing power for large models and has a high response latency, and usually it is difficult to support high-concurrency scenarios.

[0040] The present disclosure provides a method for processing multi-turn conversations, a method for model training, and a device, which are applied to the natural language processing technology in artificial intelligence and can effectively improve the accuracy and efficiency of processing multi-turn conversations.

[0041] It should be noted that the head model in this embodiment is not a model for a specific user and does not reflect the personal information of a specific user. It should be noted that the multi-turn conversations in this embodiment come from a public dataset.

[0042] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0043] To enable readers to more deeply understand the implementation principle of the present disclosure, the following Figures 1 - 9 is used to further refine the embodiments of the present disclosure.

[0044] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for processing multi-turn conversations provided by an embodiment of the present disclosure, and can be applied to an electronic device. The method includes steps S101-S103:

[0045] Step S101: In response to the multi-turn conversation data, perform a first processing on the multi-turn conversation data to obtain a first context feature vector of the multi-turn conversation, so that the comparison result between the first context feature vector and the first output feature vector reaches a preset comparison condition, where the first output feature vector is a feature vector about the downstream task extracted based on the multi-turn conversation data.

[0046] In this embodiment, the multi-turn conversation data can be real-time multi-turn conversation data input by the user to the electronic device in real time (which can be in the form of voice input, text input, etc.). Optionally, the electronic device can perform preprocessing on the multi-turn conversation data (such as the sliding window truncation processing, prompt splicing processing, etc. mentioned in the following embodiments, or other methods, such as text cleaning: deleting characters irrelevant to the conversation processing, such as punctuation marks, special symbols, and redundant spaces, etc.), and then process the preprocessed multi-turn conversation data to further improve the processing accuracy and efficiency of the multi-turn conversation data.

[0047] In this embodiment, the multi-turn dialogue data is subjected to a first process to convert the multi-turn dialogue data into a context vector representation, that is, the first context feature vector, which represents the overall semantics and context information. Exemplarily, the multi-turn dialogue data can be input into a model (such as a pre-trained embedding model), and the model processes the first data to output the first context feature vector of the multi-turn dialogue. During the training process of the model, a contrastive learning branch can be added between the context feature vector of the multi-turn dialogue data and the output feature vector of the downstream task (for example, training to make the similarity between the context feature vector and the output feature vector reach a preset threshold, which can be determined by those skilled in the art according to empirical values). When the model is applied, the comparison result between the context feature vector and the output feature vector of the multi-turn dialogue output meets a preset comparison condition (such as the similarity reaches a preset threshold).

[0048] In some embodiments, in addition to performing the first process by inputting into the model, the first process can also be performed in other ways. For example, the bag-of-words (BoW) model can be used, or the graph embedding technology (such as Node2Vec) can be used to convert the graph structure into a feature vector, and so on.

[0049] In this embodiment, the downstream task can be a specific application task to be solved after the extraction of the context feature vector. For example, intent recognition, or FAQ matching, or multi-turn dialogue rewriting, and so on.

[0050] In an alternative embodiment, the first output feature vector of the multi-turn dialogue can include one or more of the following: the sentence feature vector after processing the multi-turn dialogue data; the question feature vector corresponding to the first context feature vector; the intent recognition vector corresponding to the first context feature vector.

[0051] Among them, the first output feature vector is the feature vector obtained by processing the downstream task. The sentence feature vector after processing the multi-turn dialogue data can be the sentence feature vector after rewriting the multi-turn dialogue data. For example, it can be the sentence feature vector rewritten by using an existing generative large model, or the sentence feature vector obtained by other processing methods.

[0052] It can be understood that the first output feature vector and the second output feature vector in the following text are only used to distinguish similar objects and have no other special meanings. Among them, the first output feature vector and the second output feature vector can be the same feature vector or different feature vectors.

[0053] It can be understood that the first output feature vector or the second output feature vector hereinafter can be a feature vector generated based on multi-round dialogue data or multi-round dialogue data and used to measure the quality of the context feature vector.

[0054] In some embodiments, the above-mentioned first output feature vector can also be other feature vectors, which are determined according to the downstream task.

[0055] Exemplarily, the first output feature vector regarding the downstream task generated based on the multi-round dialogue data can be the first output feature vector regarding the downstream task directly generated using the multi-round dialogue data. For example, the above-mentioned first output feature vector can be the sentence after rewriting the multi-round dialogue output after inputting the multi-round dialogue data into the generative large model, and encoded by the encoding module into the sentence feature vector after processing the multi-round dialogue data.

[0056] By performing a first processing on the multi-round dialogue data to obtain the first context feature vector of the multi-round dialogue, such that the comparison result between the first context feature vector and the first output feature vector regarding the downstream task generated based on the multi-round dialogue data reaches a preset comparison condition, this means that the obtained context feature vector is more similar to the feature vector corresponding to the downstream task, effectively improving the accuracy of the context feature vector, so that the generated context feature vector can more accurately reflect the semantics and intentions of the dialogue, and thus perform better in the downstream task.

[0057] In other words, during the process of processing the multi-round dialogue data, using the output feature vector regarding the downstream task to measure the context feature vector in the data processing result can effectively solve the problem of context association and improve the extraction accuracy of the context feature vector.

[0058] It can be understood that "in response to" is used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations to be executed can be real-time or have a set delay; without special instructions, there is no limit on the execution order of multiple operations to be executed.

[0059] Step S102, determine the recognition result of the multi-round dialogue according to the first context feature vector.

[0060] In this embodiment, the process of recognizing the multi-round dialogue in combination with the context feature vector can perform different recognition processes according to different downstream tasks. For example, the downstream tasks can include FAQ matching, intent recognition, multi-round dialogue rewriting, and so on.

[0061] In the embodiments of the present disclosure, instead of directly rewriting multi-turn conversations using a generative large model, by processing multi-turn conversation data into context feature vectors, the multi-turn conversation data processing process can make the comparison result between the context feature vectors and the first output features of the multi-turn conversation data regarding the downstream task reach a specific comparison condition, thereby improving the semantic accuracy of the context feature vectors. Even without conversation rewriting, the user intention can be recognized in the downstream task or an answer can be matched for the user.

[0062] In an alternative embodiment, the recognition result of the multi-turn conversation may include the answer recognition result of the multi-turn conversation. Taking the FAQ matching as the downstream task as an example, the semantic recognition process may be to compare the similarity with the question vectors in the question and answer library (FAQ) to obtain a more accurate conversation answer.

[0063] Exemplarily, the method may further include the following steps: extracting the question feature vectors corresponding to each question in the question and answer library, and constructing a preset search structure according to the question feature vectors, where the search structure is used to query the question feature vector closest to the first context feature vector.

[0064] The FAQ question and answer library is a system specifically designed to automatically answer common questions from users. The FAQ question and answer library usually contains a set of predefined question and answer pairs, aiming to understand the user's query through natural language processing technology and provide corresponding answers.

[0065] In this embodiment, by extracting the question feature vectors corresponding to each question in the question and answer library, the method of extracting the question feature vectors corresponding to each question may adopt the same processing method as the first processing method or a different processing method. For example, the bag-of-words model can be used to extract the question feature vectors corresponding to the questions. Then, a search structure is constructed, and the search structure can be used to store and organize all the question feature vectors in the question and answer library. Among them, the search structure may adopt a Hierarchical Navigable Small World (HNSW) index structure. Among them, the HNSW index structure is a data structure and algorithm for efficient approximate nearest neighbor search (ANNS).

[0066] The above-mentioned step S102 determines the recognition result of the multi-turn dialogue according to the first context feature vector, and can adopt the following method: perform a nearest neighbor search on the first context feature vector in the search structure to obtain a preset number of candidate question feature vectors; determine the target question feature vector according to the cosine similarity between the first context feature vector and the candidate question feature vectors; and determine the target question from the Q&A library according to the first context feature vector, and obtain the answer recognition result corresponding to the multi-turn dialogue according to the target question.

[0067] In this embodiment, a nearest neighbor search is performed on the context feature vector and the questions in the Q&A library. In the search structure, the context feature vector is used for nearest neighbor search to find a preset number (for example, the first K) of question feature vectors that are most similar to the context feature vector as candidate question feature vectors. Then, for each candidate question feature vector, calculate the cosine similarity between it and the context feature vector. It can be understood that the cosine similarity is a commonly used similarity measure that measures the angle between two vectors, and the closer the value is to 1, the more similar they are. Next, according to the calculated cosine similarity, select the candidate question feature vector with the highest similarity (one or more, for example, select the Top-3 (similarity ≥ threshold) candidate question feature vectors as the target question feature vectors. These one or more target question feature vectors are the most matched with the context feature vector and represent the most likely relevant questions. Since the target question feature vector is the vector corresponding to the question with the highest matching similarity, the target question can be quickly determined using this vector, and thus the answer recognition result corresponding to the target question can be extracted.

[0068] Through the above technical solution, the method of converting the context feature vector into a question feature vector that matches the FAQ Q&A system can effectively solve the problem of fast and accurate matching of multi-turn dialogue answering when the downstream task is FAQ Q&A.

[0069] In another optional implementation manner, the downstream task can also be intent recognition, and the recognition result of the multi-turn dialogue includes the intent recognition result of the multi-turn dialogue. Optionally, the process of the above-mentioned step S102 for determining the recognition result of the multi-turn dialogue according to the first context feature vector can be:

[0070] Input the first context feature vector into a preset intent recognition model for recognizing dialogue intent, where the intent recognition model includes a support vector machine; process the first context feature vector based on the intent recognition model to obtain the intent recognition result of the multi-turn dialogue.

[0071] In some embodiments, the first context feature vector can be processed first (such as dimensionality reduction processing, feature selection processing, etc.) into a vector suitable for processing by the intent recognition model, so as to facilitate the processing of the intent recognition model.

[0072] Optionally, the intent recognition model can include a Support Vector Machine (SVM), also known as an SVM classifier. The support vector machine is a machine learning model for classification tasks, which is good at processing high-dimensional data and complex classification problems. It realizes data classification by finding an optimal hyperplane to distinguish data points of different categories. Specifically, the SVM calculates the position of the vector in the feature space and evaluates its distance and direction relative to the classification hyperplane. In this embodiment, the SVM classifier can process the input context feature vector. Based on the calculation results of the SVM, the model will classify the context feature vector into a certain intent category. The SVM classifies by maximizing the margin between different categories, thereby improving the accuracy and robustness of classification. And it outputs the classification result, which is usually a class label indicating the recognized user intent, such as "query weather", "book a restaurant", etc. Thus, the information in the context feature vector is effectively utilized to accurately identify the user intent in the conversation. This method is very useful in dialogue systems, customer service robots and other natural language processing applications, especially in scenarios that require processing complex semantics and context.

[0073] It can be seen that the above FAQ question and answer and intent recognition process can omit the process of multi-round dialogue rewriting and obtain relatively accurate multi-round dialogue recognition results. This effectively solves the problems of excessive computing power and high latency of the generative large model in the related technology, and further achieves the purpose of improving the processing accuracy and efficiency of multi-round dialogue.

[0074] Figure 2 It is a schematic flowchart of another multi-round dialogue processing method provided by the embodiments of the present disclosure. On the basis of the above embodiments, to further improve the processing efficiency and accuracy of multi-round dialogue, the first processing method can adopt model processing. Specifically, in addition to the above steps S101 and S102, the above step S101 is further divided into step S1011 and step S1012.

[0075] Step S1011: Input the multi-turn conversation data into a pre-set model for processing multi-turn conversations. The model is trained based on historical multi-turn conversation data and meets a preset comparison condition between the second context feature vector and the second output feature vector. Among them, the second context feature vector is the feature vector of the historical multi-turn conversation extracted based on the historical multi-turn conversation data, and the second output feature vector is the feature vector regarding the downstream task extracted based on the historical multi-turn conversation data.

[0076] Optionally, the above model can be a model trained using an embedding model. It can be understood that the embedding model is a technology that maps high-dimensional data to a low-dimensional space. It can convert high-dimensional data such as text and pictures into low-dimensional vectors, enabling the computer to better understand and process this data.

[0077] Among them, the second context feature vector extracted based on the historical multi-turn conversation data, that is, the second context feature vector of the historical multi-turn conversation extracted by inputting the historical multi-turn conversation data into the model, and the second output feature vector regarding the downstream task extracted (or generated) based on the historical multi-turn conversation can be the second output feature vector output to the downstream task after the model outputs the context feature vector, or can be the second output feature vector of the downstream task obtained by directly processing the historical multi-turn conversation data using other models or algorithms. The comparison result can be the similarity (such as cosine similarity) between the context feature vector corresponding to the historical multi-turn conversation and the second output feature vector regarding the downstream task. The higher the similarity, the more accurate the semantics represented by the context feature vector output after model processing. When the comparison result can meet the preset comparison condition (such as the similarity reaches a preset similarity threshold, and those skilled in the art can adaptively determine it according to empirical values), it indicates that the current model has a good processing effect.

[0078] Through the above model processing method, the context feature vector corresponding to the multi-turn conversation data to be processed can be efficiently output, and the comparison result between the output context feature vector and the output feature vector can reach the preset comparison condition.

[0079] It should be noted that for the unstated part of the training process of the above model, reference can be made to the model training embodiments in the following text, and no more details will be elaborated here.

[0080] Step S1012: Perform a first processing on the multi-turn conversation data according to the model to obtain the first context feature vector of the multi-turn conversation.

[0081] Optionally, the model can process the input multi-turn conversation through its architecture (such as neural network layers), capture the semantic and context information in the conversation, and generate the context feature vector of the conversation during the processing. Based on the processing of the model, it can effectively capture the important features and relationships in the conversation, omit the multi-turn conversation rewriting method, and improve the conversation recognition accuracy.

[0082] In some embodiments, considering the limitations of the model itself (such as insufficient processing power or attention shift) in the related art, it is difficult to focus on important conversations, especially the global attention to long texts is unbalanced and vulnerable to redundant conversation interference. In this embodiment, by designing an encoding module and an output module to encode the conversation separately, the processing accuracy of the conversation can be further improved. Specifically, the model can include an encoding module and an output module. The encoding module is used to encode a preset number of turns of conversation separately, and the output module is used to perform weighted average calculation on the encoding result of the encoding module and then output.

[0083] The above step S1012 performs a first processing on the multi-turn conversation according to the model to obtain the first context feature vector of the multi-turn conversation, which can be implemented as follows: according to the encoding module, encode the preset number of turns of conversation in the multi-turn conversation data separately to extract the feature vectors of the corresponding turns of conversation in the multi-turn conversation; perform weighted average calculation on the feature vectors of the preset number of turns of conversation according to the output module to obtain the first context feature vector of the multi-turn conversation.

[0084] Optionally, the encoding module can be a word embedding layer that maps the conversation information to a fixed-dimensional vector space to capture the semantic relationship between conversations.

[0085] Optionally, the preset number of turns can be set according to actual needs or empirical values for the conversation turns that need to be encoded separately. For example, each turn of conversation can be encoded separately, or it can be set to encode every other turn or multiple turns of conversation. This embodiment does not make a special limitation on this. Encoding the preset number of turns of conversation separately combines more conversation encoding results for semantic recognition, which can effectively avoid the error caused by only focusing on a specific conversation. For example, through the embodiments of the present application, the problem of losing the key clues in the conversation history caused by simply using the Embedding of the last sentence of the conversation in the related art can be effectively solved. In addition, through the predefined way of conversation turns, it supports users to design and encode the conversation turns, while considering the problems such as insufficient processing power of the model itself.

[0086] Optionally, the output module may be a weighted average layer. For example, different weights (such as attention weights) are assigned to the preset round of conversations respectively, and the contribution of the feature vectors corresponding to the preset round of conversations is dynamically adjusted. For example, in an example where each round of conversation is encoded separately, different weights can be assigned to each round of conversation, so that the finally output context feature vector can focus on the key conversations, reduce the interference problem of redundant conversations, and improve the dialogue processing accuracy of the model.

[0087] It should be noted that in this embodiment, the separate encoding of the preset round of conversations may be the separate encoding of one or more conversation rounds in the predefined round of conversations.

[0088] Further exemplarily, among them, the weighted value of the preset round of conversations is determined according to the hint information carried in the multi-round conversation data, or is determined according to the sequence of the multi-round conversation data.

[0089] For example, the multi-round conversation data may be input into the model after being concatenated by prompts (promt). The model can obtain the hint information according to this concatenation information and use this hint information to weight specific conversations (such as the last round of conversation is the most important, and a higher weight, such as 0.5, is given to the last round of conversation, which can be adjusted according to empirical values or actual applications). Or for example, it is determined according to the sequence (time) of the conversations in the multi-round conversation data. The closer to the current time (i.e., the more recent), the higher the weight, and the farther from the current time (i.e., the more distant), the lower the weight, and so on. Optionally, when the weights are assigned, the sum of the weights corresponding to each round is 1.

[0090] In the above solution, by separately encoding the preset round of conversations and assigning corresponding weights to the preset round of conversations, the semantics can be recognized in combination with historical conversations, while focusing on the key conversations, reducing the interference problem of redundant conversations, and improving the dialogue processing accuracy of the model. This process effectively solves the problem of low recognition accuracy caused by the fact that in the related art, the attention weight of the embedding model for the important conversations of the user (such as the last round of conversation) is less than 30%, or the embedding model only processes the last round of conversation.

[0091] Figure 3 It is a schematic flowchart of another processing method for multi-round conversations provided by an embodiment of the present disclosure. On the basis of the Figure 1 corresponding embodiment, in this embodiment, before the first processing of the multi-round conversation data, by preprocessing the multi-round conversation data, the dialogue processing accuracy can be further improved. Specifically, before the above step S101 performs the first processing on the multi-round conversation data, the following step S301 may further be included, and step S101 is further divided into step S1011.

[0092] Step S301: Preprocess the multi-turn dialogue data to obtain preprocessed multi-turn dialogue data, where the preprocessing method includes truncation processing and / or prompt splicing processing of the multi-turn dialogue data.

[0093] Exemplarily, for the above step S301 to preprocess the multi-turn dialogue data, the following method can be adopted:

[0094] Method 1: According to a preset sliding window, truncate the multi-turn dialogue data to retain the truncated multi-turn dialogue data; where the sliding window indicates the truncation time and / or the truncation quantity of the multi-turn dialogue.

[0095] Exemplarily, sliding window parameters can be predefined, such as the truncation time, which indicates how to select dialogue turns in the time dimension. For example, within a specific time range (such as the dialogue in the last 5 minutes) or a fixed time interval. Or, the truncation quantity, which indicates the maximum number of dialogue turns to be retained. For example, at most retain the last 10 turns of dialogue. And the sliding window can be initialized, such as determining the initial position of the sliding window. According to specific application requirements, it can be selected to start from the latest part of the dialogue (i.e., from the end forward) or from the earliest part (i.e., from the beginning backward).

[0096] Further exemplarily, the process of truncating the multi-turn dialogue using the sliding window can be to screen the dialogue turns within this time range according to the set truncation time. For example, if it is set to the last 5 minutes, only retain the dialogue that occurred within these 5 minutes. Or, on the basis of time truncation, further limit the retained dialogue turns according to the truncation quantity. For example, if it is set to at most 10 turns of dialogue, only retain the last 10 turns of dialogue data. Or, only display the retained dialogue turns according to the truncation quantity, and so on.

[0097] Through the above method of sliding window truncation, it can help the electronic device focus on the recent dialogue content, improve the response speed and relevance, and reduce the redundancy of unimportant dialogues.

[0098] Method 2: According to preset prompt information, splice the multi-turn dialogue data to obtain multi-turn dialogue data containing the prompt information; where the prompt information is used to indicate that the importance of at least one dialogue in the multi-turn dialogue is higher than that of other dialogues.

[0099] Exemplarily, the prompt information can be a short text segment to indicate the importance of certain dialogue turns. For example, "Pay special attention to the last turn of dialogue". Through the method of prompt information splicing, it is convenient to guide the electronic device or model to focus on specific dialogue content (such as the last turn of dialogue), and improve the accuracy and efficiency of multi-turn dialogue data processing.

[0100] Method 3: Preprocess multi-turn conversations by combining the above Method 1 and Method 2. For example, first truncate the multi-turn conversation data, and then splice the prompt information for the truncated multi-turn conversation data.

[0101] In some examples, in addition to the above preprocessing methods, other preprocessing methods can also be adopted, such as text cleaning: deleting characters irrelevant to conversation processing, such as punctuation marks, special symbols, and redundant spaces, etc.

[0102] Step S1011: Perform a first process on the multi-turn conversation data to obtain a first context feature vector of the multi-turn conversation, including: performing a first process on the preprocessed multi-turn conversation data to obtain a first context feature vector of the multi-turn conversation.

[0103] Combined with the above Figure 2 Corresponding to the embodiment, by inputting the truncated multi-turn conversation data or the multi-turn conversation data after prompt splicing into the model for processing, since preprocessing has been performed before the model, during the model processing, the model does not need to perform conversation truncation and can process important conversations according to the prompt. This can effectively solve the problems that the model randomly truncates conversations or simply splices historical statements, resulting in loss of key information, or it is difficult to focus on important conversations due to insufficient processing capabilities of the model and other situations after the multi-turn conversation is input into the model.

[0104] Further combined with the above Figure 2 Corresponding to the embodiment, in some scenarios, the preprocessing process can also be executed in the model, that is, the model can include a preprocessing module. This embodiment does not make special limitations on whether the preprocessing process is performed outside the model or inside the model.

[0105] Figure 4 This is a model training method provided by an embodiment of the present disclosure. As Figure 4 shown, this method includes steps S401 - S403.

[0106] Step S401: Input the historical multi-turn conversation data into an initial model for processing multi-turn conversations to extract a second context feature vector of the historical multi-turn conversation and a second output feature vector of the historical multi-turn conversation regarding the downstream task.

[0107] Optionally, the initial model can be an embedding model or other models that can be used to extract context feature data of multi-turn conversation data.

[0108] Exemplarily, historical multi-turn dialogue data can be collected, such as obtaining a large amount of historical dialogue data, which should be able to represent typical conversations in the target application scenario, and these historical dialogue data can be cleaned and formatted to ensure its suitability for model input. For example, noise can be removed, word segmentation, part-of-speech tagging, etc. By inputting the historical multi-turn dialogue data into the initial model for pre-training, feature vectors are extracted. For example, the model processes the input data to generate context feature vectors for the historical multi-turn dialogue, and these vectors capture the semantic and context information of the dialogue. And according to specific downstream tasks (such as FAQ matching, intent recognition, etc.), the model generates second output feature vectors, which can be used to represent the performance of the dialogue in the downstream tasks.

[0109] In some embodiments, the second output feature vector includes at least one of the following: the sentence feature vector after processing the historical multi-turn dialogue data; the question feature vector corresponding to the second context feature vector; the intent recognition vector corresponding to the second context feature vector.

[0110] It can be understood that the second output feature vector is similar to the first output feature vector in the above text, and no special limitation is made in the relevant description.

[0111] Step S402: Train the initial model according to the second context feature vector and the second output feature vector.

[0112] In this embodiment, the second context feature vector and the second output feature vector are used to train the model, and the training objective is to optimize the model parameters so that it can generate the context feature vector more accurately. Exemplarily, a suitable loss function design can be made to measure the gap between the model output and the expected output. For example, the loss function can combine the errors of the context feature and the downstream task feature.

[0113] It should be noted that the second context feature vector in this embodiment and the first context feature vector in the above text. The second context feature vector is a feature vector about the historical multi-turn dialogue extracted using the historical multi-turn dialogue data. In some embodiments, the first context feature vector and the second context feature vector can be the same feature vector.

[0114] In some embodiments, the initial model is trained by means of contrast learning. Specifically, the initial model includes a contrast learning branch for processing the contrast information between the context feature vector and the output feature vector, and the contrast learning branch includes a preset contrast loss function.

[0115] The above step S402 can train the initial model according to the second context feature vector and the second output feature vector in the following manner: according to the second context feature vector and the second output feature vector, train the initial model through the contrast loss function, so that the contrast loss function reaches a preset loss condition, and the loss condition is used to indicate that the contrast result between the second context feature vector and the second output feature vector reaches a preset contrast condition.

[0116] In this embodiment, a method of using contrastive learning to train a multi-turn dialogue processing model is provided.

[0117] Exemplarily, the contrast loss function can adopt contrastive loss, triplet loss, etc., and is used to measure the similarity between feature vectors. In the process of loss calculation, the difference between the second context feature vector and the second output feature vector can be calculated through the contrast loss function, and the goal is to minimize the loss between the two (to make them more similar).

[0118] Step S403, when the contrast result between the second context feature vector and the second output feature vector reaches a preset comparison condition, obtain a model for processing multi-turn dialogues.

[0119] When the preset contrast condition (such as the similarity reaches a set threshold) is reached between the second context feature vector and the second output feature vector, or when the preset maximum number of iterations (such as 50 times) is reached, the model training ends.

[0120] In some embodiments, inputting historical multi-turn dialogue data into the initial model for processing multi-turn dialogues includes: after preprocessing the historical multi-turn dialogue data, inputting it into the initial model, and the preprocessing method includes truncating the historical multi-turn dialogue data and / or prompt splicing processing.

[0121] It should be noted that the preprocessing method of the above historical multi-turn dialogue data is similar to the preprocessing method of multi-turn dialogues in the above model application process. For relevant descriptions, reference can be made to the embodiments of the multi-turn data processing method above, and details will not be elaborated here.

[0122] Figure 5 A possible application scenario provided by the embodiments of the present disclosure is as Figure 5 shown, including the following process:

[0123] Data collection: Collect multi-turn dialogue data through a voice interaction device. It is the original multi-turn dialogue text (such as the M-round interaction between the user and the customer service).

[0124] Data processing: Preprocess the multi-round dialogue data in the following ways:

[0125] 1) Sliding window truncation: Only keep the last N rounds of dialogues (configurable, N ≤ M), and discard the excess.

[0126] 2) Static Prompt concatenation: Use the prompt to prompt the model to focus on the last query of the user to enhance understanding according to the context.

[0127] Model training: Pretrain the context embedding model. Include the following ways:

[0128] 1) Model selection: You can select an embedding model, such as vector models like bge-embedding.

[0129] 2) In the embedding model: Pretrain the multi-round dialogue context; add a contrastive learning branch for the context embedding and the sentence embedding after multi-round rewriting for pre-training; encode each round of dialogue separately, and the weight of the last round is the highest when calculating the weighted average. This process is an optimization process for long texts of multi-round dialogues.

[0130] During the online inference stage, apply it to downstream tasks: The downstream task adaptation process can be as follows:

[0131] 1) FAQ matching: Precompute the FAQ library embedding, build an HNSW index; calculate the cosine similarity in real time and return the top-3 candidates (similarity ≥ threshold).

[0132] 2) Intent classification: Combine an SVM classifier, etc. for intent recognition.

[0133] Data feedback: Use the data in the inference stage (such as the input data and output data of the inference) to optimize the model.

[0134] The embodiments of the present disclosure improve the multi-round rewriting effect and reduce the computing power by combining a pre-trained embedding model with stronger generalization ability. In the modeling process, joint modeling of sentence discrimination (context semantics) and downstream tasks (such as sentence rewriting) is combined, that is, both the discrimination (context semantics) of the initial multi-round dialogue model and downstream tasks (such as sentence rewriting) are trained, improving the dialogue rewriting effect of the multi-round dialogue model, thereby improving the response speed of the overall process and reducing the hardware computing power cost in public cloud and private cloud scenarios.

[0135] Figure 6 It is a multi-round dialogue analysis device provided by the embodiments of the present disclosure. As Figure 6 shown, the device 600 includes a first processing unit 601 and an identification unit 602.

[0136] The first processing unit 601 is configured to perform a first processing on the multi-turn dialogue data in response to the multi-turn dialogue data, so as to obtain a first context feature vector of the multi-turn dialogue, such that a comparison result between the first context feature vector and a first output feature vector reaches a preset comparison condition, where the first output feature vector is a feature vector regarding a downstream task extracted based on the multi-turn dialogue data;

[0137] The recognition unit 602 is configured to determine a recognition result of the multi-turn dialogue according to the first context feature vector.

[0138] In an optional implementation manner, the first output feature vector includes at least one of the following: a sentence feature vector obtained by processing the multi-turn dialogue data; a question feature vector corresponding to the first context feature vector; an intention recognition vector corresponding to the first context feature vector.

[0139] In an optional implementation manner, the first processing unit 601 includes:

[0140] A first input module, configured to input the multi-turn dialogue data into a preset model for processing multi-turn dialogues, where the model is trained based on historical multi-turn dialogues and a comparison result between a second context feature vector and a second output feature vector reaches a preset comparison condition, where the second context feature vector is a feature vector of the historical multi-turn dialogue extracted based on historical multi-turn dialogue data, and the second output feature vector is a feature vector regarding a downstream task extracted based on historical multi-turn dialogue data;

[0141] A processing module, configured to perform a first processing on the multi-turn dialogue data according to the model to obtain the first context feature vector of the multi-turn dialogue.

[0142] In an optional implementation manner, the model includes an encoding module and an output module, where the encoding module is configured to separately encode a preset number of turns of dialogue, and the output module is configured to perform a weighted average calculation on the encoding result of the encoding module and then output;

[0143] The processing module is specifically configured to separately encode a preset number of turns of dialogue according to the encoding module to extract feature vectors of the preset number of turns of dialogue in the multi-turn dialogue data; perform a weighted average calculation on the feature vectors of the preset number of turns of dialogue according to the output module to obtain the first context feature vector of the multi-turn dialogue.

[0144] In an optional implementation manner, the weighting value of the preset number of turns of dialogue is determined according to the prompt information carried in the multi-turn dialogue, or determined according to the sequence of the multi-turn dialogue.

[0145] In an alternative embodiment, the recognition result of the multi-turn dialogue includes the answer recognition result of the multi-turn dialogue, and the apparatus further includes:

[0146] A search structure construction module, configured to extract the question feature vector corresponding to each question in the Q&A library, and construct a preset search structure according to the question feature vector, where the search structure is used to query the question feature vector closest to the first context feature vector;

[0147] The recognition unit 602 includes:

[0148] A search module, configured to perform a nearest neighbor search on the first context feature vector in the search structure to obtain a preset number of candidate question feature vectors;

[0149] A determination module, configured to determine a target question feature vector according to the cosine similarity between the first context feature vector and the candidate question feature vectors;

[0150] A first acquisition module, configured to determine a target question from the Q&A library according to the target question feature vector, and obtain the answer recognition result corresponding to the multi-turn dialogue according to the target question.

[0151] In an alternative embodiment, the recognition result of the multi-turn dialogue includes the intent recognition result of the multi-turn dialogue;

[0152] The recognition module 602 includes:

[0153] A second input module, configured to input the first context feature vector into a preset intent recognition model for recognizing dialogue intent, where the intent recognition model includes a support vector machine;

[0154] A second acquisition module, configured to process the first context feature vector based on the intent recognition model to obtain the intent recognition result of the multi-turn dialogue.

[0155] In an alternative embodiment, the apparatus 600 may further include:

[0156] A preprocessing unit, configured to preprocess the multi-turn dialogue data to obtain preprocessed multi-turn dialogue data, where the preprocessing method includes truncation processing and / or prompt splicing processing of the multi-turn dialogue data;

[0157] The first processing unit 601 is specifically configured to perform a first processing on the preprocessed multi-turn dialogue data.

[0158] In an alternative embodiment, the preprocessing unit is specifically configured to:

[0159] According to a preset sliding window, truncate the multi-turn dialogue data to retain the truncated multi-turn dialogue data; wherein, the sliding window indicates the truncation time and / or the truncation quantity of the multi-turn dialogue.

[0160] And / or,

[0161] According to a preset prompt message, splice the multi-turn dialogue data to obtain multi-turn dialogue data including the prompt message; wherein, the prompt message is used to indicate that the importance of at least one dialogue in the multi-turn dialogue is higher than that of other dialogues.

[0162] Figure 7 It is a schematic flowchart of a model training device provided by an embodiment of the present disclosure. As Figure 7 shown, the device 700 may include an input unit 701, a training unit 702, and an acquisition unit 703.

[0163] Wherein,

[0164] The input unit 701 is configured to input historical multi-turn dialogue data into an initial model for processing multi-turn dialogue to extract a second context feature vector of the historical multi-turn dialogue and a second output feature vector of the historical multi-turn dialogue regarding a downstream task.

[0165] The training unit 702 is configured to train the initial model according to the second context feature vector and the second output feature vector.

[0166] The acquisition unit 703 is configured to obtain a model for processing multi-turn dialogue when a comparison result between the second context feature vector and the second output feature vector reaches a preset comparison condition.

[0167] In an optional implementation manner, the second output feature vector includes at least one of the following: a sentence feature vector after processing the historical multi-turn dialogue data; a question feature vector corresponding to the second context feature vector; an intention recognition vector corresponding to the second context feature vector.

[0168] In an optional implementation manner, the initial model includes a contrast learning branch for processing comparison information between the context feature vector and the output feature vector, and the contrast learning branch includes a preset contrast loss function.

[0169] The training unit 702 is specifically configured to train the initial model according to the second context feature vector and the second output feature vector through the contrast loss function, so that the contrast loss function reaches a preset loss condition, and the loss condition is used to indicate that a comparison result between the second context feature vector and the second output feature vector reaches a preset comparison condition.

[0170] In an alternative embodiment, the input unit 701 is specifically configured to preprocess historical multi-turn conversation data and then input it into the initial model. The preprocessing method includes truncating the historical multi-turn conversation data and / or performing prompt splicing processing.

[0171] In an alternative embodiment, the initial model includes an embedded model.

[0172] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0173] It should be noted that the device, electronic device, readable storage medium, and computer program product of the present disclosure have the beneficial effects corresponding to the above method embodiments, and reference may be made to the relevant descriptions of the above method embodiments, which will not be elaborated herein.

[0174] Figure 8 An electronic device provided by an embodiment of the present disclosure is as Figure 8 shown. The electronic device may include: at least one processor 801; and a memory 802 communicatively connected to the at least one processor; wherein,

[0175] The memory 802 stores instructions executable by the at least one processor 801. The instructions are executed by the at least one processor 801 to enable the at least one processor to execute the above multi-turn conversation processing method or the above model training method. In addition, a transceiver 803 may be included for receiving and transmitting data, such as receiving multi-turn conversation data.

[0176] The corresponding embodiment of the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the multi-turn conversation processing method provided by the above embodiment or the model training method provided by the above embodiment.

[0177] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product. The computer program product includes: a computer program. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the solution provided by any of the above embodiments.

[0178] Figure 9FIG. shows a schematic block diagram of an exemplary electronic device 900 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0179] As Figure 9 shown, the device 900 includes a computing unit 901 that may perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 may also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0180] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 806, such as, for example, a keyboard, a mouse, etc.; an output unit 907, such as, for example, various types of displays, speakers, etc.; a storage unit 908, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 909, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0181] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as method XXX. For example, in some embodiments, method XXX can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the above methods in any other suitable manner (e.g., by means of firmware).

[0182] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0183] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0184] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0186] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0187] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system or a server combined with a blockchain.

[0188] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.

[0189] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for processing a multi-round dialogue, comprising: In response to the multi-round dialogue data, performing a first processing on the multi-round dialogue data to obtain a first context feature vector of the multi-round dialogue, so that a comparison result between the first context feature vector and a first output feature vector meets a preset comparison condition, wherein the first output feature vector is a feature vector about a downstream task extracted based on the multi-round dialogue data; Determine recognition results of the multiple rounds of conversations according to the first context feature vector.

2. The method according to claim 1, wherein: The first output feature vector includes at least one of the following: a sentence feature vector after processing multiple rounds of dialogue data; a question feature vector corresponding to the first context feature vector; and an intention recognition vector corresponding to the first context feature vector.

3. The method according to claim 1 or 2, wherein: The first processing of the multi-round dialogue data includes: Inputting the multi-round dialogue data into a preset model for processing multi-round dialogues, wherein the model is trained based on the historical multi-round dialogue data, and the comparison result between the second context feature vector and the second output feature vector meets a preset comparison condition; wherein the second context feature vector is a feature vector of the historical multi-round dialogue extracted based on the historical multi-round dialogue data, and the second output feature vector is a feature vector of the downstream task extracted based on the historical multi-round dialogue data; The multi-round dialogue data is first processed according to the model to obtain a first context feature vector of the multi-round dialogue.

4. The method according to claim 3, wherein: The model includes an encoding module and an output module, wherein the encoding module is used to encode the preset round dialogues separately, and the output module is used to perform weighted average calculation on the encoding results of the encoding module and then output; Performing a first process on the multi-round dialogue data according to the model to obtain a first context feature vector of the multi-round dialogue includes: According to the encoding module, a preset round of dialogue in the multi-round dialogue data is separately encoded to extract a feature vector corresponding to the round of dialogue in the multi-round dialogue; The output module performs weighted average calculation on the feature vectors of the preset rounds of dialogue to obtain the first context feature vectors of the multi-round dialogue.

5. The method according to claim 4, wherein: The weighted value of the preset round of dialogue is determined according to the prompt information carried in the multi-round dialogue data, or according to the dialogue sequence of the multi-round dialogue data.

6. The method according to any one of claims 1 to 5, wherein: The recognition results of the multiple rounds of dialogue include recognition results of answers to the multiple rounds of dialogue, and the method further includes: Extracting a question feature vector corresponding to each question in the question-answer database, and constructing a preset search structure based on the question feature vector, wherein the search structure is used to query the question feature vector closest to the first context feature vector; The determining the recognition results of the multiple rounds of dialogue according to the first context feature vector includes: Performing a nearest neighbor search on the first context feature vector in the search structure to obtain a preset number of candidate question feature vectors; Determining a target question feature vector according to the cosine similarity between the first context feature vector and the candidate question feature vector; According to the target question feature vector, a target question is determined from the question and answer library, and answer recognition results corresponding to the multiple rounds of dialogues are obtained according to the target question.

7. The method according to any one of claims 1 to 5, wherein: The recognition results of the multiple rounds of conversations include intention recognition results of the multiple rounds of conversations; The determining the recognition results of the multiple rounds of dialogue according to the first context feature vector includes: Inputting the first context feature vector into a preset intention recognition model for recognizing conversation intention, wherein the intention recognition model includes a support vector machine; The first context feature vector is processed based on the intention recognition model to obtain intention recognition results of the multiple rounds of dialogue.

8. The method according to any one of claims 1 to 7, before performing the first processing on the multi-round dialogue data, further comprising: Preprocessing the multi-round conversation data to obtain preprocessed multi-round conversation data, wherein the preprocessing method includes truncation processing and / or prompt splicing processing of the multi-round conversation; The performing a first processing on the multi-round dialogue data includes: performing a first processing on the pre-processed multi-round dialogue data.

9. The method according to claim 8, wherein: The preprocessing of the multi-round dialogue data includes: According to a preset sliding window, the multi-round dialogue data is truncated to retain the truncated multi-round dialogue data; wherein the sliding window indicates the truncation time and / or truncation quantity of the multi-round dialogue; and / or, According to the preset prompt information, the multiple rounds of dialogue data are concatenated to obtain the multiple rounds of dialogue data containing the prompt information; wherein the prompt information is used to prompt that the importance of at least one dialogue in the multiple rounds of dialogue is higher than the importance of other dialogues.

10. A model training method, comprising: Inputting the historical multi-round dialogue data into an initial model for processing multi-round dialogues to extract a second context feature vector of the historical multi-round dialogues and a second output feature vector of the historical multi-round dialogues related to a downstream task; Training the initial model according to the second context feature vector and the second output feature vector; When the comparison result between the second context feature vector and the second output feature vector meets a preset comparison condition, a model for processing multiple rounds of dialogue is obtained.

11. According to the method of claim 10, the second output feature vector includes at least one of the following: a sentence feature vector after processing the historical multi-round dialogue data; a question feature vector corresponding to the second context feature vector; and an intention recognition vector corresponding to the second context feature vector.

12. The method according to claim 10 or 11, wherein: The initial model includes a contrastive learning branch for processing contrast information between the context feature vector and the output feature vector, and the contrastive learning branch includes a preset contrastive loss function; The step of training the initial model according to the second context feature vector and the second output feature vector comprises: According to the second context feature vector and the second output feature vector, the initial model is trained by the contrast loss function so that the contrast loss function reaches a preset loss condition, and the loss condition is used to indicate that the contrast result between the second context feature vector and the second output feature vector reaches a preset contrast condition.

13. The method according to any one of claims 10 to 12, wherein: Input historical multi-round dialogue data into the initial model for processing multi-round dialogues, including: The historical multi-round dialogue data is pre-processed and then input into the initial model, wherein the pre-processing method includes truncation processing and / or prompt splicing processing of the historical multi-round dialogue data.

14. The method according to any one of claims 10 to 13, wherein: The initial model includes an embedded model.

15. A multi-round conversation analysis device, comprising: a first processing unit, configured to, in response to the multi-round dialogue data, perform a first processing on the multi-round dialogue data to obtain a first context feature vector of the multi-round dialogue, so that a comparison result between the first context feature vector and a first output feature vector meets a preset comparison condition, wherein the first output feature vector is a feature vector about a downstream task extracted based on the multi-round dialogue data; A recognition unit is used to determine the recognition results of the multiple rounds of dialogue according to the first context feature vector.

16. A model training device, comprising: An input unit, used to input the historical multi-round dialogue data into an initial model for processing multi-round dialogues to extract a second context feature vector of the historical multi-round dialogues and a second output feature vector of the historical multi-round dialogues related to a downstream task; A training unit, configured to train the initial model according to the second context feature vector and the second output feature vector; An acquisition unit is used to obtain a model for processing multiple rounds of dialogue when a comparison result between the second context feature vector and the second output feature vector meets a preset comparison condition.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.

19. A computer program product, comprising a computer program, which implements the steps of the method according to any one of claims 1 to 14 when executed by a processor.