Multi-round dialogue classification method and device, equipment and storage medium

By obtaining the feature vectors of multiple rounds of dialogue text and inputting them into the text classification model, candidate classification categories and question information are generated, and inputting them into the large language model to determine the target classification category, the problem of low accuracy of multi-round dialogue classification is solved and the accuracy of classification is improved.

CN120216694APending Publication Date: 2025-06-27PING AN INT FINANCIAL LEASING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510274114.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately classify multiple rounds of conversations, resulting in low accuracy of multiple rounds of text classification, especially in intelligent consultation and customer service scenarios in the fields of medical care and health and finance.

Method used

By obtaining the first text feature vector of multiple rounds of dialogue text, input it into the preset text classification model for processing, and obtaining multiple candidate classification categories. Questioning information is generated based on these candidate categories and input into a large language model to determine the target classification category for multiple rounds of conversation text.

Benefits of technology

By streamlining classification categories and using large language models to determine target classification categories, the accuracy of multiple rounds of dialogue classification is improved, and is suitable for complex intelligent consultation and customer service scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216694A_ABST
    Figure CN120216694A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to the field of medical health and the field of financial science and technology, and provides a multi-round dialogue classification method and device, equipment and a storage medium, and the method comprises the steps: obtaining a first text feature vector of a to-be-recognized multi-round dialogue text; inputting the first text feature vector into a preset text classification model for processing to obtain a plurality of candidate classification categories of the multi-round dialogue text; on the basis of the multiple rounds of dialogue texts and the multiple candidate classification categories of the multiple rounds of dialogue texts, question information used for being input into a large language model is generated; and inputting the question information into the large language model, so that the large language model determines a target classification category of the multi-round dialogue text based on the question information. The invention aims to improve the accuracy of multi-round text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a multi-turn dialogue classification method, device, equipment, and storage medium. Background Art

[0002] Classifying various types of dialogues is a common task in the field of natural language processing. Among them, multi-turn dialogues usually contain relatively long text, so the classification difficulty is relatively large compared to conventional single-turn dialogues.

[0003] To solve the above problems, the long text of the multi-turn dialogue can be divided into multiple short texts and then classified one by one to obtain the mapping relationship between the multi-turn dialogue and multiple classification labels. However, the context complexity of the multi-turn dialogue is relatively high, and each round of user input may depend on the previous context. It is difficult to accurately classify the multi-turn dialogue based on the mapping relationship between the multi-turn dialogue and multiple classification labels, and the accuracy of multi-turn text classification is relatively low. For example, in the online intelligent consultation scenario in the field of medical and health, patients may ask some questions related to their condition, symptoms, historical treatment records, or family medical history, and follow up according to the system's answers. In this context, each reply from the patient will depend on the previous consultation information, and it is often difficult to accurately judge the actual needs of the patient based on a single symptom. In the online intelligent customer service scenario in the financial field, customers may ask questions related to account balances, transaction records, investment financial products, or loan situations, and conduct further consultations according to the system's answers. It is also difficult to accurately understand the true needs of customers based on a single question or data point.

[0004] Therefore, how to improve the accuracy of multi-turn text classification is an urgent problem to be solved at present. Summary of the Invention

[0005] The main purpose of this application is to provide a multi-turn dialogue classification method, device, equipment, and storage medium, aiming to improve the accuracy of multi-turn text classification.

[0006] In a first aspect, this application provides a multi-turn dialogue classification method, including:

[0007] Obtain a first text feature vector of the multi-turn dialogue text to be recognized;

[0008] Input the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text;

[0009] Generate question information for input to a large language model based on the multi-turn dialogue text and the multiple candidate classification categories of the multi-turn dialogue text;

[0010] Input the question information into the large language model for the large language model to determine the target classification category of the multi-turn dialogue text based on the question information.

[0011] In a second aspect, the present application further provides a multi-turn dialogue classification device, which includes:

[0012] An acquisition module, configured to acquire a first text feature vector of a multi-turn dialogue text to be recognized;

[0013] A classification module, configured to input the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text;

[0014] A generation module, configured to generate question information for input into a large language model based on the multi-turn dialogue text and the multiple candidate classification categories of the multi-turn dialogue text;

[0015] An input module, configured to input the question information into the large language model for the large language model to determine the target classification category of the multi-turn dialogue text based on the question information.

[0016] In a third aspect, the present application further provides a computer device, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the multi-turn dialogue classification method described above are implemented.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-turn dialogue classification method described above are implemented.

[0018] The present application provides a multi-turn dialogue classification method, device, equipment, and storage medium. The present application first acquires a first text feature vector of a multi-turn dialogue text to be recognized; inputs the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text; generates question information for input into a large language model based on the multi-turn dialogue text and the multiple candidate classification categories of the multi-turn dialogue text; and inputs the question information into the large language model for the large language model to determine the target classification category of the multi-turn dialogue text based on the question information. The present application first screens out multiple candidate classification categories of the multi-turn dialogue text to be recognized to reduce the subsequent classification difficulty by streamlining the classification categories, and uses the question information generated based on the multiple candidate classification categories to determine the target classification category of the multi-turn dialogue text from the multiple candidate classification categories through the large language model, thereby improving the accuracy of multi-turn dialogue classification. Description of the Drawings

[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of the steps of a multi-turn dialogue classification method provided by an embodiment of the present application;

[0021] Figure 2 For Figure 1 It is a schematic sub-step flowchart of the multi-turn dialogue classification method in

[0022] Figure 3 For Figure 1 It is another schematic sub-step flowchart of the multi-turn dialogue classification method in

[0023] Figure 4 It is a schematic block diagram of a multi-turn dialogue classification device provided by an embodiment of the present application;

[0024] Figure 5 For Figure 4 It is a schematic block diagram of a sub-module of the multi-turn dialogue classification device in

[0025] Figure 6 For Figure 4 It is another schematic block diagram of a sub-module of the multi-turn dialogue classification device in

[0026] Figure 7 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.

[0027] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments and with reference to the drawings. Detailed implementation manners

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0029] The flowchart shown in the drawings is only an example, and does not necessarily include all the contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed or partially combined, so the actual execution order may be changed according to the actual situation.

[0030] In the natural language processing task of dialogue classification, multi-turn dialogues usually contain long and relatively complex text. Therefore, the classification difficulty is relatively greater than that of conventional single-turn dialogues.

[0031] In the prior art, dialogue classification tasks usually use small pre-trained language models, large language models, or a combination of both.

[0032] When using a small pre-trained language model for dialogue classification tasks, methods such as the truncation method and the pooling method are usually adopted. Among them, the truncation method truncates the text length of the input multi-turn dialogue and retains the important part to adapt to the input length of the pre-trained model. However, the truncation method will cause loss of some dialogue information, thus affecting the accuracy of text classification; the pooling method divides the long text of multi-turn dialogues into multiple short texts, places multiple short texts in the same batch and inputs them into the small pre-trained language model to obtain the vector representation corresponding to each short text, and performs a pooling operation on the short text representation vectors belonging to the same multi-turn dialogue, and then obtains the long text representation vector corresponding to the multi-turn dialogue. However, the pooling method only establishes a mapping relationship between multi-turn dialogues and classification labels, without establishing a semantic connection between multi-turn dialogues and classification labels, and cannot obtain the long-term dependence relationship between each dialogue turn in multi-turn dialogues.

[0033] When using a large language model for dialogue classification tasks, the long text length of multi-turn dialogues and the excessive number of dialogue categories will generate overly long query information, causing the large language model to focus too much on classification categories that are not relevant to the dialogue text, increasing the classification difficulty of multi-turn dialogues, resulting in poor prediction accuracy of the large language model, and even outputting unseen categories, resulting in low accuracy of multi-turn text classification.

[0034] When using a combination of a small pre-trained language model and a large language model for dialogue classification tasks, multiple text combinations are formed by the user's dialogue history and the last query. The small pre-trained language model is used to generate semantic feature vector representations for these text combinations to represent the text, and an intent classification is predicted for each text combination. The classification labels of multiple text combinations form candidate classifications, single-turn dialogue examples are constructed according to the similarity with the multi-turn dialogue, and an input prompt is constructed: <task instruction, example, dialogue context, and user query>, and predictions are generated based on the large language model. However, this method only performs intent classification on the last user input.

[0035] Based on this, the embodiments of the present application provide a multi-round dialogue classification method, apparatus, device, and storage medium to solve at least one of the technical problems in the above-mentioned prior art. Among them, the multi-round dialogue classification method can be applied to a terminal device or a server. The terminal device can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device; the server can be a single server or a server cluster composed of multiple servers. Hereinafter, an example will be given to explain the multi-round dialogue classification method applied to a server.

[0036] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the steps of a multi-round dialogue classification method provided by an embodiment of the present application.

[0038] As Figure 1 shown, the multi-round dialogue classification method includes steps S101 to S104.

[0039] Step S101: Obtain a first text feature vector of the multi-round dialogue text to be recognized.

[0040] Among them, the multi-round dialogue text to be recognized is a dialogue text composed of multiple sentences, such as a dialogue text composed of multiple interactive questions and answers between a user and a system. The first text feature vector is used to represent the multi-round dialogue text, and can be a comprehensive representation of all the text content in the multi-round dialogue text, or a partial representation of some key dialogue rounds or some specific information extracted from the multi-round dialogue text. For example, in the application in the field of medical and health, the multi-round dialogue text to be recognized may come from the question-and-answer interaction between a patient and a system on issues related to the condition, symptoms, historical treatment records, or family medical history on an online intelligent consultation platform. In the application in the financial field, the multi-round dialogue text to be recognized may come from the question-and-answer interaction between a customer and a system on aspects such as account information, transaction records, and loan situations on an online intelligent customer service platform.

[0041] In one embodiment, step S101 includes sub-steps S1011 to S1012.

[0042] Sub-step S1011: Convert the multi-round dialogue text to be recognized into a second text feature vector.

[0043] Among them, the second text feature vector can be a comprehensive representation of all the text content in the multi-turn dialogue text, or a partial representation of some partial text content obtained by extracting certain key dialogue turns or certain specific information from the multi-turn dialogue text.

[0044] In one embodiment, converting the multi-turn dialogue text to be recognized into a second text feature vector includes: obtaining the dialogue turns of the multi-turn dialogue text to be recognized; based on the dialogue turns, performing a segmentation process on the multi-turn dialogue text to obtain a plurality of single-turn dialogue texts; converting the plurality of single-turn dialogue texts into a plurality of target text feature vectors; based on the dialogue turns, performing a splicing process on the plurality of target text feature vectors to obtain the second text feature vector corresponding to the multi-turn dialogue text.

[0045] Among them, the dialogue turn refers to an interaction between participants (such as a user and a system) during a dialogue process. For example, a multi-turn dialogue text includes three single-turn dialogue texts, and the specific content is as follows:

[0046] User: "Hello."

[0047] System: "Hello! How can I help you?"

[0048] User: "I want to know the weather."

[0049] System: "Please tell me the city you are in."

[0050] User: "Beijing."

[0051] System: "It is sunny turning to cloudy in Beijing today, and the temperature is 20-28 degrees."

[0052] That is, each turn of the dialogue is the user's speech and the system's response, and they take turns. The content of each turn of the dialogue usually depends on the previous dialogue turns (i.e., the context). For example, the system's answer may be based on the user's previous question. It can be understood that when segmenting a multi-turn dialogue, single-turn dialogue texts containing irrelevant greetings may be obtained, and these texts can be filtered according to actual needs to improve the efficiency of subsequent processing.

[0053] Exemplarily, the multi-turn dialogue text Q to be recognized includes n turns of dialogue, expressed as: Q = {(q1, a1), (q2, a2),..., (q n , a n )}. The multi-turn dialogue text Q is segmented into n single-turn dialogue texts: (q i , a i ), i ∈ 1, 2,..., n.

[0054] In one embodiment, after splitting the multi-turn dialogue text based on the dialogue turn to obtain multiple single-turn dialogue texts, the method further includes: detecting the number of characters in the multiple single-turn dialogue texts; determining target single-turn dialogue texts with the number of characters greater than or equal to a set number of characters from the multiple single-turn dialogue texts; and splitting the target single-turn dialogue texts so that the number of characters in the target single-turn dialogue texts is less than the set number of characters.

[0055] Among them, the set number of characters can be determined according to the input limit of the conversion model into which the single-turn dialogue text is input for converting into the target text feature vector. For example, some conversion models (such as BERT, GPT, etc.) usually have a maximum input length, and texts exceeding this length cannot be directly input into the model. Therefore, the set number of characters can be determined according to the maximum input length of the model, such as 1 / 2 of the maximum input length, etc. Correspondingly, the splitting of the target single-turn dialogue text can also be adjusted according to the set number of characters or the maximum input length, as long as it can ensure that the text can be completely input into the conversion model. Those skilled in the art can adjust it by themselves and no specific limitation is made here.

[0056] It should be noted that by further splitting the text according to the set number of characters, it is possible to avoid situations such as being unable to be input due to exceeding the input limit when converting multiple single-turn dialogue texts, or losing important information after input, so as to ensure the smooth progress of subsequent dialogue classification or other tasks based on the output multiple target text feature vectors.

[0057] Sub-step S1012: Input the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-turn dialogue text.

[0058] Specifically, the rotational position encoding is used to implement by rotating the positions of the respective feature elements in the second text feature vector. The encoding of each position can be adjusted based on the relative position, and the rotation parameter depends on the relative distance between each position. It should be noted that through the rotational position encoding, the multi-head attention network can better learn the relative relationship between different positions in the text feature, so as to better capture the long-distance dependency relationship, and further improve the expression ability of the text feature. That is to say, after the rotational position encoding, the first text feature vector can express the multi-turn dialogue text more accurately and clearly.

[0059] In one embodiment, the multi-head attention network includes a position transformation layer, a position encoding layer, a vector calculation layer, a vector splicing layer, and a position inverse transformation layer.

[0060] Among them, the position transformation layer is used to map the input feature vector to another space to adjust the relationship between the position information and the input features through transformation, so as to facilitate capturing dependencies in a higher dimension.

[0061] The position encoding layer is used to encode the positions of the feature elements through rotation operations to further capture the relative relationships between the feature elements in the feature vector by adjusting the relative position relationships.

[0062] The vector calculation layer is used to calculate the weights of the input feature vector to determine the relative importance of each element according to the relative relationships between the feature elements in the input feature vector.

[0063] The vector concatenation layer is used to concatenate the outputs of the vector calculation layer to obtain a more comprehensive feature representation.

[0064] The position inverse transformation layer is used to perform an inverse position transformation on the feature vector to restore the feature vector output by the vector concatenation layer to the same position structure as when input to the position transformation layer.

[0065] In one embodiment, inputting the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-turn dialogue text, including:

[0066] Performing multiple position transformations on the second text feature vector based on the first mapping matrix of the position transformation layer to obtain a first transformed vector injected with different first text position information, and performing multiple position transformations on the second text feature vector based on the second mapping matrix of the position transformation layer to obtain a second transformed vector injected with different second text position information;

[0067] Performing rotational position encoding on multiple first transformed vectors based on the rotation matrix of the position encoding layer to obtain multiple third transformed vectors fused with rotational text position information;

[0068] Performing weight calculation on multiple third transformed vectors based on the attention mechanism of the vector calculation layer to obtain multiple attention weights, and obtaining multiple target transformed vectors based on the product of each attention weight and each second transformed vector;

[0069] Performing a concatenation process on multiple target transformed vectors based on the vector concatenation layer to obtain a target text feature vector;

[0070] Performing a position inverse transformation on the target text feature vector based on the inverse transformation matrix of the position inverse transformation layer to obtain the first text feature vector.

[0071] Exemplarily, the processing process of inputting the second text feature vector E corresponding to the multi-turn dialogue text into the multi-head attention network is as follows:

[0072] The first mapping matrix W based on the position transformation layer Q and W K and the second mapping matrix W V , perform a position transformation on E:

[0073] Q = EW Q , K = EW K , V = EW V

[0074] Based on the position encoding layer, perform rotational position encoding on Q and K:

[0075] Q' = RoPE(Q) = Q · Rot(θ pos ), K' = RoPE(K) = K · Rot(θ pos )

[0076] Specifically, use the rotation matrix to multiply the adjacent two-dimensional values on each token of Q and K

[0077] where p represents the position of the token, i represents the dimensional position, and the token is a basic unit in the text, usually a word, punctuation mark, or sub-word, etc.

[0078] Based on the attention mechanism of the vector calculation layer, calculate the attention weights for multiple Q' and multiple K' respectively:

[0079]

[0080] where D k is the scaling factor used to stabilize the gradient.

[0081] Based on the product of multiple attention weights and multiple second transformation vectors, obtain multiple target transformation vectors respectively:

[0082]

[0083] where h ∈ 1, 2,..., H, and H is the number of weights AttentionScores.

[0084] Based on the vector concatenation layer, perform concatenation processing on multiple heads h :

[0085] Head = concat(head1, head2,..., head H )

[0086] Based on the inverse transformation matrix W of the position inverse transformation layer KPerform an inverse position transformation on Head to obtain the first text feature vector M that is consistent with the input in terms of the E dimension. head :

[0087] M head = Head · W O

[0088] In one embodiment, after performing an inverse position transformation on the target text feature vector based on the inverse transformation matrix of the position inverse transformation layer to obtain the first text feature vector, it further includes: performing a residual connection and layer normalization processing on the first text feature vector to update the first text feature vector.

[0089] Exemplarily, perform a residual connection on the first text feature vector E to obtain the updated first text feature vector O. M :

[0090] O M = M head + E

[0091] Perform layer normalization on the updated first text feature vector O. M to obtain the first text feature vector O' that is updated again. M :

[0092] O' M = layerNorm(O M ) = layerNorm(M head + E)

[0093] It should be noted that the residual connection combines the original second text feature vector with the first text feature vector output by the multi-head attention network through an addition operation, making it easier to learn differential features subsequently. Layer normalization standardizes the output of each layer, enabling the updated first text feature vector to be processed more effectively. It can be seen that updating the first text feature vector through residual connection and layer normalization processing can capture text features more deeply, thereby effectively enhancing the expression ability of text features.

[0094] In one embodiment, after inputting the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-round dialogue text, it further includes: inputting the first text feature vector into a preset feed-forward neural network for feature extraction to update the first text feature vector.

[0095] Exemplarily, input the first text feature vector O'. M Input it into the feed-forward neural network FFN to obtain the updated first text feature vector O. F :

[0096] O F = FFN(O' M ) = ReLU(O' M W1 + B1)W2 + B

[0097] Wherein, W1 and W2 are weight matrices, B1 and B2 are bias values, and ReLU is an activation function.

[0098] It should be noted that through the non-linear transformation and activation function in the feed-forward neural network, the second text feature vector that may still contain some redundant or irrelevant information can be abstracted, thereby enhancing the expression ability of the features, so that the updated first text feature vector retains the feature data more closely related to the subsequent classification task, thereby improving the accuracy of classification.

[0099] In one embodiment, after inputting the first text feature vector into a preset feed-forward neural network for feature extraction to update the first text feature vector, it further includes: performing residual connection and layer normalization processing on the updated first text feature vector to update the first text feature vector again.

[0100] Wherein, the specific operations and functions of updating the first text feature vector through residual connection and layer normalization processing can refer to the process of updating the first text feature vector by the above-mentioned multi-head attention network, which will not be elaborated here.

[0101] In one embodiment, after obtaining the first text feature vector of the multi-turn dialogue text to be recognized, it further includes: performing an average operation on the first text feature vector to update the first text feature vector.

[0102] Exemplarily, the representation of the entire multi-turn dialogue is obtained by calculating the average of the representation vectors of each single-turn dialogue text in the multi-turn dialogue text, as follows:

[0103]

[0104] Wherein, n is the number of dialogue turns, and O j represents the representation vector of the j-th single-turn dialogue text.

[0105] It should be noted that the average operation can aggregate the representations of each single-turn dialogue text in the multi-turn dialogue text to obtain a more global dialogue representation, so as to more comprehensively summarize the semantic information of the entire multi-turn dialogue text.

[0106] It should be noted that, to further ensure the privacy and security of the above-mentioned multi-turn dialogue text and other relevant information to be recognized, the above-mentioned multi-turn dialogue text and other relevant information to be recognized can also be stored in a node of a blockchain. The technical solution of this application can also be applied to adding other data files stored on the blockchain. The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.

[0107] Step S102: Input the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text.

[0108] Among them, the preset text classification model can be a deep learning model (such as CNN, RNN, Transformer, etc.) or a traditional machine learning model (such as SVM, Naive Bayes, etc.). The specific number of multiple candidate classification categories is associated with the representation degree of the first text feature vector for the multi-turn dialogue text to be classified. In some cases, for example, when only the feature data closely related to the classification task is retained in the first text feature vector, the multiple candidate classification categories can be all the classification categories corresponding to the first text feature vector. In other cases, for example, when there is still redundant or irrelevant information retained in the first text feature vector, the multiple candidate classification categories can also be partial classification categories corresponding to the first text feature vector.

[0109] In one embodiment, the text classification model includes a linear layer, an activation function layer, and an output layer. Step S102 includes sub-steps S1021 to S1023.

[0110] Sub-step S1021: Perform classification prediction on the first text feature vector based on the linear layer to obtain prediction results of multiple classification labels.

[0111] Exemplarily, assume that the linear layer has 5 output neurons corresponding to 5 classification labels, and the first text feature vector O final is a vector containing 100-dimensional features. Input it into this linear layer, and after linear transformation, 5 values (logits) are obtained, respectively representing the prediction results for each classification label. The specific calculation process is as follows:

[0112] logits = O final W f + B f

[0113] Among them, W f is the weight matrix (5 * 100 dimensions), B f is the bias value (5 dimensions), and logits are the 5 values obtained for subsequent classification prediction.

[0114] Exemplarily, assume the obtained logits are: logits = [2.5, -1.0, 3.0, 0.5, 1.2]. Among them, the respective values in logits represent the original predicted values of each classification label.

[0115] Sub-step S1022: Based on the activation function layer, perform probability prediction on the prediction results of multiple classification labels to obtain distribution probability data of multiple classification labels.

[0116] Exemplarily, based on the previously obtained logits = [2.5, -1.0, 3.0, 0.5, 1.2], convert them into a probability distribution through the activation function Softmax. The calculation results are as follows:

[0117] [P(y1) = 0.35, P(y2) = 0.05, P(y3) = 0.42, P(y4) = 0.08, P(y5) = 0.10]

[0118] Among them, P(y i ) represents the probability that the classification category is y i . For example, P(y1) = 0.35 means the probability of the category being y1 is 35%.

[0119] Sub-step S1023: Based on the output layer, screen the distribution probability data of multiple classification labels to select multiple candidate classification categories of the multi-turn dialogue text from multiple classification labels.

[0120] Among them, the candidate classification categories can be obtained by setting a threshold to screen multiple classification labels, or can be determined by selecting the k categories with the highest probabilities from multiple classification labels.

[0121] Exemplarily, if the top 3 categories by probability are selected, the classification categories corresponding to P(y1), P(y3), and P(y5) will be selected as candidate classification categories.

[0122] Step S103: Based on the multi-turn dialogue text and multiple candidate classification categories of the multi-turn dialogue text, generate question information for input to the large language model.

[0123] Exemplarily, assume in the field of medical and health, in the health consultation scenario of an online intelligent consultation platform, the system and the user have had multiple rounds of conversations. The results of the multi-turn dialogue text and multiple candidate classification categories of the multi-turn dialogue text are as follows:

[0124] Multi-turn dialogue text:

[0125] User: Recently, I always feel extremely tired, sleep poorly at night, and always have headaches when I wake up in the morning.

[0126] System: Have you adjusted your daily routine recently? Have there been any changes in your diet?

[0127] User: I'm very busy at work now, often stay up late, and eat irregularly.

[0128] System: Understood. Staying up late and irregular diet may be the causes of fatigue and headaches. Have you tried any ways to improve your sleep or adjust your diet?

[0129] User: I've tried, but the effect is not obvious. I wonder if I need to do more exercise to relieve this situation.

[0130] Candidate classification categories: Healthy diet, Exercise and fitness, Stress management, and Sleep improvement.

[0131] The question information is: Based on the above conversation information, the user seems to be experiencing fatigue and sleep problems caused by staying up late and irregular diet, and hopes to know whether exercise can relieve this situation. Please determine the most appropriate classification category according to this information.

[0132] It can be understood that by combining multi-round conversation texts and multiple candidate classification labels to generate a structured and clear question information, it can help the large language model better understand the needs of the questioner and make accurate answers.

[0133] Step S104: Input the question information into the large language model for the large language model to determine the target classification category of the multi-round conversation text based on the question information.

[0134] Exemplarily, in the field of medical and health, according to the user's description in the above question, although diet and sleep problems may have an impact on their health, the user clearly states that they want to know whether they need to exercise to relieve fatigue. Therefore, the model determines that "Exercise and fitness" in the candidate classification categories is the target classification category of this multi-round conversation text, which helps to more accurately identify the actual needs of the user. Similarly, in the application of the fintech field, the model can identify the customer's interest in short-term financial management, long-term investment, or risk management products, so as to accurately recommend corresponding products, improve the customer experience and satisfaction, and at the same time improve the targeting and conversion rate of marketing.

[0135] The multi-turn dialogue classification method provided by the above embodiment first obtains the first text feature vector of the multi-turn dialogue text to be recognized; inputs the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text; generates question information for input into a large language model based on the multi-turn dialogue text and the multiple candidate classification categories of the multi-turn dialogue text; and inputs the question information into the large language model for the large language model to determine the target classification category of the multi-turn dialogue text based on the question information. In this application, multiple candidate classification categories of the multi-turn dialogue text to be recognized are first screened out to reduce the subsequent classification difficulty by streamlining the classification categories, and the question information generated based on the multiple candidate classification categories is used to determine the target classification category of the multi-turn dialogue text from the multiple candidate classification categories through the large language model, thereby improving the accuracy of multi-turn dialogue classification.

[0136] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a multi-turn dialogue classification device provided by an embodiment of the present application.

[0137] As Figure 4 shown, the multi-turn dialogue classification device 200 includes:

[0138] An acquisition module 201, configured to obtain the first text feature vector of the multi-turn dialogue text to be recognized.

[0139] A classification module 202, configured to input the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text.

[0140] A generation module 203, configured to generate question information for input into a large language model based on the multi-turn dialogue text and the multiple candidate classification categories of the multi-turn dialogue text.

[0141] An input module 204, configured to input the question information into the large language model for the large language model to determine the target classification category of the multi-turn dialogue text based on the question information.

[0142] In one embodiment, as Figure 5 shown, the acquisition module 201 includes:

[0143] A vector conversion sub-module 2011, configured to convert the multi-turn dialogue text to be recognized into a second text feature vector.

[0144] A position rotation sub-module 2012, configured to input the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-turn dialogue text.

[0145] In one embodiment, the vector conversion sub-module 2011 is further configured to: obtain the dialogue turns of the multi-turn dialogue text to be recognized; based on the dialogue turns, perform segmentation processing on the multi-turn dialogue text to obtain multiple single-turn dialogue texts; convert the multiple single-turn dialogue texts into multiple target text feature vectors; based on the dialogue turns, perform splicing processing on the multiple target text feature vectors to obtain a second text feature vector corresponding to the multi-turn dialogue text.

[0146] In one embodiment, the position rotation sub-module 2012 is further configured to: input the first text feature vector into a preset feed-forward neural network for feature extraction to update the first text feature vector.

[0147] In one embodiment, the multi-head attention network includes a position transformation layer, a position encoding layer, a vector calculation layer, a vector splicing layer, and a position inverse transformation layer. The position rotation sub-module 2012 is further configured to: perform multiple position transformations on the second text feature vector based on the first mapping matrix of the position transformation layer to obtain a first transformation vector injected with different first text position information, and perform multiple position transformations on the second text feature vector based on the second mapping matrix of the position transformation layer to obtain a second transformation vector injected with different second text position information; perform rotational position encoding on the multiple first transformation vectors based on the rotation matrix of the position encoding layer to obtain multiple third transformation vectors fused with rotational text position information; perform weight calculation on the multiple third transformation vectors based on the attention mechanism of the vector calculation layer to obtain multiple attention weights, and obtain multiple target transformation vectors based on the product of each attention weight and each second transformation vector; perform splicing processing on the multiple target transformation vectors based on the vector splicing layer to obtain a target text feature vector; perform position inverse transformation on the target text feature vector based on the inverse transformation matrix of the position inverse transformation layer to obtain the first text feature vector.

[0148] In one embodiment, the position rotation sub-module 2012 is further configured to: perform residual connection and layer normalization processing on the first text feature vector to update the first text feature vector.

[0149] In one embodiment, the text classification model includes a linear layer, an activation function layer, and an output layer. As Figure 6 shown, the classification module 202 includes:

[0150] A classification prediction sub-module 2021, configured to perform classification prediction on the first text feature vector based on the linear layer to obtain prediction results of multiple classification labels.

[0151] A probability prediction sub-module 2022, configured to perform probability prediction on the prediction results of multiple classification labels based on the activation function layer to obtain distribution probability data of multiple classification labels.

[0152] The category screening sub-module 2023 is used to screen the distribution probability data of multiple classification labels based on the output layer, so as to select multiple candidate classification categories of the multi-turn dialogue text from multiple classification labels.

[0153] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing embodiments of the multi-turn dialogue classification method, and will not be described herein again.

[0154] The device provided by the above embodiment can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 7 shown.

[0155] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.

[0156] As shown in Figure 7 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a storage medium and an internal memory, and the storage medium can be non-volatile or volatile.

[0157] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any multi-turn dialogue classification method.

[0158] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0159] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any multi-turn dialogue classification method.

[0160] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 7 the structure shown in

[0161] It should be understood that the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0162] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0163] Obtain the first text feature vector of the multi-turn dialogue text to be recognized.

[0164] Input the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text.

[0165] Generate question information for input to the large language model based on the multi-turn dialogue text and the multiple candidate classification categories of the multi-turn dialogue text.

[0166] Input the question information into the large language model for the large language model to determine the target classification category of the multi-turn dialogue text based on the question information.

[0167] In one embodiment, when the processor implements obtaining the first text feature vector of the multi-turn dialogue text to be recognized, it is used to implement:

[0168] Convert the multi-turn dialogue text to be recognized into a second text feature vector;

[0169] Input the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-turn dialogue text.

[0170] In one embodiment, when the processor implements converting the multi-turn dialogue text to be recognized into a second text feature vector, it is used to implement:

[0171] Obtain the dialogue turns of the multi-turn dialogue text to be recognized; based on the dialogue turns, perform segmentation processing on the multi-turn dialogue text to obtain multiple single-turn dialogue texts; convert the multiple single-turn dialogue texts into multiple target text feature vectors; based on the dialogue turns, perform splicing processing on the multiple target text feature vectors to obtain the second text feature vector corresponding to the multi-turn dialogue text.

[0172] In one embodiment, after the processor implements inputting the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-turn dialogue text, it is further used to implement:

[0173] Input the first text feature vector into a preset feed-forward neural network for feature extraction to update the first text feature vector.

[0174] In one embodiment, the multi-head attention network includes a position transformation layer, a position encoding layer, a vector calculation layer, a vector concatenation layer, and a position inverse transformation layer. When the processor implements inputting the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-turn dialogue text, it is used to implement:

[0175] Perform multiple position transformations on the second text feature vector based on the first mapping matrix of the position transformation layer to obtain a first transformed vector injected with different first text position information, and perform multiple position transformations on the second text feature vector based on the second mapping matrix of the position transformation layer to obtain a second transformed vector injected with different second text position information;

[0176] Based on the rotation matrix of the position encoding layer, perform rotational position encoding on multiple first transformed vectors to obtain multiple third transformed vectors fused with rotational text position information;

[0177] Based on the attention mechanism of the vector calculation layer, calculate weights for multiple third transformed vectors to obtain multiple attention weights, and obtain multiple target transformed vectors based on the product of each attention weight and each second transformed vector;

[0178] Based on the vector concatenation layer, perform concatenation processing on multiple target transformed vectors to obtain a target text feature vector;

[0179] Based on the inverse transformation matrix of the position inverse transformation layer, perform position inverse transformation on the target text feature vector to obtain the first text feature vector.

[0180] In one embodiment, after the processor implements performing position inverse transformation on the target text feature vector based on the inverse transformation matrix of the position inverse transformation layer to obtain the first text feature vector, it is further used to implement:

[0181] Perform residual connection and layer normalization processing on the first text feature vector to update the first text feature vector.

[0182] In one embodiment, the text classification model includes a linear layer, an activation function layer, and an output layer. When the processor implements inputting the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-turn dialogue text, it is used to implement:

[0183] Classify and predict the first text feature vector based on the linear layer to obtain the prediction results of multiple classification labels;

[0184] Perform probability prediction on the prediction results of multiple classification labels based on the activation function layer to obtain the distribution probability data of multiple classification labels;

[0185] Filter the distribution probability data of multiple classification labels based on the output layer to select multiple candidate classification categories of the multi-turn dialogue text from multiple classification labels.

[0186] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described computer device can refer to the corresponding process in the foregoing embodiments of the multi-turn dialogue classification method, and will not be elaborated herein.

[0187] This application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0188] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed can refer to each embodiment of the multi-turn dialogue classification method of the present application.

[0189] Among them, the computer-readable storage medium may be the internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.

[0190] Furthermore, the computer-usable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc. The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0191] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0192] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any and all possibilities of one or more of the related listed items, and includes these. It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including the element.

[0193] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A multi-round dialogue classification method, comprising: Obtaining a first text feature vector of a multi-round conversation text to be recognized; Inputting the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-round dialogue text; Based on the multiple rounds of dialogue texts and multiple candidate classification categories of the multiple rounds of dialogue texts, generating question information for input into a large language model; The question information is input into the large language model so that the large language model determines a target classification category of the multi-round dialogue text based on the question information.

2. The multi-round dialogue classification method according to claim 1, characterized in that: The step of obtaining a first text feature vector of the multi-round conversation text to be recognized includes: Converting the multi-round conversation texts to be recognized into a second text feature vector; The second text feature vector is input into a preset multi-head attention network for rotational position encoding to obtain a first text feature vector of the multi-round dialogue text.

3. The multi-round dialogue classification method as claimed in claim 2, characterized in that: After the second text feature vector is input into a preset multi-head attention network for rotational position encoding to obtain the first text feature vector of the multi-round dialogue text, the method further includes: The first text feature vector is input into a preset feedforward neural network for feature extraction to update the first text feature vector.

4. The multi-round dialogue classification method as claimed in claim 2, characterized in that: The multi-head attention network includes a position transformation layer, a position encoding layer, a vector calculation layer, a vector splicing layer and a position inverse transformation layer; The step of inputting the second text feature vector into a preset multi-head attention network for rotational position encoding to obtain a first text feature vector of the multi-round dialogue text includes: Performing multiple position transformations on the second text feature vector based on the first mapping matrix of the position transformation layer to obtain a first transformation vector injected with different first text position information, and performing multiple position transformations on the second text feature vector based on the second mapping matrix of the position transformation layer to obtain a second transformation vector injected with different second text position information; Based on the rotation matrix of the position encoding layer, performing rotation position encoding on the plurality of the first transformation vectors to obtain a plurality of third transformation vectors fused with the rotation text position information; Based on the attention mechanism of the vector calculation layer, weight calculation is performed on the plurality of third transformation vectors to obtain a plurality of attention weights, and based on the product of each of the attention weights and each of the second transformation vectors, a plurality of target transformation vectors are obtained; Based on the vector concatenation layer, a plurality of the target transformation vectors are concatenated to obtain a target text feature vector; Based on the inverse transformation matrix of the position inverse transformation layer, the target text feature vector is subjected to position inverse transformation to obtain a first text feature vector.

5. The multi-round dialogue classification method according to claim 4, characterized in that: After performing position inverse transformation on the target text feature vector based on the inverse transformation matrix of the position inverse transformation layer to obtain the first text feature vector, the method further includes: Perform residual connection and layer normalization processing on the first text feature vector to update the first text feature vector.

6. The multi-round dialogue classification method according to claim 2, characterized in that: The step of converting the multi-round conversation texts to be recognized into a second text feature vector comprises: Obtaining the conversation turns of the multi-round conversation text to be recognized; Based on the conversation rounds, segmenting the multi-round conversation texts to obtain multiple single-round conversation texts; Converting the plurality of single-turn conversation texts into a plurality of target text feature vectors; Based on the dialogue rounds, a plurality of the target text feature vectors are concatenated to obtain second text feature vectors corresponding to the plurality of dialogue rounds of text.

7. The multi-round dialogue classification method according to any one of claims 1 to 6, characterized in that: The text classification model includes a linear layer, an activation function layer and an output layer; The step of inputting the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-round dialogue texts includes: Performing classification prediction on the first text feature vector based on the linear layer to obtain prediction results of multiple classification labels; Based on the activation function layer, probability prediction is performed on the prediction results of the multiple classification labels to obtain distribution probability data of the multiple classification labels; The distribution probability data of the multiple classification labels are screened based on the output layer to select multiple candidate classification categories of the multiple rounds of dialogue texts from the multiple classification labels.

8. A multi-round dialogue classification device, characterized in that: The multi-round dialogue classification device comprises: An acquisition module, used for acquiring a first text feature vector of a multi-round conversation text to be recognized; A classification module, used for inputting the first text feature vector into a preset text classification model for processing to obtain multiple candidate classification categories of the multi-round dialogue text; A generating module, configured to generate question information for input into a large language model based on the multi-round dialogue texts and a plurality of candidate classification categories of the multi-round dialogue texts; An input module is used to input the question information into the large language model, so that the large language model determines the target classification category of the multi-round dialogue text based on the question information.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the multi-round dialogue classification method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the multi-round dialogue classification method according to any one of claims 1 to 7 is implemented.