Text intent classification method, device, computer equipment and storage medium
User conversations are processed through text completion and fragment recognition models, and matched with preset intention clusters, which solves the problem that intelligent customer service cannot recognize user intentions and improves user experience.
Patent Information
- Application Number
- CN202210897512.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-28
AI Technical Summary
In the prior art, due to the existence of pronouns in user conversations, intelligent customer service cannot accurately identify user intentions, resulting in the inability to pop up related business speeches in a timely manner, affecting the user experience.
By obtaining the current statement text for text completion, using the fragment recognition model for fragment recognition, and clustering the recognition results, combining preset general and scene intention clusters to match, obtain intent classification results.
Improve user experience, avoid intent recognition efficiency caused by referential pronouns, and achieve more accurate intent classification.
Smart Images

Figure CN115203372B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a text intent classification method, device, computer equipment and storage medium. Background Art
[0002] With the rapid development of the Internet, artificial intelligence has developed rapidly and has been widely used in various fields. Especially in intelligent customer service, it usually recognizes the intention of user questions and automatically pops up relevant business words to assist manual customer service or directly answer user questions.
[0003] In the prior art, when there are pronouns in the user's conversation, the intelligent customer service is often unable to accurately identify the user's intention, resulting in the intelligent customer service being unable to pop up relevant business scripts in a timely manner, or even unable to answer the user's questions, which seriously affects the user experience. Summary of the invention
[0004] The present invention provides a text intent classification method, device, computer equipment and storage medium, which solves the problem of low recognition accuracy caused by the existence of pronouns in text intent classification.
[0005] A text intent classification method, comprising:
[0006] Acquire the current sentence text, and perform text completion on the current sentence text to obtain the target sentence text;
[0007] Inputting the target sentence text into a segment recognition model, performing segment recognition on the target sentence text through the segment recognition model, and obtaining a first segment recognition result, a second segment recognition result, and a third segment recognition result;
[0008] Clustering the first segment recognition result, the second segment recognition result, and the third segment recognition result respectively to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result, and a third text intent cluster corresponding to the third segment recognition result;
[0009] Obtain a preset general intent cluster and a preset scenario intent cluster, and match the first text intent cluster, the second text intent cluster, and the third text intent cluster according to the preset general intent cluster and the preset scenario intent cluster to obtain an intent classification result corresponding to the current sentence text.
[0010] A text intent classification device, comprising:
[0011] An acquisition module is used to acquire the current sentence text and perform text completion on the current sentence text to obtain the target sentence text;
[0012] a recognition module, configured to input the target sentence text into a segment recognition model, perform segment recognition on the target sentence text through the segment recognition model, and obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result;
[0013] a clustering module, used to cluster the first segment recognition result, the second segment recognition result and the third segment recognition result respectively to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result and a third text intent cluster corresponding to the third segment recognition result;
[0014] The result module is used to obtain a preset general intent cluster and a preset scenario intent cluster, and match the first text intent cluster, the second text intent cluster and the third text intent cluster according to the preset general intent cluster and the preset scenario intent cluster to obtain the intent classification result corresponding to the current sentence text.
[0015] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned text intent classification method when executing the computer program.
[0016] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text intent classification method as described above is implemented.
[0017] The present invention provides a text intent classification method, device, computer equipment and storage medium. The present invention achieves the acquisition of the target sentence text by performing text completion on the current sentence text, thereby avoiding the low efficiency of intent recognition caused by the presence of pronouns in the current sentence text. By inputting the target sentence text into a fragment recognition model and performing fragment recognition on the target sentence text according to the fragment recognition model, the acquisition of the fragment recognition result is achieved. By clustering the first fragment recognition result, the second fragment recognition result and the third fragment recognition result respectively, the acquisition of the text intent cluster is achieved. According to the preset general intent cluster and the preset scene intent cluster, the first text intent cluster, the second text intent cluster and the third text intent cluster are matched to achieve the acquisition of the intent classification result, further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0019] Figure 1 is a schematic diagram of an application environment of a text intent classification method in one embodiment of the present invention;
[0020] Figure 2 is a flow chart of a text intent classification method according to an embodiment of the present invention;
[0021] Figure 3 is a flowchart of step S1 of a text intent classification method in one embodiment of the present invention;
[0022] Figure 4 is a flowchart of step S2 of the text intent classification method according to an embodiment of the present invention;
[0023] Figure 5 is a principle block diagram of a text intent classification device according to an embodiment of the present invention;
[0024] Figure 6 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] The text intent classification method provided by the embodiment of the present invention can be applied as follows: Figure 1 Specifically, the text intent classification method is applied in a text intent classification device, and the text intent classification device includes: Figure 1The client and server shown in the figure communicate with each other through the network, which is used to solve the problem of low accuracy of intent recognition due to the presence of pronouns in text intent classification in the prior art. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The client can be installed on, but not limited to, various personal computers, laptops, smart phones, tablets, and portable wearable devices.
[0027] In one embodiment, if Figure 2 As shown, a text intent classification method is provided, which is applied in Figure 1 The server in the example is used as an example to illustrate the process, including the following steps:
[0028] S1, obtaining a current sentence text, and performing text completion on the current sentence text to obtain a target sentence text.
[0029] Specifically, all current sentence texts are retrieved from the server, all current sentence texts are spliced in text order, and the pronouns in the current sentence text are replaced, and the pronouns are replaced with referent entities. After all the pronouns are replaced in sequence, the target sentence text can be obtained. Among them, pronouns are also called demonstrative pronouns, which are words or phrases with referential functions, such as this, that, you and his words or phrases. The referent entity is the word or word or thing or person represented by the pronoun. For example, the referent entity of this can be an object or a person, and the specific meaning is determined in combination with context information. The current sentence text is the text of the conversation between the current user and the manual customer service. The target sentence text is the text after splicing and completing the current sentence text.
[0030] S2, inputting the target sentence text into a segment recognition model, performing segment recognition on the target sentence text through the segment recognition model, and obtaining a first segment recognition result, a second segment recognition result, and a third segment recognition result.
[0031] Understandably, the fragment recognition model is a model for recognizing the target sentence text, and the model can be a neural network model obtained by training with a large amount of data. The first fragment recognition result is the result of recognizing the main process intention. The second fragment recognition result is the result of recognizing the ambiguous process intention. The third fragment recognition result is the result of recognizing the end process intention. The main process intention is the process of replying in a preset order, such as the process of introducing a product. The ambiguous process intention is the process of replying that deviates from the main process intention, such as the process of asking about other products. The end process intention is the process of not being interested in the question asked, such as the process where the user answers that he is busy.
[0032] Specifically, after obtaining the target sentence text, the target sentence text is input into the fragment recognition model, and the target sentence text is processed by different modules in the fragment recognition model. When the target sentence text conforms to the main process intention, a first fragment recognition result representing the main process intention is obtained. When the conversation sentence text in the target sentence text conforms to the ambiguous process intention, a second fragment recognition result representing the ambiguous process intention is obtained. When the conversation sentence text in the target sentence text conforms to the end process intention, a third fragment recognition result representing the end process intention is obtained.
[0033] S3, clustering the first segment recognition result, the second segment recognition result and the third segment recognition result respectively to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result and a third text intent cluster corresponding to the third segment recognition result.
[0034] Specifically, after obtaining the fragment recognition result, the first fragment recognition result, the second fragment recognition result and the third fragment recognition result are clustered in turn. First, all the target sentence texts in the first fragment recognition result are clustered by a clustering algorithm to obtain a first text intent cluster corresponding to the first fragment recognition result. Secondly, all the target sentence texts in the first fragment recognition result are clustered by a clustering algorithm to obtain a second text intent cluster corresponding to the second fragment recognition result. Finally, all the target sentence texts in the first fragment recognition result are clustered by a clustering algorithm to obtain a third text intent cluster corresponding to the third fragment recognition result. Among them, the text intent cluster is a collection of all sentence texts with the same process meaning, such as sentence texts such as inconvenient and busy. A clustering algorithm is an algorithm used to cluster texts or words or vectors with the same meaning.
[0035] S4, obtain a preset general intent cluster and a preset scenario intent cluster, match the first text intent cluster, the second text intent cluster and the third text intent cluster according to the preset general intent cluster and the preset scenario intent cluster, and obtain the intent classification result corresponding to the current sentence text.
[0036] Specifically, after obtaining the text intent cluster, the pre-set general intent cluster and the pre-set scene intent cluster are retrieved from the server, and each text intent cluster is first matched with the preset general intent cluster, that is, it is determined whether the similarity between the text intent cluster and the preset general intent cluster exceeds the preset threshold. When the similarity between the text intent cluster and the preset general intent cluster exceeds the preset threshold, an intent matching result representing a successful match is obtained. When the similarity between the text intent cluster and the preset general intent cluster does not exceed the preset threshold, an intent matching result representing a failed match is obtained. Among them, the preset threshold is a threshold set in advance for judging the similarity between intent clusters. The preset general intent cluster is an intent cluster set in advance that can be used in multiple scenarios. The preset scene intent cluster is an intent cluster set in advance that can be used in a certain scene. The intent matching result is a result used to characterize the similarity between intent clusters.
[0037] In this way, the present invention achieves the acquisition of the target sentence text by completing the text of the current sentence, thereby avoiding the inability to recognize the user's intention due to the presence of pronouns in the current sentence text. By inputting the target sentence text into a fragment recognition model and performing fragment recognition on the target sentence text according to the fragment recognition model, the acquisition of the fragment recognition result is achieved. By clustering the first fragment recognition result, the second fragment recognition result, and the third fragment recognition result respectively, the acquisition of the text intent cluster is achieved. According to the preset general intent cluster and the preset scene intent cluster, the first text intent cluster, the second text intent cluster, and the third text intent cluster are matched to achieve the acquisition of the intent classification result, further improving the user experience.
[0038] In one embodiment, if Figure 3 As shown, in step S1, the current sentence text is completed to obtain the target sentence text, including:
[0039] S11, obtaining a historical sentence text and a historical reply text corresponding to the historical sentence text; the historical sentence text refers to the text of the previous round of the current sentence text; the historical sentence text and the historical reply text correspond to a historical text label;
[0040] S12, concatenating the historical sentence text, the historical reply text and the current sentence text to obtain an initial text.
[0041] Understandably, the historical sentence text refers to the text of the previous round of the current sentence text, that is, the current sentence text of the previous round. The historical reply text is the text corresponding to the customer service reply to the historical sentence text. The historical text label is a label of the historical sentence text used to distinguish the historical sentence text from the current sentence text, such as token type 0. The historical sentence text and the historical reply text correspond to one historical text label.
[0042] Specifically, after obtaining the current sentence text, the historical sentence text and the historical reply text corresponding to the historical sentence text are retrieved from the server, and historical text tags are set for the historical sentence text and the historical reply text, such as setting tokentype to 0, that is, the text tags of the historical sentence text and the historical reply text are both 0. All historical sentence texts, historical reply texts and the current sentence text are connected in the order of time obtained, and the historical sentence texts, historical reply texts and the current sentence text are separated by separators (such as spaces or / , etc.), and the initial text can be obtained.
[0043] S13, obtaining a current text label corresponding to the current sentence text, and concatenating the historical text label and the current text label to obtain an initial label.
[0044] Understandably, the current text label is a label of the current sentence text used to distinguish the historical sentence text from the current sentence text, such as token type is 1. The initial label is a label obtained by connecting the historical text label and the current text label, such as the initial label is 0, 0, 1.
[0045] Specifically, after obtaining the initial text, the label of the current sentence text is set, such as setting the token type to 1, and the label setting of the current sentence text is completed. According to the order of the historical sentence text, the historical reply text and the current sentence text in the initial text, the historical text label and the current text label are sequentially spliced to obtain the initial label.
[0046] S14, inputting the initial text and the initial label into a preset text query model, obtaining the referent entity position corresponding to the initial text output by the preset text query model, and the to-be-completed position corresponding to the current sentence text.
[0047] Understandably, the preset text query model is a model set in advance for finding the position of the referent entity and the position to be completed in the initial text. The referent entity position is the position of the entity referred to by the pronoun in the initial text. The position to be completed is the position that needs to be completed in the initial text or the position of the pronoun in the initial text.
[0048] Specifically, after obtaining the initial label, the initial text and the initial label are input together as input into the preset text query model. The preset text query model first encodes the initial text, and then maps it to a high-dimensional space through a three-layer transformation module. Then, the fully connected layer and the normalized layer are used to query the referent entity position and the position to be completed in the current sentence text to obtain the referent entity position in the initial text and the position to be completed in the current sentence text. The probability of the position to be completed and the referent entity position corresponding to the position to be completed is calculated, and the position to be completed with the highest probability is associated with the referent entity position, and all the positions to be completed in the current sentence text and the referent entity positions corresponding to each position to be completed are associated in turn. Among them, a referent entity position can be associated with at least one position to be completed, and a position to be completed can only be associated with one referent entity position.
[0049] S15, extracting the referential entity text corresponding to the referential entity position from the initial text, and performing text completion on the current sentence text according to the referential entity text and the position to be completed to obtain a target sentence text.
[0050] Understandably, the entity-referring text is the text that refers to the entity position. The target sentence text is the current sentence text after completing the position to be completed.
[0051] Specifically, after obtaining the referent entity position and the position to be completed, the referent entity text corresponding to the referent entity position is extracted from the initial text, that is, the referent entity text is extracted from the referent entity position. First, determine the starting position and the ending position of the referent entity text to obtain the length of the referent entity text. Then determine the starting position and the ending position of the position to be completed to obtain the length of the position to be completed. Then, based on the obtained associated information, all the positions to be completed associated with the referent entity position are obtained, and the current sentence text is completed, that is, the referent entity text is filled or replaced to the position to be completed. After all the positions to be completed are filled or replaced in turn, the target sentence text can be obtained.
[0052] Furthermore, when the user and customer service are discussing helmets, the user asks how much it costs or how much this blue one costs. In the absence of a subject, the intelligent customer service cannot recognize it and needs to fill in or replace the helmet to the position to be completed, such as how much the helmet costs or how much the blue helmet costs. The user's intention can be quickly identified.
[0053] The embodiment of the present invention obtains the current sentence text, the historical sentence text, and the historical reply text corresponding to the historical sentence text, and concatenates and sets labels to achieve the acquisition of the initial text and the initial label. The initial text and the initial label are simultaneously input into the preset text query model, and the reference entity position and the position to be completed are obtained through the preset text query model, thereby achieving the acquisition of the reference entity position in the initial text and the position to be completed in the current sentence text. The target sentence text is acquired by extracting the reference entity text from the reference entity position and supplementing the reference entity text to the position to be completed according to the reference completion position. This further avoids the problem that the intelligent customer service cannot reply to the user due to the missing subject.
[0054] In one embodiment, before step S1, that is, before obtaining the current sentence text, the following steps are included:
[0055] S16, obtaining an initial sentence text, performing word segmentation processing on the initial sentence text, and obtaining at least one word to be processed in the initial sentence text.
[0056] Specifically, the initial sentence text is retrieved from the server, the initial sentence text is segmented by the Chinese word segmentation algorithm, and the initial sentence text is fully segmented and segmented according to the connection of context features to obtain at least one word to be processed corresponding to the initial sentence text. The full segmentation path selection segmentation process is to list all possible segmentation results, select the best segmentation path from them, and form a directed acyclic graph with all the segmentation results. The segmentation results can be used as nodes, and the edges between words are weighted. The path with the smallest weight is the final result. For example, the word frequency can be used as the weight, and a path with the largest total word frequency can be found to be the best path. Among them, the words to be processed are the results of the segmentation of the initial call text, the segmentation results are the words to be processed obtained after segmentation, and the directed acyclic graph is a graph without loops and directions. The initial sentence text is the text of the conversation between the user and the customer service.
[0057] S17, performing entity recognition on the word to be processed to obtain an entity recognition result corresponding to the word to be processed.
[0058] Specifically, after obtaining the words to be processed, all the words to be processed are tagged with parts of speech through a part-of-speech coding table, and each word or phrase is tagged with a part-of-speech label, such as adjective, verb, noun, etc., so that the words to be processed can be integrated with more useful information in the subsequent processing. The text of the sentence to be processed after the part-of-speech tagging of each word to be processed is input into the entity recognition model, and the entity recognition model is used to perform entity recognition on the sentence text to be processed, such as determining the entity type of each word to be processed according to the part-of-speech of each word to be processed, and then determining the entity type as the entity recognition result, that is, according to the context features, the connection between the parts of speech of sentences and words, extract important entity information from the given sentence text, such as time, place, person, etc., time can be a time entity, place can be a place entity, and person can be a name entity, etc. Among them, the entity recognition result is the entity information extracted from the sentence text. The entity recognition model can be obtained by supervised training of a model built based on a neural network. Part-of-speech tagging is to set a part-of-speech label for a word according to the part-of-speech coding table. Entity recognition is the process of extracting entity information from the sentence text.
[0059] S18, filtering the initial sentence text according to the entity recognition results corresponding to each word to be processed to obtain the current sentence text.
[0060] Specifically, after obtaining the entity recognition results corresponding to each word to be processed, the stop words and modal particles, noise words, low-frequency words, and incoherent sentences in all entity recognition results are filtered out through a pre-set dictionary library, and all entity recognition results obtained after filtering are sorted into sentence texts to obtain the target call text. Among them, the deletion of stop words is determined according to the specific scenario. For example, in some sentence texts of sentiment analysis, modal particles and exclamation marks should be retained because they have certain significance in expressing the degree of tone and emotional color.
[0061] The embodiment of the present invention achieves the acquisition of the words to be processed in the initial sentence text by performing word segmentation processing on the initial call text. Part-of-speech tagging is performed on the words to be processed through a part-of-speech coding table and entity recognition is performed on the words to be processed through an entity recognition model, thereby achieving the acquisition of entity recognition results. The initial sentence text is filtered through the entity recognition results corresponding to each word to be processed, thereby achieving the acquisition of the current sentence text.
[0062] In one embodiment, if Figure 4 As shown, in step S2, the target sentence text is subjected to segment recognition by the segment recognition model to obtain a first segment recognition result, a second segment recognition result and a third segment recognition result, including:
[0063] S21, encoding the target sentence text through the encoding module in the fragment recognition model to obtain a target word vector.
[0064] It can be understood that the encoding module is a module in the segment recognition model for encoding text. The target word vector is a vector obtained by encoding the words in the target sentence text.
[0065] Specifically, after obtaining the target sentence text, the target call text is encoded by the encoding module in the fragment recognition model. When the encoding module is word embedding encoding, the target call text is first segmented, and the same words are eliminated to obtain at least one segmentation result in the target sentence text. According to the number of segmentation results M, N-dimensional vectors of different dimensions are used to represent them. N can be 64, 128, 256, 512, etc., that is, each segmentation result is represented by multiple numbers, and then the numbers are represented by vectors to obtain the target word vector.
[0066] S22, transforming the target word vector through the transformation module in the fragment recognition model to obtain a target sentence vector.
[0067] It can be understood that the transformation module is a module in the segment recognition model for transforming the word vector. The target word vector is the result of transforming the target word vector.
[0068] Specifically, after obtaining the target word vector, the target word vector is transformed by the transformation module in the segment recognition model. When the transformation module is a transformer encoder, all target word vectors are input into the transformer encoder, and the target sentence vector is obtained after being processed by three layers of transformer encoders. The transformation process of this embodiment is to add positional encoding to the target word vector of each word segmentation result. Then all target word vectors are passed through multiple groups of three weight matrices W. Q , W K , W V Calculate and obtain multiple sets of Query, Keys, and Values vectors.
[0069] Furthermore, the correlation score between the target word vectors is calculated using the dot product method, that is, the dot product is calculated using each target word vector in Q and each target word vector in K. Specifically, in the form of a matrix: score = Q*K T , where socre is a (2, 2) matrix. Normalize the correlation scores between target word vectors in the input sequence, score = score*sqrt(d k ), d kis the dimension of K. When d k When the dimension is 128, the score vector between the target word vectors is converted into a probability distribution between [0, 1] through the softmax function. After softmax, the score is converted into a (2, 2)α probability distribution matrix with values distributed between [0, 1]. According to the probability distribution between the target word vectors, the corresponding Values value is multiplied, and α is dot-producted with V, Z = softmax(score)*V, the dimension of V is (2, 128), (2, 2)*(2, 128), and the final Z is a (2, 128)-dimensional matrix. The multiple Z matrices obtained are concatenated, and the above process is repeated three times to obtain the target sentence vector.
[0070] S23, obtaining a target position vector, and performing segment recognition on the target sentence text according to the target sentence vector and the target position vector to obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result.
[0071] It can be understood that the target position vector is a vector used to represent the position of the target sentence vector in the target sentence text.
[0072] Specifically, after obtaining the target sentence vector, positional encoding is added to the target word vector of each word segmentation result, and the target sentence vector and target position vector are used as inputs through three layers of transformer encoder, and then input into the fully connected layer and softmax layer, and the target sentence vector and target position vector are processed by the fully connected layer and softmax layer to obtain the fragment recognition result. Among them, the process of obtaining the target position vector is to obtain the position of the target sentence vector through sine and cosine position encoding, and represent the position of the target sentence vector as a vector to obtain the target position vector.
[0073] Furthermore, in this embodiment, the position encoding in the transformer encoder is generated by using sine and cosine functions of different frequencies, and then added to the target word vector of the corresponding position. The dimension of the position vector must be consistent with the dimension of the target word vector. The processing process of the fully connected layer and the softmax layer is to first linearly change the matrix Z, then nonlinearly change it, and then linearly transform it to obtain a higher-dimensional space. The transformed matrix is normalized by the softmax layer to obtain the first segment recognition result, the second segment recognition result, and the third segment recognition result.
[0074] The embodiment of the present invention performs encoding and change processing on the target sentence text through the encoding module in the segment recognition model, thereby obtaining the target word vector and the target sentence vector corresponding to the target sentence text. The segment recognition result is obtained by obtaining the target position vector in the target sentence text and performing segment recognition on the target sentence text according to the target sentence vector and the target position vector.
[0075] In one embodiment, the step S3, i.e. clustering the first segment recognition result, the second segment recognition result and the third segment recognition result, includes:
[0076] S31, input the first fragment recognition result, the second fragment recognition result and the third fragment recognition result into a preset encoding model, and encode the target sentence text corresponding to each fragment recognition result through the preset encoding model to obtain a first text semantic vector corresponding to the first fragment recognition result, a second text semantic vector corresponding to the second fragment recognition result, and a third text semantic vector corresponding to the third fragment recognition result.
[0077] Understandably, the preset encoding model is a model set in advance for converting the segment recognition result into a vector, such as a BERT model. The text semantic vector extracts semantic information from the target sentence text corresponding to the segment recognition result and expresses it in vector form.
[0078] Specifically, after obtaining the fragment recognition result, the first fragment recognition result, the second fragment recognition result and the third fragment recognition result are input into the preset coding model in sequence, that is, the fragment recognition result representing the main process intention, the fragment recognition result representing the ambiguous process intention and the fragment recognition result representing the end process intention are input into the preset coding model, and the preset coding model is used to extract semantic information of the target sentence text corresponding to the first fragment recognition result, the second fragment recognition result and the third fragment recognition result, and then the semantic information is encoded to obtain the first text semantic vector corresponding to the first fragment recognition result, the second text semantic vector corresponding to the second fragment recognition result and the third text semantic vector corresponding to the third fragment recognition result.
[0079] Furthermore, when the preset encoding model is the BERT model, BERT is used as a feature extractor, a target sentence text S is input, and the target sentence text is segmented through the components in the BERT model to obtain S = [CLS, tok1, tok2, ..., tokN, SEP]. The result is then subjected to Dense processing to make the originally sparse text dense, and the text is vectorized to output a vector sequence. The vector sequence is used as a text semantic vector to obtain the text semantic vector corresponding to the target sentence text of each segment recognition result.
[0080] S32, clustering the text semantic vectors based on a clustering algorithm to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result, and a third text intent cluster corresponding to the third segment recognition result.
[0081] It can be understood that the clustering algorithm is an algorithm that merges texts, sentences or vectors with the same meaning, such as the k-means clustering algorithm.
[0082] Specifically, after obtaining the text semantic vector, the first text semantic vector corresponding to the first fragment recognition result, the second text semantic vector corresponding to the second fragment recognition result, and the third text semantic vector corresponding to the third fragment recognition result are clustered in turn through a clustering algorithm to obtain a first text intent cluster corresponding to the first fragment recognition result, a second text intent cluster corresponding to the second fragment recognition result, and a third text intent cluster corresponding to the third fragment recognition result.
[0083] Furthermore, when the clustering algorithm is the k-means clustering algorithm, the process of this embodiment is to divide the text semantic vectors corresponding to the fragment recognition results into K groups, and randomly select N (N is less than or equal to K) text semantic vectors as the initial cluster centers. Then, the distance between each text semantic vector and each cluster center is calculated by Euclidean distance or cosine similarity, and each text semantic vector is assigned to the cluster center closest to it. The cluster center and the assigned text semantic vector represent a cluster. Each time a text semantic vector is assigned, the cluster center of the cluster will be recalculated based on the existing text semantic vectors in the cluster. This process will be repeated until a certain termination condition is met. The termination condition can be that no (or a minimum number) objects are reassigned to different clusters, or that no (or a minimum number) cluster centers change again, or that the sum of squared errors is locally minimized.
[0084] The embodiment of the present invention encodes the target sentence text corresponding to the segment recognition result through a preset encoding model, thereby obtaining the text semantic vector corresponding to each segment recognition result. The text semantic vector is clustered through a clustering algorithm, thereby obtaining the text intention cluster corresponding to each segment recognition result.
[0085] In one embodiment, in step S4, matching the first text intent cluster, the second text intent cluster, and the third text intent cluster to obtain the intent classification result corresponding to the current sentence text includes:
[0086] S41, performing vector extraction on the preset general intent cluster and each text intent cluster to obtain a general semantic vector corresponding to the preset general intent cluster and a text semantic vector corresponding to each text intent cluster.
[0087] Understandably, the general semantic vector is the central semantic vector of a preset general intent cluster. The text semantic vector is the central semantic vector of a text intent cluster. The central semantic vector is in the form of a vector representing the semantic information of the intent cluster. The semantic information is a word or phrase or sentence or text that can represent the intent cluster.
[0088] Specifically, after obtaining the text intent clusters, the central semantic vectors of the first text intent cluster, the second text intent cluster, and the third text intent cluster are extracted. The TextRank algorithm can be used to first extract the keywords or key sentences in the first text intent cluster, the second text intent cluster, and the third text intent cluster, and then the keywords or key sentences are vectorized to obtain the text semantic vector. The first text intent cluster, the second text intent cluster, and the third text intent cluster can also be vectorized first, and then the vectorized text intent clusters are feature extracted through the Mel-frequency cepstral coefficients to obtain the text semantic vector. Similarly, the central semantic vector of the preset general intent cluster is extracted to obtain a general semantic vector. Among them, keyword or key sentence extraction through the TextRank algorithm and feature extraction through the Mel-frequency cepstral coefficients are common methods and will not be repeated here.
[0089] S42, matching the general semantic vector with all the text semantic vectors to obtain a general classification result.
[0090] Specifically, after obtaining the universal semantic vector and the text semantic vector, the text semantic vector and the universal semantic vector are matched for similarity, and all the text semantic vectors and the universal semantic vector are matched for similarity in turn. When the cosine similarity is adopted in this embodiment, the cosine value between the universal semantic vector and the text semantic vector is calculated, and it is determined whether the cosine value between the universal center semantic vector and the text center semantic vector exceeds a preset threshold. And it is determined in turn whether the cosine value between all the text semantic vectors and the universal semantic vector exceeds a preset threshold.
[0091] Furthermore, when the cosine value between the general semantic vector and the text semantic vector is greater than or equal to a preset threshold, an intention matching result representing a successful match is obtained, that is, the text intent cluster is determined as a through-use intent cluster. When the cosine value between the general semantic vector and the text semantic vector is less than the preset threshold, an intention matching result representing a failed match is obtained, that is, the text intent cluster does not belong to the through-use intent cluster. Among them, the general classification result is a result used to characterize the similarity between the general semantic vector and the text semantic vector. The preset threshold is a threshold set in advance for judging the similarity between vectors. The vector threshold can be set according to actual conditions, such as when the Euclidean distance is used for calculation, it is set to the corresponding distance threshold, and when the cosine similarity is used for calculation, it is set to the corresponding cosine threshold.
[0092] The embodiment of the present invention achieves acquisition of general semantic vectors and text semantic vectors by performing vector extraction on preset general intent clusters and text intent clusters, and acquires general classification results by performing similarity matching on general semantic vectors and text semantic vectors.
[0093] In one embodiment, the step S4, that is, matching the first text intent cluster, the second text intent cluster, and the third text intent cluster to obtain the intent classification result corresponding to the current sentence text, further includes:
[0094] S51, record the text intent cluster corresponding to the intent classification result representing the matching failure as a matching intent cluster; perform vector extraction on the preset scene intent cluster and each matching intent cluster to obtain a scene semantic vector corresponding to the preset scene intent cluster and a matching semantic vector corresponding to each matching intent cluster.
[0095] It can be understood that the matching intent cluster is the text intent cluster corresponding to the intent classification result that represents the failed matching. The scene semantic vector is the central semantic vector of the preset scene intent cluster. The matching semantic vector is the central semantic vector of the matching intent cluster.
[0096] Specifically, after obtaining the general classification result, the text intent cluster corresponding to the intent matching result representing the failed match is determined as the matching intent cluster, that is, all text intent clusters that do not belong to the used intent cluster are determined as matching intent clusters. The matching intent cluster and the preset scene intent cluster are input into a preset label model (such as the mT5 model), and the keywords or key words of the matching intent cluster and the preset scene intent cluster are extracted through the preset label model, and the extracted keywords or key words are used as the intent labels of the matching intent cluster and the preset scene intent cluster to obtain matching intent labels and scene intent labels. The obtained matching intent labels and scene intent labels are vectorized to obtain the scene semantic vector corresponding to the preset scene intent cluster and the matching semantic vector corresponding to each matching intent cluster. All matching intent clusters and all preset scene intent clusters are input into the preset label model in turn to obtain the matching intent labels corresponding to each matching intent cluster and the scene intent labels corresponding to each preset scene intent cluster. All matching intent labels are vectorized to obtain the matching semantic vectors corresponding to each matching intent cluster.
[0097] S52, matching the scene semantic vector with the matching semantic vector to obtain an intent classification result.
[0098] Specifically, after obtaining the scene semantic vector and the matching semantic vector, all the scene semantic vectors and all the matching semantic vectors are input into a preset matching model (such as the RE2 model), and the similarity between each scene semantic vector and each matching semantic vector is calculated by the preset matching model. The similarity between the scene semantic vector and each matching semantic vector can be determined by calculating the Euclidean distance (or cosine similarity) between the scene intent cluster and each matching intent cluster, and by judging whether the Euclidean distance (or cosine similarity) exceeds a preset threshold, it is determined whether the matching intent cluster is recorded as a scene intent cluster. The similarity between all scene semantic vectors and matching semantic vectors is calculated in sequence.
[0099] Furthermore, when the Euclidean distance (or cosine similarity) is greater than or equal to a preset threshold, an intent classification result representing a successful match is obtained, that is, the matching intent cluster is recorded as a scene intent cluster. When the Euclidean distance (or cosine similarity) is less than the preset threshold, an intent classification result representing a failed match is obtained, that is, the text intent cluster neither meets the general intent cluster nor the scene intent cluster, and the text intent cluster is then determined as a new scene intent cluster. When matching the text intent clusters in the next round, it is necessary to perform a similarity match between the text intent clusters in the next round and the new scene intent clusters to determine whether they belong to the new scene intent cluster. Among them, the intent classification result is a result used to characterize the similarity between the scene semantic vector and the matching semantic vector.
[0100] The embodiment of the present invention determines the text intent cluster corresponding to the intent matching result representing the failed match as the matching intent cluster, generates the matching intent label of the matching intent cluster and the scene intent label of the scene intent cluster according to the preset label model, and vectorizes the matching intent label and the scene intent label, thereby achieving the acquisition of the scene semantic vector and the matching semantic vector. The scene semantic vector and the matching semantic vector are matched through the preset matching model, thereby achieving the acquisition of the intent classification result.
[0101] In one embodiment, a text intent classification device is provided, which corresponds one-to-one to the text intent classification method in the above embodiment. Figure 5 As shown, the text intent classification device includes an acquisition module 11, a recognition module 12, a clustering module 13 and a result module 14. The functional modules are described in detail as follows:
[0102] The acquisition module 11 is used to acquire the current sentence text and perform text completion on the current sentence text to obtain the target sentence text;
[0103] A recognition module 12 is used to input the target sentence text into a segment recognition model, perform segment recognition on the target sentence text through the segment recognition model, and obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result;
[0104] A clustering module 13, used to cluster the first segment recognition result, the second segment recognition result and the third segment recognition result respectively to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result and a third text intent cluster corresponding to the third segment recognition result;
[0105] The result module 14 is used to obtain a preset general intent cluster and a preset scenario intent cluster, and match the first text intent cluster, the second text intent cluster and the third text intent cluster according to the preset general intent cluster and the preset scenario intent cluster to obtain the intent classification result corresponding to the current sentence text.
[0106] In one embodiment, the acquisition module 11 includes:
[0107] An acquisition unit, used for acquiring a historical sentence text and a historical reply text corresponding to the historical sentence text; the historical sentence text refers to the text of the previous round of the current sentence text; the historical sentence text and the historical reply text correspond to a historical text label;
[0108] A concatenation unit, used for concatenating the historical sentence text, the historical reply text and the current sentence text to obtain an initial text;
[0109] A label unit, used to obtain a current text label corresponding to the current sentence text, and to concatenate the historical text label and the current text label to obtain an initial label;
[0110] A query unit, used to input the initial text and the initial label into a preset text query model, and obtain the reference entity position corresponding to the initial text output by the preset text query model, and the position to be completed corresponding to the current sentence text;
[0111] The completion unit is used to extract the referring entity text corresponding to the referring entity position from the initial text, and complete the current sentence text according to the referring entity text and the position to be completed to obtain the target sentence text.
[0112] In one embodiment, the acquisition module 11 further includes:
[0113] A word segmentation unit, used for acquiring an initial sentence text, performing word segmentation processing on the initial sentence text, and obtaining at least one word to be processed in the initial sentence text;
[0114] An entity recognition unit, used for performing entity recognition on the word to be processed to obtain an entity recognition result corresponding to the word to be processed;
[0115] The filtering unit is used to filter the initial sentence text according to the entity recognition results corresponding to each word to be processed to obtain the current sentence text.
[0116] In one embodiment, the identification module 12 includes:
[0117] An encoding unit, configured to encode the target sentence text through an encoding module in the segment recognition model to obtain a target word vector;
[0118] A transformation unit, configured to transform the target word vector through a transformation module in the segment recognition model to obtain a target sentence vector;
[0119] The result unit is used to obtain a target position vector, perform segment recognition on the target sentence text according to the target sentence vector and the target position vector, and obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result.
[0120] In one embodiment, the clustering module 13 further includes:
[0121] a semantic vector unit, configured to input the first segment recognition result, the second segment recognition result, and the third segment recognition result into a preset coding model, and encode the target sentence text corresponding to each segment recognition result respectively through the preset coding model to obtain a first text semantic vector corresponding to the first segment recognition result, a second text semantic vector corresponding to the second segment recognition result, and a third text semantic vector corresponding to the third segment recognition result;
[0122] A vector clustering unit is used to cluster the text semantic vectors based on a clustering algorithm to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result, and a third text intent cluster corresponding to the third segment recognition result.
[0123] In one embodiment, the result module 14 includes:
[0124] An extraction unit, used to extract vectors from the preset general intent cluster and each text intent cluster to obtain a general semantic vector corresponding to the preset general intent cluster and a text semantic vector corresponding to each text intent cluster;
[0125] The matching unit is used to match the general semantic vector with all the text semantic vectors to obtain a general classification result.
[0126] In one embodiment, the result module 14 further includes:
[0127] A recording unit, used to record the text intent cluster corresponding to the intent classification result representing the failed match as a matching intent cluster; perform vector extraction on the preset scene intent cluster and each matching intent cluster to obtain a scene semantic vector corresponding to the preset scene intent cluster and a matching semantic vector corresponding to each matching intent cluster;
[0128] The classification unit is used to match the scene semantic vector with the matching semantic vector to obtain an intent classification result.
[0129] For the specific limitations of the text intent classification device, please refer to the limitations of the text intent classification method above, which will not be repeated here. Each module in the above-mentioned text intent classification device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0130] In one embodiment, a computer device is provided. The computer device may be a client or a server. The internal structure diagram thereof may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the readable storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a text intent classification method is implemented.
[0131] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the text intent classification method in the above embodiment is implemented.
[0132] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the text intent classification method in the above embodiment is implemented.
[0133] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0134] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0135] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A text intent classification method, characterized in that: include: Acquire the current sentence text, and perform text completion on the current sentence text to obtain the target sentence text; Inputting the target sentence text into a segment recognition model, performing segment recognition on the target sentence text through the segment recognition model, and obtaining a first segment recognition result, a second segment recognition result, and a third segment recognition result; Clustering the first segment recognition result, the second segment recognition result, and the third segment recognition result respectively to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result, and a third text intent cluster corresponding to the third segment recognition result; Obtaining a preset general intent cluster and a preset scene intent cluster, and matching the first text intent cluster, the second text intent cluster, and the third text intent cluster according to the preset general intent cluster and the preset scene intent cluster to obtain an intent classification result corresponding to the current sentence text; The performing segment recognition on the target sentence text by using the segment recognition model to obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result includes: The target sentence text is encoded by the encoding module in the segment recognition model to obtain a target word vector; The target word vector is transformed by the transformation module in the segment recognition model to obtain a target sentence vector; Acquire a target position vector, perform segment recognition on the target sentence text according to the target sentence vector and the target position vector, and obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result; The clustering the first segment recognition result, the second segment recognition result and the third segment recognition result comprises: Inputting the first segment recognition result, the second segment recognition result, and the third segment recognition result into a preset encoding model, encoding the target sentence text corresponding to each segment recognition result respectively through the preset encoding model, and obtaining a first text semantic vector corresponding to the first segment recognition result, a second text semantic vector corresponding to the second segment recognition result, and a third text semantic vector corresponding to the third segment recognition result; The text semantic vectors are clustered based on a clustering algorithm to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result, and a third text intent cluster corresponding to the third segment recognition result.
2. The text intent classification method according to claim 1, characterized in that: The performing text completion on the current sentence text to obtain the target sentence text includes: Obtain a historical sentence text and a historical reply text corresponding to the historical sentence text; the historical sentence text refers to the text of the previous round of the current sentence text; the historical sentence text and the historical reply text correspond to a historical text label; Performing text splicing on the historical sentence text, the historical reply text and the current sentence text to obtain an initial text; Obtaining a current text label corresponding to the current sentence text, and concatenating the historical text label and the current text label to obtain an initial label; Inputting the initial text and the initial label into a preset text query model, obtaining the referent entity position corresponding to the initial text output by the preset text query model, and the position to be completed corresponding to the current sentence text; The referential entity text corresponding to the referential entity position is extracted from the initial text, and the current sentence text is completed according to the referential entity text and the position to be completed to obtain a target sentence text.
3. The text intent classification method according to claim 1, characterized in that: The matching of the first text intent cluster, the second text intent cluster, and the third text intent cluster to obtain an intent classification result corresponding to the current sentence text includes: Performing vector extraction on the preset general intent cluster and each text intent cluster to obtain a general semantic vector corresponding to the preset general intent cluster and a text semantic vector corresponding to each text intent cluster; The general semantic vector is matched with all the text semantic vectors to obtain a general classification result.
4. The text intent classification method as claimed in claim 3, characterized in that: The matching of the first text intent cluster, the second text intent cluster, and the third text intent cluster to obtain an intent classification result corresponding to the current sentence text also includes: The text intent cluster corresponding to the intent classification result representing the failed matching is recorded as the matching intent cluster; Performing vector extraction on the preset scene intention cluster and each matching intention cluster to obtain a scene semantic vector corresponding to the preset scene intention cluster and a matching semantic vector corresponding to each matching intention cluster; The scene semantic vector and the matching semantic vector are matched to obtain an intent classification result.
5. The text intent classification method according to claim 1, characterized in that: Before obtaining the current sentence text, the following steps are included: Acquire an initial sentence text, perform word segmentation processing on the initial sentence text, and obtain at least one word to be processed in the initial sentence text; Performing entity recognition on the word to be processed to obtain an entity recognition result corresponding to the word to be processed; According to the entity recognition results corresponding to each word to be processed, the initial sentence text is filtered to obtain the current sentence text.
6. A text intent classification device, characterized in that: include: An acquisition module is used to acquire the current sentence text and perform text completion on the current sentence text to obtain the target sentence text; a recognition module, configured to input the target sentence text into a segment recognition model, perform segment recognition on the target sentence text through the segment recognition model, and obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result; a clustering module, used to cluster the first segment recognition result, the second segment recognition result and the third segment recognition result respectively to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result and a third text intent cluster corresponding to the third segment recognition result; A result module is used to obtain a preset general intent cluster and a preset scene intent cluster, and match the first text intent cluster, the second text intent cluster and the third text intent cluster according to the preset general intent cluster and the preset scene intent cluster to obtain an intent classification result corresponding to the current sentence text; The identification module comprises: An encoding unit, configured to encode the target sentence text through an encoding module in the segment recognition model to obtain a target word vector; A transformation unit, configured to transform the target word vector through a transformation module in the segment recognition model to obtain a target sentence vector; A result unit is used to obtain a target position vector, perform segment recognition on the target sentence text according to the target sentence vector and the target position vector, and obtain a first segment recognition result, a second segment recognition result, and a third segment recognition result; The clustering module comprises: a semantic vector unit, configured to input the first segment recognition result, the second segment recognition result, and the third segment recognition result into a preset coding model, and encode the target sentence text corresponding to each segment recognition result respectively through the preset coding model to obtain a first text semantic vector corresponding to the first segment recognition result, a second text semantic vector corresponding to the second segment recognition result, and a third text semantic vector corresponding to the third segment recognition result; A vector clustering unit is used to cluster the text semantic vectors based on a clustering algorithm to obtain a first text intent cluster corresponding to the first segment recognition result, a second text intent cluster corresponding to the second segment recognition result, and a third text intent cluster corresponding to the third segment recognition result.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the text intent classification method as described in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text intent classification method as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
User intention identification method and device and electronic equipment
CN111581388A
Intention recognition method and device, computer equipment and storage medium
CN114528844A