A semantic classification method and device

By adjusting the attention algorithm of the classification model, calculating the weights based on the token length ratio of the target text and training samples, and generating the target feature vector, the problem of low classification accuracy of long texts is solved, and higher classification accuracy and generalization ability are achieved.

CN117033630BActive Publication Date: 2025-09-16HISENSE GRP HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310955924.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2025-09-16
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

In the existing technology, the classification model has low classification accuracy for long texts, especially when the token length of the target text exceeds 512, the classification effect drops significantly.

Method used

By obtaining the first eigenvector of the target text, determining the target token length and the average token length of the training sample text, calculating the target weight of the scaling factor, and combining the preset attention algorithm and the second eigenvector, the target feature vector for classification is generated to improve the accuracy and generalization ability of the classification model.

Benefits of technology

By adjusting the entropy distribution of the attention algorithm, the degree of attention distraction in feature extraction is reduced, the classification accuracy and generalization ability of the classification model for long texts are improved, and the robustness and interpretability of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033630B_ABST
    Figure CN117033630B_ABST
Patent Text Reader

Abstract

The present application relates to the field of natural language processing technology, and in particular to a semantic classification method and device. In an embodiment of the present application, the classification model determines the target weight of the scaling factor based on the target symbol token length corresponding to the target text and the average token length of the training sample text, and determines the target feature vector of the target text based on the target weight, the second feature vector and the preset attention algorithm, so that when texts with different token lengths are subjected to feature extraction, the feature extraction distribution is close, that is, the degree of attention distraction of feature extraction is reduced, and the accuracy and generalization ability of the classification model are improved, so that the embodiment of the present application has robustness, interpretability and reliability, and meets the trustworthy characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a semantic classification method and device. Background Art

[0002] With the development of general artificial intelligence, the application scope of language big models is becoming wider and wider. Language big models can provide different functions based on different prompt words. Therefore, before inputting text into the language big model, a classification model can be used to classify the text, and the prompt words can be determined based on the classification results. Then, the prompt words and text are input into the language big model together, so that the language big model can provide corresponding services.

[0003] In the prior art, a classification model is trained using training text. The classification model segments the input training text into words and treats each word as a token. The classification model is then trained based on the segmented training text. Because the maximum token length of the training text is generally 512, the trained classification model performs well for text within 512 tokens. However, in actual applications, such as in the summary extraction and text translation functions provided by large models, the token length of the target text to be identified can easily exceed 500. This results in a significant decrease in the classification accuracy of the classification model for long texts with more than 512 tokens. Summary of the Invention

[0004] The present application provides a semantic classification method and device to solve the problem of low accuracy in classifying long texts in the semantic classification methods of the prior art.

[0005] In a first aspect, an embodiment of the present application provides a semantic classification method, the method comprising:

[0006] Obtaining a target text to be classified, performing word segmentation on the target text, and determining a first feature vector corresponding to the target text based on a pre-saved correspondence between word segmentation and word vectors;

[0007] The first feature vector is input into the classification model. The classification model determines the target token length and the second feature vector contained in the target text based on the first feature vector, and determines the target weight of the scaling factor based on the target token length and the average token length of the training sample text. According to the scaling factor, the target weight, the second feature vector and the preset attention algorithm, the target feature vector for classification corresponding to the target text is determined, and according to the target feature vector, the target classification result of the target text is determined.

[0008] In a second aspect, an embodiment of the present application further provides an electronic device, which includes a processor, and the processor is used to implement the steps of the semantic classification method as described above when executing a computer program stored in a memory.

[0009] In an embodiment of the present application, an electronic device obtains a target text to be classified, and performs word segmentation on the target text, and determines a first feature vector corresponding to the target text based on a pre-saved correspondence between the word segmentation and the word vector; the first feature vector is input into a classification model, and the classification model determines a target token length and a second feature vector contained in the target text based on the first feature vector, and determines a target weight of a scaling factor based on the target token length and the average token length of a training sample text, and determines a target feature vector for classification corresponding to the target text based on the scaling factor, the target weight, the second feature vector, and a preset attention algorithm, and determines a target classification result of the target text based on the target feature vector. That is to say, in an embodiment of the present application, the classification model determines the target weight of the scaling factor based on the target symbol token length corresponding to the target text and the average token length of the training sample text, and determines the target feature vector of the target text based on the target weight, the second feature vector and the preset attention algorithm, so that when feature extraction is performed on texts with different token lengths, the feature extraction distribution is close, that is, the degree of attention distraction in feature extraction is reduced, and the accuracy and generalization ability of the classification model are improved, so that the embodiment of the present application is robust, interpretable and reliable, and meets the trustworthy characteristics. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A schematic diagram of a semantic classification process provided in an embodiment of the present application;

[0012] Figure 2 A schematic diagram of the application of the classification model provided in the embodiment of the present application;

[0013] Figure 3 A schematic diagram of the existing classification model training process provided by the prior art;

[0014] Figure 4 A schematic diagram of the training process of the classification model provided in the embodiment of the present application;

[0015] Figure 5A flowchart of semantic classification provided in an embodiment of the present application;

[0016] Figure 6 A schematic diagram of the structure of a semantic classification device provided in an embodiment of the present application;

[0017] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0019] When training classification models, existing techniques often discrepancy between the token lengths of sample texts from different categories in the training set and the actual text token lengths used in classification tasks. This results in the accuracy and generalization of the trained classification models decreasing in real-world applications as the token length of the input text changes. The reasons for the trained classification models' inability to accurately classify long texts are, firstly, the use of positional encodings for which they were not trained, and, secondly, the increased amount of information contained in the text used in real-world applications.

[0020] Based on this, in order to improve the accuracy and generalization ability of the classification model, the embodiment of the present application provides a semantic classification method and device.

[0021] In an embodiment of the present application, an electronic device obtains a target text to be classified, and performs word segmentation on the target text, and determines a first feature vector corresponding to the target text based on a pre-saved correspondence between the word segmentation and the word vector; the first feature vector is input into a classification model, and the classification model determines a target token length and a second feature vector contained in the target text based on the first feature vector, and determines a target weight of a scaling factor based on the target token length and the average token length of a training sample text, and determines a target feature vector for classification corresponding to the target text based on the scaling factor, the target weight, the second feature vector, and a preset attention algorithm, and determines a target classification result of the target text based on the target feature vector.

[0022] Figure 1 A schematic diagram of a semantic classification process provided in an embodiment of the present application includes:

[0023] S101: Obtain a target text to be classified, perform word segmentation on the target text, and determine a first feature vector corresponding to the target text based on a pre-stored correspondence between word segmentation and word vectors.

[0024] A semantic classification method provided in an embodiment of the present application is applied to an electronic device, which may be a PC or a server.

[0025] In an embodiment of the present application, the electronic device obtains a target text to be classified, performs word segmentation on the target text, and determines a first feature vector corresponding to the target text.

[0026] The electronic device obtains a target text to be classified, wherein the target text may be obtained by the electronic device performing semantic recognition on received voice information. After the electronic device obtains the target text, the electronic device performs word segmentation on the target text, dividing the target text into at least one word segment.

[0027] Specifically, in an embodiment of the present application, the electronic device can segment the target text using a pre-configured segmentation algorithm. The electronic device can also input the target text into a segmentation model and determine the result output by the segmentation model as the segmentation result of the target text.

[0028] After the electronic device performs word segmentation on the target text, for each word segmentation, the electronic device determines the word vector corresponding to the word segmentation based on the correspondence between the word segmentation and the word vector. The electronic device determines the first feature vector corresponding to the target text based on the word vector of each word segmentation. In this embodiment of the present application, each word segmentation is regarded as a token.

[0029] In this embodiment of the present application, the word vector corresponding to each word segmentation can be a d-dimensional vector, where d is a positive integer.

[0030] For example, in the embodiment of the present application, the electronic device determines that the target text contains n word segments, and then the electronic device determines that the first feature vector corresponding to the target text is x={x1, x2, x3, ..., x n}.

[0031] S102: Input the first feature vector into the classification model, and the classification model determines the target token length and the second feature vector contained in the target text based on the first feature vector, and determines the target weight of the scaling factor based on the target token length and the average token length of the training sample text, and determines the target feature vector for classification corresponding to the target text based on the scaling factor, the target weight, the second feature vector and a preset attention algorithm, and determines the target classification result of the target text based on the target feature vector.

[0032] When existing classification models classify text based on the first feature vector obtained from word segmentation, they directly use existing attention algorithms to extract features from this first feature vector, determine the feature vector used for classification, and then use a classifier to determine the classification result. However, the first feature vector input contains more semantic information about each word segmentation itself, and directly using the first feature vector for feature extraction can lead to insufficient global information when processing long texts.

[0033] In addition, the attention algorithm used by the existing classification model is:

[0034]

[0035] Among them, Q, K, and V are the query corresponding to the first eigenvector, the key corresponding to the first eigenvector, and the value corresponding to the first eigenvector, respectively. They can be regarded as n vectors of dimension d, where the size of n is the token length corresponding to the first eigenvector of the input. is the scaling factor, d=d k is a hyperparameter.

[0036] Assume that the i-th text to be classified is x i , the text x i The token length is j (j <= n), denoted by x i,j The text x i The process of extracting features through the classification model and determining the feature vector used for classification is expressed as f(x i,j ), that is, Attention(Q,K,V) is f(x i,j ), p(f(x i,j )) is the probability distribution of the feature extraction result. Assuming that the probability distribution is continuous, the entropy of the feature extraction of the model is recorded as:

[0037]

[0038] This shows that in the entropy calculation process of the attention algorithm used in existing classification models, an increase in the number of tokens leads to an increase in entropy. For example, this means that attention is concentrated when extracting features from short texts, and attention is dispersed when extracting features from long texts.

[0039] Based on this, in an embodiment of the present application, after the electronic device determines the first feature vector corresponding to the target text, the electronic device inputs the first feature vector into the classification model. The classification model processes the first feature vector to obtain a second feature vector, so that the second feature vector contains not only the semantic information of each word segment, but also the global information of the target text.

[0040] Furthermore, in an embodiment of the present application, the classification model also determines the target token length corresponding to the target text, and determines the target weight of the scaling factor based on the target token length and the average token length of the training sample text corresponding to the classification model. Different token lengths correspond to different weights.

[0041] The classification model extracts features from the second feature vector according to the attention algorithm, the target weight, and the scaling factor to obtain a target feature vector for classification; and classifies the target text according to the target feature vector to obtain a target classification result for the target text.

[0042] Specifically, in an embodiment of the present application, the classification model obtains a pre-configured average token length of training sample text used to train the classification model, and determines the relationship between the target token length of the target text and the average token length. The classification model obtains a target weight corresponding to the stored size relationship, and uses the target weight, a scaling factor, and a preset attention algorithm to perform feature extraction on the second feature vector to determine the target feature vector for classification.

[0043] Among them, the attention algorithm in the classification model provided in the embodiment of the present application is:

[0044] Attention(Q,K,V)=softmax(λ*QK T )V

[0045] Wherein, λ is the scaling factor after adding the target weight, and the meanings of other characters are the same as those in the prior art.

[0046] Figure 2 This is a schematic diagram of the application of the classification model provided in the embodiment of the present application, such as the Figure 2 As shown, the classification model in the electronic device obtains the first feature vector (first feature vector 1, first feature vector 2, first feature vector 3) of the input target text, and the classification model outputs the target classification result corresponding to the target text. If the electronic device is based on the task corresponding to the target classification result and is implemented by the language large model, the electronic device determines the prompt word corresponding to the target classification result, and sends the prompt word and the target text to the language large model, wherein the prompt word can be a prompt word for the translation task, a prompt word for the summary task, or a prompt word for the question-and-answer task. If the electronic device determines that the task corresponding to the target classification result is implemented by other NLP services or models, the electronic device sends the target text to the corresponding NLP service or model (NLP service / model 1, NLP service / model n).

[0047] In an embodiment of the present application, the classification model determines the target weight of the scaling factor based on the target symbol token length corresponding to the target text and the average token length of the training sample text, and determines the target feature vector of the target text based on the target weight, the second feature vector and the preset attention algorithm, so that when feature extraction is performed on texts with different token lengths, the feature extraction distribution is close, that is, the degree of attention distraction in feature extraction is reduced, thereby improving the accuracy and generalization ability of the classification model.

[0048] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiment, in the embodiment of the present application, the classification model determines the target token length contained in the target text according to the first feature vector and the second feature vector includes:

[0049] The hidden layer of the classification model processes the first feature vector to obtain a third feature vector, and determines the dimension of the first feature vector as the target token length;

[0050] The normalization layer of the classification model combines the first eigenvector and the third eigenvector to determine the second eigenvector.

[0051] In an embodiment of the present application, the classification model includes a hidden layer and a normalization layer, wherein the hidden layer is a fully connected structure, and the hidden layer processes the first feature vector to obtain a third feature vector. The third feature vector contains the semantic information represented by each token in the target text. In addition, in an embodiment of the present application, the hidden layer also determines the dimension of the first feature vector and determines the dimension as the target token length of the target text.

[0052] In the embodiment of the present application, the hidden layer is consistent with the existing hidden layer structure.

[0053] In an embodiment of the present application, after the hidden layer of the classification model outputs the third eigenvector, the normalization layer of the classification model merges the input first eigenvector of the classification model and the third eigenvector to obtain a second eigenvector.

[0054] In this embodiment of the present application, the third eigenvector output by the hidden layer can be expressed as x′=FCN(x), where x′ is the third eigenvector and x is the first eigenvector.

[0055] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in an embodiment of the present application, the normalization layer of the classification model merges the first eigenvector and the third eigenvector to determine the second eigenvector, including:

[0056] The normalization layer performs normalization processing on the first eigenvector and the third eigenvector respectively; and determines the mean vector of the normalized first eigenvector and the normalized third eigenvector as the second eigenvector.

[0057] In an embodiment of the present application, the normalization layer of the classification model uses the Softmax function to normalize the first eigenvector and the third eigenvector respectively, and determines the mean vector of the normalized first eigenvector and the normalized third eigenvector as the second eigenvector.

[0058] In this embodiment of the present application, the normalization layer can determine the second eigenvector by the following formula:

[0059]

[0060] Where x is the first eigenvector, x′ is the second eigenvector, and X is the second eigenvector.

[0061] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in the embodiment of the present application, the target weight of the scaling factor is determined according to the target token length and the average token length of the training sample text, including:

[0062] If the target token length does not exceed the average token length, the attention layer of the classification model determines a value as an exponential value of the ratio of the target token length to the average token length with a preset value as the base, and determines the value as the target weight;

[0063] If the target token length exceeds the average token length, the attention layer of the classification model determines the ratio of the target token length to the logarithm of the average token length, and determines the ratio as the target weight.

[0064] In an embodiment of the present application, when the classification model determines the target weight corresponding to the scaling factor, the classification model determines the target weight based on the relationship between the target token length and the average token length.

[0065] Specifically, in an embodiment of the present application, the classification model sets the weight of the scaling factor as a piecewise function. If the target token length exceeds the average token length, the attention layer of the classification model uses a preset first function to determine the target weight. If the target token length does not exceed the average token length, the attention layer of the classification model uses a preset second function to determine the target weight.

[0066] Among them, in an embodiment of the present application, if the target token length does not exceed the average token length, the attention layer of the classification model determines the ratio of the target token length to the average token length, and determines the value of the exponent of the ratio with a preset value as the base, and determines the value as the target weight.

[0067] If the target token length exceeds the average token length, the attention layer of the classification model determines the ratio of the logarithm of the target token length to the average token length, and determines the ratio as the target weight.

[0068] Among them, the attention layer can determine the scaling factor by the following formula:

[0069]

[0070] Among them, λ is the weighted target factor, il is the target token length, tl is the average token length, is the scaling factor.

[0071] In the embodiment of the present application, by weighting the scaling factor, a parameter is added to reflect the ratio of the target token length of the target text to be recognized and the average token length of the training sample text. When the target token length is greater than the average token length, the original scaling factor is weighted to determine the target weight. When the target token length is less than the average token length, the original scaling factor is weighted to determine the target weight.

[0072] In an embodiment of the present application, the value of λ is always related to the ratio of il and tl or the ratio of the logarithm of il and tl, so that the attention (Attention) of the attention layer of the classification model when determining the target feature vector for classification becomes a local Attention, achieving the approximation of the Attention results during training and application. In an embodiment of the present application, compared with the attention algorithm of the prior art, the effect of the token length on the entropy of the attention algorithm is reduced by setting the target weight of the scaling factor. Based on this, in an embodiment of the present application, the attention layer of the classification model performs feature extraction on target texts with different target token lengths, and the extracted feature vectors are close in distribution, that is, the degree of attention dispersion is reduced.

[0073] In an embodiment of the present application, when processing long texts, the attention layer performs an Attention operation on the second eigenvector. The attention layer weights the second eigenvector using the attention mechanism, which can effectively improve the accuracy when processing long texts and reduce the disadvantage of the model focusing too much on local features.

[0074] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in an embodiment of the present application, determining the target feature vector for classification corresponding to the target text according to the scaling factor, the target weight, the second feature vector, and the preset attention algorithm includes:

[0075] The attention layer of the classification model is iterated a preset number of times, wherein each iteration includes:

[0076] Obtain a weight matrix corresponding to the iteration, process the second eigenvector according to the weight matrix, and determine an intermediate matrix; determine an intermediate eigenvector according to the scaling factor, the target weight, the intermediate matrix, and a preset attention algorithm; if the iteration is not the last iteration, use the intermediate eigenvector to update the second eigenvector; if the iteration is the last iteration, determine the intermediate eigenvector as the target eigenvector.

[0077] In the embodiment of the present application, the classification model includes multiple attention layers. When the classification model extracts feature vectors, the attention layer performs a number of prediction iterations, wherein any one iteration includes:

[0078] Obtain the weight matrix corresponding to the iteration, perform a linear transformation on the second eigenvector according to the weight matrix, and determine an intermediate matrix; determine the intermediate eigenvector according to the scaling factor, the target weight, the intermediate matrix, and the preset attention algorithm; if the iteration is not the last iteration, use the intermediate eigenvector to update the second eigenvector, and use the updated second eigenvector for the next iteration; if the iteration is the last iteration, determine the intermediate eigenvector as the target eigenvector.

[0079] In the embodiment of the present application, the weight matrix corresponding to each iteration may be the same or different.

[0080] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in an embodiment of the present application, the preset number of times is the number of network layers of the attention layer included in the classification model.

[0081] In this embodiment of the present application, when the attention layer of the classification model performs the prediction iteration number, the iteration number is consistent with the number of layers of the attention layer. In other words, in this embodiment of the present application, each attention layer performs one iteration.

[0082] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in the embodiment of the present application, the second eigenvector is processed according to the weight matrix to determine the intermediate matrix, which includes:

[0083] According to a first sub-weight matrix corresponding to the query Query included in the weight matrix, determining a first dot product of the second eigenvector and the first sub-weight matrix as a first sub-intermediate matrix;

[0084] Determine, according to a second sub-weight matrix corresponding to a key Key included in the weight matrix, a first dot product of the second eigenvector and the second sub-weight matrix as a second sub-intermediate matrix;

[0085] According to the third sub-weight matrix corresponding to the value Value included in the weight matrix, a first dot product of the second eigenvector and the third sub-weight matrix is ​​determined as a third sub-intermediate matrix.

[0086] In an embodiment of the present application, when the attention layer processes the second eigenvector according to the weight matrix to obtain the intermediate matrix, the attention layer determines the first dot product of the second eigenvector and the first sub-weight matrix as the first sub-intermediate matrix according to the first sub-weight matrix corresponding to the Query contained in the weight matrix; determines the first dot product of the second eigenvector and the second sub-weight matrix as the second sub-intermediate matrix according to the second sub-weight matrix corresponding to the Key contained in the weight matrix; and determines the first dot product of the second eigenvector and the third sub-weight matrix as the third sub-intermediate matrix according to the third sub-weight matrix corresponding to the Value contained in the weight matrix.

[0087] In this embodiment of the present application, the attention layer can determine the intermediate matrix by the following formula:

[0088]

[0089] Among them, X is the second eigenvector, W1 is the first sub-weight matrix, Q is the first sub-intermediate matrix, W2 is the second sub-weight matrix, K is the second sub-intermediate matrix, W3 is the third sub-weight matrix, and V is the third sub-intermediate matrix.

[0090] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in the embodiment of the present application, the training process of the classification model includes:

[0091] According to the sample first feature vector of the sample text, the sample second feature vector of the sample text is determined, and according to the preset sample token length and the sample average token length of the trained sample text, the sample target weight of the scaling factor is determined, according to the scaling factor, the sample target weight, the sample second feature vector and the preset attention algorithm, the sample target feature vector for classification corresponding to the sample text is determined, and according to the sample target feature vector, the sample classification result of the sample text is determined; according to the actual classification result carried in the sample text and the sample classification result, the loss value is calculated; according to the current training round and the loss value, the parameters to be adjusted corresponding to the current training round are adjusted.

[0092] In an embodiment of the present application, when training the classification model, the sample second eigenvector of the sample text is determined based on the sample first eigenvector of the sample text, and the sample target weight of the scaling factor is determined based on the preset sample token length and the sample average token length of the trained sample text. According to the scaling factor, the sample target weight, the sample second eigenvector and the preset attention algorithm, the sample target feature vector for classification corresponding to the sample text is determined, and according to the sample target feature vector, the sample classification result of the sample text is determined; the loss value is calculated based on the actual classification result carried in the sample text and the sample classification result; and according to the current training round and the loss value, the parameters to be adjusted corresponding to the current training round are adjusted.

[0093] Among them, in the embodiment of the present application, the process of determining the second eigenvector of the sample by the classification model is consistent with the process of determining the second eigenvector, and the process of determining the target eigenvector of the sample is consistent with the process of determining the target eigenvector, which will not be repeated here.

[0094] It should be noted that, in the embodiment of the present application, the electronic device trains the parameters of the hidden layer and the attention layer in the classification model in an alternating iterative manner.

[0095] Specifically, in an embodiment of the present application, the electronic device stores a correspondence between each training round and the parameter to be adjusted. After the electronic device determines the loss value, the electronic device obtains the parameter to be adjusted corresponding to the current training round and adjusts the parameter to be adjusted based on the loss value. If the number of times the loss value of the classification model is less than a preset threshold reaches a preset number threshold, or if other manually set conditions are met, the electronic device determines that the classification model training is complete.

[0096] It should be noted that the parameter to be adjusted can be the weight matrix of the attention layer or the parameter of the hidden layer.

[0097] Furthermore, in the embodiments of this application, when constructing a training set, some text may be added to the training set to extend its length through repeated concatenation. The training set must contain a sufficient amount of short conversational text, translated medium-length text, and long summary-type text, with a balance between the various data sets.

[0098] Figure 3 A schematic diagram of the existing classification model training process provided by the existing technology, such as Figure 3 As shown in the figure, when the existing classification model is trained, the sample text is labeled with a bag of words to determine the sample feature vector, and the final feature vector of the sample is obtained in the hidden layer through multiple hidden layers, attention mechanisms and activation functions; the final feature vector enters the classifier through the forward fully connected layer (FCN), and the loss value is calculated by the difference between the classification result and the labeling result, and the parameters of the hidden layer are adjusted through back propagation.

[0099] Figure 4 A schematic diagram of the training process of the classification model provided in the embodiment of the present application, such as the Figure 4 As shown, the sample second eigenvector of the sample text is determined based on the sample first eigenvector and the sample third eigenvector of the sample text, and the intermediate eigenvector of the sample text is determined based on the preset weight matrix and the preset improved attention algorithm. The sample target eigenvector for classification is then determined based on the intermediate eigenvector. The sample classification result of the sample text is determined based on the sample target eigenvector; the loss value is calculated based on the actual classification result carried in the sample text and the sample classification result; and the parameters to be adjusted corresponding to the current training round are adjusted based on the current training round and the loss value.

[0100] In order to improve the accuracy and generalization ability of the classification model, based on the above embodiments, in the embodiment of the present application, determining the target classification result of the target text according to the target feature vector includes:

[0101] The classifier of the classification model determines a target classification result of the target text according to the target feature vector.

[0102] In an embodiment of the present application, the classifier of the classification model obtains the target feature vector output by the last attention layer, and the classifier classifies the target text according to the target feature vector to determine the target classification result of the target text.

[0103] Figure 5 The flowchart of the semantic classification provided in the embodiment of the present application is as follows: Figure 5 As shown, the process includes:

[0104] S501: Obtain a target text to be classified, perform word segmentation on the target text, and determine a first feature vector corresponding to the target text based on a pre-saved correspondence between word segmentation and word vectors.

[0105] S502: The hidden layer of the classification model processes the first eigenvector to obtain a third eigenvector, and determines the dimension of the first eigenvector as the target token length.

[0106] S503: The normalization layer performs normalization processing on the first eigenvector and the third eigenvector respectively; and determines the mean vector of the normalized first eigenvector and the normalized third eigenvector as the second eigenvector.

[0107] S504: Determine a target weight of the scaling factor according to the target token length and the average token length of the training sample text.

[0108] S505: Determine a target feature vector for classification corresponding to the target text according to the scaling factor, the target weight, the second feature vector, and a preset attention algorithm.

[0109] S506: Determine the target classification result of the target text according to the target feature vector.

[0110] Based on the above embodiments, Figure 6 A schematic diagram of the structure of a semantic classification device provided in an embodiment of the present application, the device comprising:

[0111] The word segmentation module 601 is used to obtain a target text to be classified, perform word segmentation on the target text, and determine a first feature vector corresponding to the target text based on a pre-stored correspondence between word segmentation and word vectors;

[0112] Classification module 602 is used to input the first feature vector into a classification model, and the classification model determines the target token length and the second feature vector contained in the target text based on the first feature vector, and determines the target weight of the scaling factor based on the target token length and the average token length of the training sample text, and determines the target feature vector for classification corresponding to the target text based on the scaling factor, the target weight, the second feature vector and a preset attention algorithm, and determines the target classification result of the target text based on the target feature vector.

[0113] In one possible implementation, the classification module 602 is specifically used for the hidden layer of the classification model to process the first feature vector to obtain a third feature vector, and determine the dimension of the first feature vector as the target token length; the normalization layer of the classification model merges the first feature vector and the third feature vector to determine the second feature vector.

[0114] In a possible implementation, the classification module 602 is specifically configured to perform normalization processing on the first eigenvector and the third eigenvector respectively at the normalization layer; and determine the mean vector of the normalized first eigenvector and the normalized third eigenvector as the second eigenvector.

[0115] In one possible embodiment, the classification module 602 is specifically used to determine, by the attention layer of the classification model, a value of an exponential value of the ratio of the target token length to the average token length with a preset value as the base if the target token length does not exceed the average token length, and determine the value as the target weight; if the target token length exceeds the average token length, then the attention layer of the classification model determines the ratio of the target token length to the logarithm of the average token length, and determines the ratio as the target weight.

[0116] In one possible implementation, the classification module 602 is specifically used to iterate the attention layer of the classification model a preset number of times, wherein any iteration includes: obtaining a weight matrix corresponding to the iteration, processing the second eigenvector according to the weight matrix, and determining an intermediate matrix; determining an intermediate eigenvector according to the scaling factor, the target weight, the intermediate matrix, and a preset attention algorithm; if the iteration is not the last iteration, using the intermediate eigenvector to update the second eigenvector; if the iteration is the last iteration, determining the intermediate eigenvector as the target eigenvector.

[0117] In one possible implementation, the preset number is the number of network layers of the attention layer included in the classification model.

[0118] In one possible embodiment, the classification module 602 is specifically used to determine the first dot product of the second eigenvector and the first sub-weight matrix as a first sub-intermediate matrix based on the first sub-weight matrix corresponding to the query Query contained in the weight matrix; determine the first dot product of the second eigenvector and the second sub-weight matrix as a second sub-intermediate matrix based on the second sub-weight matrix corresponding to the key Key contained in the weight matrix; and determine the first dot product of the second eigenvector and the third sub-weight matrix as a third sub-intermediate matrix based on the third sub-weight matrix corresponding to the value Value contained in the weight matrix.

[0119] In a possible implementation, the device further includes:

[0120] Training module 603 is used to determine the sample second feature vector of the sample text based on the sample first feature vector of the sample text, and determine the sample target weight of the scaling factor based on the preset sample token length and the sample average token length of the trained sample text, determine the sample target feature vector for classification corresponding to the sample text based on the scaling factor, the sample target weight, the sample second feature vector and the preset attention algorithm, and determine the sample classification result of the sample text based on the sample target feature vector; calculate the loss value based on the actual classification result carried in the sample text and the sample classification result; and adjust the parameters to be adjusted corresponding to the current training round based on the current training round and the loss value.

[0121] In a possible implementation, the classification module 602 is specifically configured to use a classifier of the classification model to determine a target classification result of the target text according to the target feature vector.

[0122] Based on the above embodiments, the present application also provides an electronic device, Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, it includes: a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704;

[0123] The memory 703 stores a computer program. When the program is executed by the processor 701 , the processor 701 executes the steps of the semantic classification method provided in the above embodiments.

[0124] Since the principle of solving the problem by the above electronic device is similar to that of the semantic classification method, the implementation of the above electronic device can refer to the embodiment of the method, and the repeated parts will not be repeated.

[0125] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 702 is used for communication between the above-mentioned electronic device and other devices. The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.

[0126] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it can also be a digital signal processing processor (DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0127] Based on the above embodiments, an embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program that can be executed by a processor. When the program runs on the processor, the processor implements the steps of the semantic classification method provided in the above embodiments.

[0128] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to that of the semantic classification method, the implementation of the above-mentioned computer-readable storage medium can refer to the embodiment of the method, and the repeated parts will be omitted.

[0129] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0130] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0131] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0133] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A semantic classification method, characterized in that: The method comprises: Obtaining a target text to be classified, performing word segmentation on the target text, and determining a first feature vector corresponding to the target text based on a pre-saved correspondence between word segmentation and word vectors; Inputting the first feature vector into a classification model, the classification model determining a target token length and a second feature vector contained in the target text based on the first feature vector, and determining a target weight of a scaling factor based on the target token length and an average token length of training sample texts, determining a target feature vector for classification corresponding to the target text based on the scaling factor, the target weight, the second feature vector, and a preset attention algorithm, and determining a target classification result for the target text based on the target feature vector; The classification model determines the target token length and the second feature vector contained in the target text based on the first feature vector, including: The hidden layer of the classification model processes the first feature vector to obtain a third feature vector, and determines the dimension of the first feature vector as the target token length; The normalization layer of the classification model combines the first eigenvector and the third eigenvector to determine the second eigenvector; The step of determining the target weight of the scaling factor according to the target token length and the average token length of the training sample text includes: If the target token length does not exceed the average token length, the attention layer of the classification model determines a value as an exponential value of the ratio of the target token length to the average token length with a preset value as the base, and determines the value as the target weight; If the target token length exceeds the average token length, the attention layer of the classification model determines the ratio of the target token length to the logarithm of the average token length, and determines the ratio as the target weight.

2. The method according to claim 1, characterized in that The normalization layer of the classification model combines the first eigenvector and the third eigenvector to determine the second eigenvector, including: The normalization layer performs normalization processing on the first eigenvector and the third eigenvector respectively; and determines the mean vector of the normalized first eigenvector and the normalized third eigenvector as the second eigenvector.

3. The method according to claim 1, characterized in that Determining the target feature vector for classification corresponding to the target text according to the scaling factor, the target weight, the second feature vector, and a preset attention algorithm includes: The attention layer of the classification model is iterated a preset number of times, wherein each iteration includes: Obtain a weight matrix corresponding to the iteration, process the second eigenvector according to the weight matrix, and determine an intermediate matrix; determine an intermediate eigenvector according to the scaling factor, the target weight, the intermediate matrix, and a preset attention algorithm; if the iteration is not the last iteration, use the intermediate eigenvector to update the second eigenvector; if the iteration is the last iteration, determine the intermediate eigenvector as the target eigenvector.

4. The method according to claim 3, characterized in that The preset number is the number of network layers of the attention layer included in the classification model.

5. The method according to claim 3, characterized in that The processing of the second eigenvector according to the weight matrix to determine the intermediate matrix includes: According to a first sub-weight matrix corresponding to the query Query included in the weight matrix, determining a first dot product of the second eigenvector and the first sub-weight matrix as a first sub-intermediate matrix; Determine, according to a second sub-weight matrix corresponding to a key Key included in the weight matrix, a first dot product of the second eigenvector and the second sub-weight matrix as a second sub-intermediate matrix; According to the third sub-weight matrix corresponding to the value Value included in the weight matrix, a first dot product of the second eigenvector and the third sub-weight matrix is ​​determined as a third sub-intermediate matrix.

6. The method according to claim 5, characterized in that The training process of the classification model includes: According to the sample first feature vector of the sample text, the sample second feature vector of the sample text is determined, and according to the preset sample token length and the sample average token length of the trained sample text, the sample target weight of the scaling factor is determined, according to the scaling factor, the sample target weight, the sample second feature vector and the preset attention algorithm, the sample target feature vector for classification corresponding to the sample text is determined, and according to the sample target feature vector, the sample classification result of the sample text is determined; according to the actual classification result carried in the sample text and the sample classification result, the loss value is calculated; according to the current training round and the loss value, the parameters to be adjusted corresponding to the current training round are adjusted.

7. The method according to claim 1, characterized in that Determining a target classification result of the target text according to the target feature vector includes: The classifier of the classification model determines a target classification result of the target text according to the target feature vector.

8. An electronic device, characterized in that: The electronic device includes a processor, and the processor is used to implement the steps of the semantic classification method according to any one of claims 1 to 7 when executing a computer program stored in a memory.

Citation Information

Patent Citations

  • Answer generation method based on deep learning, electronic device and readable storage medium

    CN111241304A

  • Classification model training method, text classification method, device and equipment

    CN115391542A