A text classification method, device, equipment and computer storage medium

By performing word segmentation processing and semantic analysis model input on text, the conditional probability distribution results of hidden variables are obtained, text similarity is calculated and clustered analysis is performed, and the problem of high text matching error rate caused by ignoring word semantics in the existing technology is solved, and more accurate text classification and matching is achieved.

CN115309891BActive Publication Date: 2025-05-13LIAONING MOBILE COMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110502245.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-08
Publication Date
2025-05-13
Estimated Expiration
2041-05-08

AI Technical Summary

Technical Problem

The prior art tends to ignore the semantics between words and the synonyms or polysenses of the word itself when processing text information, resulting in a high error rate of text matching.

Method used

By performing word segmentation on the target text, a co-occurrence matrix is ​​constructed, and inputting it into a pre-constructed semantic analysis model, the conditional probability distribution results of multiple hidden variables are obtained, the text similarity is calculated, and finally text classification is used using the clustering algorithm.

Benefits of technology

By identifying the implicit semantics between words and the word itself, the error rate of text matching is reduced and the accuracy of clustering analysis is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115309891B_ABST
    Figure CN115309891B_ABST
Patent Text Reader

Abstract

The present application discloses a text classification method, device, equipment and computer storage medium. The method comprises: performing word segmentation processing on the acquired target text to obtain a co-occurrence matrix and a plurality of different word segmentations; inputting the co-occurrence matrix and the plurality of different word segmentations into a pre-built semantic analysis model to obtain a conditional probability distribution result of a plurality of hidden variables on the target text, wherein the plurality of hidden variables include a plurality of different word segmentations; calculating the text similarity between the target text and each preset text in the preset text library according to the conditional probability distribution result, and determining the similarity matrix; performing cluster analysis on the target text according to the clustering algorithm and the similarity matrix, and obtaining the text classification result. According to the text classification method of the embodiment of the present application, it is possible to accurately extract hidden synonymous or near-synonymous junk information, thereby reducing the error rate of text matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of text analysis technology, and in particular relates to a text classification method, apparatus, device and computer storage medium. Background Art

[0002] With the rapid growth of online information and the increasing maturity of technologies such as search engines, the primary challenge facing human society is no longer a lack of information, but rather how to improve the efficiency of information acquisition and access. Text clustering technology, with its flexibility and automated processing capabilities in information analysis, has become an important means of effectively organizing and navigating text information.

[0003] In recent years, developers have been committed to optimizing clustering algorithms, such as structural clustering, dispersed clustering, spectral clustering, hierarchical clustering, density clustering, balanced iterative reduction and clustering, and mean shift clustering. These clustering algorithms have been well applied to text analysis.

[0004] However, when processing text information through these traditional clustering algorithms, it is easy to ignore the semantics between words and the synonymy or polysemy of words themselves, resulting in a high text matching error rate. Summary of the Invention

[0005] The embodiments of the present application provide a text classification method, apparatus, device, and computer storage medium, which can accurately extract hidden synonymous or near-synonymous spam information, thereby reducing the error rate of text matching.

[0006] In a first aspect, an embodiment of the present application provides a text classification method, the method comprising:

[0007] Performing word segmentation processing on the acquired target text to obtain a co-occurrence matrix and a plurality of different word segments, wherein the co-occurrence matrix is ​​a matrix composed of the weights of each of the plurality of different word segments in the target text;

[0008] Inputting the co-occurrence matrix and the multiple different word segmentations into a pre-built semantic analysis model to obtain conditional probability distribution results of multiple hidden variables on the target text, wherein the conditional probability distribution results are used to indicate implicit semantic information of the target text, and the multiple hidden variables include the multiple different word segmentations;

[0009] Calculating the text similarity between the target text and each preset text in the preset text library based on the conditional probability distribution result to determine a similarity matrix;

[0010] According to the clustering algorithm and the similarity matrix, cluster analysis is performed on the target text to obtain a text classification result.

[0011] In a second aspect, an embodiment of the present application provides a text classification device, comprising:

[0012] A word segmentation module is used to perform word segmentation processing on the acquired target text to obtain a co-occurrence matrix and multiple different word segments, wherein the co-occurrence matrix is ​​a matrix composed of the weights of each of the multiple different word segments in the target text;

[0013] a semantic analysis module, configured to input the co-occurrence matrix and the plurality of different word segmentations into a pre-built semantic analysis model to obtain conditional probability distribution results of a plurality of hidden variables on the target text, wherein the conditional probability distribution results are used to indicate implicit semantic information of the target text, and the plurality of hidden variables include the plurality of different word segmentations;

[0014] a matrix determination module, configured to calculate the text similarity between the target text and each preset text in the preset text library based on the conditional probability distribution result, and determine a similarity matrix;

[0015] The classification module is used to perform cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain a text classification result.

[0016] In a third aspect, an embodiment of the present application provides a text classification device, comprising:

[0017] a processor and a memory storing computer program instructions;

[0018] When the processor executes the computer program instructions, the text classification method as described in any one of the above embodiments is implemented.

[0019] In a fourth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the text classification method described in any one of the above embodiments is implemented.

[0020] The text classification method, device, equipment and computer storage medium of the embodiment of the present application are as follows: after the acquired target text is segmented, the co-occurrence matrix and multiple different segmentations are input into a pre-trained semantic analysis model to obtain a conditional probability distribution result of the semantic relationship between multiple hidden variables and the target text, the conditional probability distribution result is the probability calculated taking into account semantic problems such as polysemy of the segmentation word, and the calculation result is more accurate; then, based on the conditional probability distribution result, the similarity between the target text and each preset text in the preset text library is calculated, the similarity matrix is ​​determined, and finally, the target text is clustered according to the clustering algorithm and the similarity matrix. In this way, through the text classification method of the present application, according to the pre-built semantic analysis model, the hidden meanings between words and words and the words themselves can be more comprehensively identified, the potential semantic information of the segmentation words in the target text can be determined, and then the similarity between texts can be calculated based on the conditional probability distribution result of the hidden vector in the target text, so that the clustering result of clustering using the clustering algorithm is more accurate, thereby reducing the error rate of text matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 This is a flowchart of a text classification method provided by an embodiment of the present application;

[0023] Figure 2 This is a structural diagram of a text classification device provided by an embodiment of the present application;

[0024] Figure 3 It is a structural diagram of a text classification device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0026] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0027] In order to solve the problems of the existing technology, the embodiments of the present application provide a text classification method, device, equipment and computer storage medium, which can accurately identify the implicit semantic information between words in the text or between words themselves through a semantic analysis model, so that the results of clustering analysis of the text through a clustering algorithm are more accurate.

[0028] It should be noted that the text classification method provided in the embodiments of this application requires the use of a pre-built semantic analysis model to analyze the implicit semantic information of each word in the text. Therefore, before using the semantic analysis model for semantic analysis, it is necessary to first build a semantic analysis model. The following first describes the specific implementation method of the semantic analysis model construction method provided in the embodiments of this application.

[0029] The present invention provides a method for constructing a semantic analysis model, which can be implemented by the following steps:

[0030] 1. Obtain a sample set, which includes a sample text set, a sample word set, a sample co-occurrence matrix and a sample hidden variable set. The sample text set contains multiple sample texts, the sample word set contains multiple sample word segments, the sample co-occurrence matrix is ​​a matrix composed of the weights of the multiple sample word segments in the multiple sample texts, and the sample hidden variable set contains multiple sample hidden variables.

[0031] In the embodiment of the present application, the above sample set can be obtained through a local database of a computer.

[0032] In an example, the above sample text set can be represented as D. Specifically, D={d1,d2,…,d m}, where m represents the number of sample texts;

[0033] The above sample word set can be expressed as S. Specifically, S={s1,s2,…,s n}, where n represents the number of sample segmentations;

[0034] The above sample co-occurrence matrix can be expressed as A. Specifically, A=|a ij | n×m , where a ij Represents sample segmentation s j In the sample text d i The weight value in ;

[0035] The above sample hidden variable set can be expressed as Z. Specifically, Z={z1,z2,…,z r}, where r represents the number of sample hidden variables.

[0036] 2. Based on the multiple sample texts, the multiple sample word segmentations, the sample co-occurrence matrix and the multiple sample hidden variables, calculate a first probability distribution of the multiple sample word segmentations on the multiple sample texts, a second probability distribution of the multiple sample hidden variables on the multiple sample texts and a third probability distribution of the multiple sample word segmentations on the multiple sample hidden variables.

[0037] In an embodiment of the present application, an initial semantic analysis model can be preset, and the above-mentioned multiple sample texts, multiple sample word segmentations, sample co-occurrence matrices and multiple sample hidden variables can be input into the initial semantic analysis model to calculate the probability relationship between the above-mentioned multiple sample texts, multiple sample word segmentations, sample co-occurrence matrices and multiple sample hidden variables.

[0038] In one example, the sample co-occurrence matrix can be expressed as a probability distribution. Specifically, the sample co-occurrence matrix can be expressed as P(d i ,s j ), can be expressed in the sample text d i In the sample segmentation s j The probability of existence.

[0039] The first probability distribution can be expressed as P(s j |d i ), which can be expressed as i Under the condition of j Probability of existence;

[0040] The second probability distribution can be expressed as P(z r |d i ), which can be expressed as i Under the condition of r Probability of existence;

[0041] The third probability distribution can be expressed as P(s j |z r ), which can be expressed as zr Under the condition of j Probability of existence;

[0042] In one example, the first probability distribution, the second probability distribution, and the third probability distribution can be obtained by the following formula 1 and formula 2:

[0043] P(d i ,s j )=P(d i )P(s j |d i ) Formula 1

[0044] P(s j |d i )=∑ seZ P(s j |z)P(z|d i ) Formula 2

[0045] In addition, because the above initial semantic analysis model is a hybrid model, under given conditions, the above sample hidden variable z r and sample segmentation s j are all multinomial distributions, so in order to approximate the distribution of sample text and sample word segmentation to the greatest extent, it is necessary to find P(d i ,s j ) is the maximum value of the likelihood function. Specifically, P(d i ,s j The maximum value of the likelihood function of ) can be obtained by the following formula 3:

[0046]

[0047] Among them, g(d i ,s j ) means that the sample text is d i In the sample segmentation s j Number of occurrences.

[0048] 3. For the first probability distribution, the second probability distribution and the third probability distribution, a maximum expectation algorithm is used to iteratively calculate the second probability distribution, the third probability distribution and the fourth probability distribution until a preset iteration stop condition is met, thereby obtaining a constructed semantic analysis model, wherein the fourth probability distribution is used to indicate the probability distribution of the multiple sample hidden variables on the multiple sample texts and the multiple sample word segments, and the fourth probability distribution is calculated based on the first probability distribution, the second probability distribution and the third probability distribution.

[0049] In the embodiment of the present application, P(d i ,sj ) is not the optimal solution. Therefore, in probabilistic latent semantic analysis, it is also necessary to use the expectation-maximization algorithm (EM algorithm) to repeatedly estimate the parameters in the above initial semantic analysis model, that is, repeatedly calculate the above second probability distribution and third probability distribution until the iteration stopping condition is reached.

[0050] In one example, the fourth probability distribution can be expressed as P(z r |d i ,s j ), which can be expressed as i , the sample segmentation is s j Under the condition of r The probability of existence can also be understood as the sample hidden variable z r The posterior probability of can be calculated by the following formula 4.

[0051] In one example, the EM algorithm is divided into the following two steps:

[0052] 3.1, E-step, using the current parameter P(s j |d i )、P(z r |d i ) and P(s j |z r ), the sample hidden variable z can be calculated by the following formula 4 r The posterior probability of :

[0053]

[0054] 3.2, M-step, based on the above posterior probability, the following formula 5 and formula 6 can be used to recalculate P(s) in the initial semantic analysis model j |z r ) and P(z r |d i ) is estimated to be:

[0055]

[0056]

[0057] Based on the above EM algorithm, P(z r |d i ,s j ), P(s j |z r ) and P(z r |d i ) until P(zr |d i ) converges, or the expected value of the likelihood function increases less than the preset threshold, the iteration is stopped, and P(z r |d i ) to obtain the optimal solution of the constructed semantic analysis model.

[0058] It should be noted that the above-mentioned iteration stopping condition can be set according to specific circumstances and is not limited here.

[0059] The above is a specific implementation of the method for constructing a semantic analysis model provided in the embodiments of this application. The semantic analysis model constructed above can be applied to the text classification method provided in the following embodiments.

[0060] Based on this, the semantic analysis model constructed by the above method can finally obtain the optimal solution of the conditional probability distribution of hidden variables in the text. The conditional probability distribution (z r |d i ) can represent the implicit semantic information and semantic structure in the text, and can solve semantic problems such as the inability to identify hidden semantic information between words or the polysemy of a word.

[0061] The following is combined with Figure 1 The specific implementation of the text classification method provided in this application is described in detail.

[0062] Figure 1 FIG. 1 shows a flow chart of a text classification method provided by an embodiment of the present application. Figure 1 As shown, the following steps are included:

[0063] Step 101: performing word segmentation processing on the acquired target text to obtain a co-occurrence matrix and a plurality of different word segments, wherein the co-occurrence matrix is ​​a matrix composed of the weights of each of the plurality of different word segments in the target text;

[0064] Step 102: Input the co-occurrence matrix and the multiple different word segmentations into a pre-built semantic analysis model to obtain conditional probability distribution results of multiple hidden variables on the target text, wherein the conditional probability distribution results are used to indicate implicit semantic information of the target text, and the multiple hidden variables include the multiple different word segmentations;

[0065] Step 103, calculating the text similarity between the target text and each preset text in the preset text library based on the conditional probability distribution result, and determining a similarity matrix;

[0066] Step 104: Perform cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain a text classification result.

[0067] Based on this, after the acquired target text is segmented, the co-occurrence matrix and multiple different segmentations are input into a pre-trained semantic analysis model to obtain the conditional probability distribution results of the semantic relationship between multiple hidden variables and the target text. The conditional probability distribution results are the probabilities calculated taking into account semantic issues such as polysemy of the segmentation word, and the accuracy of the calculation results is higher; then, based on the conditional probability distribution results, the similarity between the target text and each preset text in the preset text library is calculated, and the similarity matrix is ​​determined. Finally, the target text is clustered according to the clustering algorithm and the similarity matrix. In this way, through the text classification method of the present application, according to the pre-built semantic analysis model, it is possible to more comprehensively identify the hidden meanings between words and the words themselves, determine the potential semantic information of the segmentation words in the target text, and then calculate the similarity between texts based on the conditional probability distribution results of the hidden vectors in the target text, so that the clustering results of clustering using the clustering algorithm are more accurate, thereby reducing the error rate of text matching.

[0068] In the above step 101, first, a target text is obtained and word segmentation is performed on the target text, and finally a co-occurrence matrix and a plurality of different word segments are obtained.

[0069] In target texts with implicit semantics, a large number of special characters or meaningless special words generally interfere with similarity judgment. Therefore, in an embodiment of the present application, the target text can be segmented to remove these interfering special characters and extract feature items from multiple segmentations.

[0070] Specifically, the word segmentation processing of the acquired target text can be completed by the following steps:

[0071] Segmenting the target text according to a preset segmentation rule to obtain multiple segmentations;

[0072] Dimensionality reduction and weighting processing are performed on the multiple segmentations to obtain a co-occurrence matrix and multiple different segmentations.

[0073] Based on this, the target text is segmented and the obtained segmentations are processed through dimensionality reduction and weighting to extract multiple different segmentations. Since interfering characters are eliminated and features are extracted for the segmentations, the text is clustered based on the extracted multiple different segmentations, and the clustering results are more accurate.

[0074] In an embodiment of the present application, the above-mentioned preset word segmentation rules can be a semantic classification algorithm in a preset corpus. In one example, according to semantic classification, each of the above-mentioned multiple word segmentations can be a character, a word, or a sentence; or, the above-mentioned preset word segmentation rules can also be word segmentation rules set according to specific application scenarios. In one example, it is stipulated as needed that the number of words in a word segmentation cannot exceed 5 words.

[0075] For example, if the target text contains long and short sentences, they will be segmented according to the preset classification rules. For example, "The weather is very good today" can be segmented into three words: "today", "weather", and "very good"; for another example, if the target text contains useless words, they will be removed according to the preset classification rules. For example, in "Have you seen the photo album Wuhu? See the URL: xxx", the stop word "Wuhu" needs to be removed.

[0076] In addition, after the target text is segmented, the multiple segmented words need to be reduced in dimension and weighted. In one example, the dimensionality reduction and weighting of the multiple segmented words can be processed using an information-incremental feature extraction method. Specifically, the information gain calculation formula can be used, that is, Formula 7:

[0077]

[0078] Among them, T represents the word segmentation feature item, c i is the text category, P(c i ) means c i The probability of a class text appearing in the preset text library, P(T) is the probability of a text containing the word segmentation feature item T, P(c i |T) indicates that the text containing the word segmentation feature item T belongs to c i The probability of the class.

[0079] It should be noted that the above-mentioned dimensionality reduction and weighting processing of the multiple word segmentations can be understood as a process of feature extraction of multiple word segmentations. Therefore, in addition to the above-mentioned information-incremental feature extraction method, other feature extraction methods commonly used in this field can also be selected, which are not limited here.

[0080] In the above step 102, the co-occurrence matrix obtained in the above step 101 and a plurality of different word segmentations are input into a pre-built semantic analysis model, and conditional probability distribution results of a plurality of hidden variables on the target text can be obtained.

[0081] In the embodiment of the present application, when constructing a semantic analysis model, a hidden variable set is included in the semantic analysis model, that is, the hidden variable set can be understood as the above-mentioned sample hidden variable set Z. Moreover, the hidden variable set Z can include the above-mentioned multiple different word segmentations.

[0082] In addition, since the above conditional probability distribution result is the conditional probability distribution result of multiple hidden variables in the hidden variable set Z on the target text, the conditional probability distribution result can be used to represent the implicit semantic information of the target text.

[0083] In one example, the target text can be represented as d0. After the text d0 is segmented in step 101, a set S0 of multiple different segmented words and a co-occurrence matrix A0 are obtained. S0 and A0 are input into the pre-built semantic analysis model to obtain the conditional probability distribution P(z) of r latent variables in the hidden variable set Z in the target text d0. r |d0).

[0084] In the above step 103, the similarity between the target text and each preset text in the preset text library can be calculated based on the conditional probability distribution result obtained in the above step 102 to determine a similarity matrix.

[0085] It should be noted that, in the embodiment of the present application, the above-mentioned preset text library can be understood as the sample text set D obtained in the above-mentioned model construction process, and the similarity matrix is ​​determined by calculating the similarity between the target text and each sample text in the sample text set D.

[0086] Specifically, the above-mentioned calculation of the text similarity between the target text and each preset text in the preset text library based on the conditional probability distribution result and determination of the similarity matrix may include the following steps:

[0087] Obtaining a probability value corresponding to each hidden variable in the plurality of hidden variables from the conditional probability distribution result;

[0088] Constructing a hidden variable vector of the target text according to the probability value, wherein the hidden variable vector is used to indicate the target text vector, and the target text vector is used to indicate implicit semantic information of the target text;

[0089] Calculating the text similarity between the target text and each preset text in the preset text library based on the target text vector to obtain multiple text similarities;

[0090] A similarity matrix is ​​determined according to the multiple text similarities.

[0091] Based on this, the target text vector is constructed through the probability value corresponding to each hidden variable, and the text similarity between the target text vector and each preset text in the preset text library is calculated, so that the obtained similarity matrix is ​​more accurate and the implicit semantics in the clustered target text is more precise.

[0092] In one example, the probability value corresponding to each of the above hidden variables can be expressed as p 0,r , which means hidden variable z r The probability value in d0 of the target text has r hidden variables, so r probability values ​​can be obtained. Therefore, the hidden variable vector corresponding to the target text can be expressed as dz0=(p 0,1 ,p 0,2 ,…,p 0,r ), the hidden variable vector dz0 is the target text vector.

[0093] After determining the target text vector, calculating the text similarity between the target text and each preset text in the preset text library based on the target text vector to obtain multiple text similarities may include the following steps:

[0094] Obtaining a preset text vector for each preset text in the preset text library, wherein the preset text vector is used to indicate a semantic relationship between the plurality of hidden variables and the preset text;

[0095] The similarity between the target text vector and the preset text vector of each preset text is calculated using the angle cosine similarity calculation formula to obtain multiple vector similarities, wherein the vector similarities are used to indicate the text similarity.

[0096] Based on this, the similarity between the target text vector and the text vector in the preset text library is calculated using the angle cosine formula. Compared with the existing similarity calculation method, this method eliminates the step of manually selecting parameters, eliminates the influence of uncertain factors, and improves the stability and accuracy of clustering.

[0097] In one example, the preset text vector of each preset text in the above text library can be expressed as dz i =(p i,1 ,p i,2 ,…,p i,r ), which means the preset text d i The corresponding preset text vector.

[0098] The target text vector dz0 and each preset text vector dz i Substitute the angle cosine similarity calculation formula, that is, the following formula 8, to calculate dz0 and dz i The vector similarity between the target text d0 and the preset text d i The text similarity.

[0099]

[0100] Since the above-mentioned sample text set D has a total of m preset texts, m text similarities can be obtained, and the similarity matrix W can be determined based on the m text similarities.

[0101] In the above step 104, based on the similarity matrix determined in the above step 103, a clustering algorithm is used to perform cluster analysis on the target text to obtain a text classification result.

[0102] In the embodiment of the present application, the above-mentioned clustering algorithm can be a K-means clustering algorithm, or a hierarchical clustering algorithm, etc. Since the concept of similarity matrix is ​​applied, the above-mentioned clustering algorithm can be any algorithm that applies the concept of similarity matrix, and is not limited here.

[0103] Specifically, performing cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain a text classification result may include the following steps:

[0104] Performing Laplace transform on the similarity matrix to obtain a Laplace matrix;

[0105] Based on a preset feature vector extraction rule, obtaining a target feature vector from the Laplacian matrix;

[0106] The target text is clustered according to a clustering algorithm and the target feature vector to obtain a clustering result, where the clustering result is used to indicate a plurality of implicit semantic categories of the target text.

[0107] Based on this, we perform a Laplace transform on the similarity matrix to determine a normalized Laplace matrix. This Laplace matrix is ​​then processed to extract the target feature vector. The target text is then clustered based on the clustering algorithm and the target feature vector. By performing a Laplace transform on the similarity matrix to extract the target feature vector, the accuracy of the clustering results can be improved, thereby improving the accuracy of text matching.

[0108] In the embodiment of the present application, the Laplace transform of the similarity matrix is ​​performed to obtain a Laplace matrix, which can be calculated using a normalized Laplace formula. In one example, the normalized Laplace matrix can be obtained using the following formula 9:

[0109]

[0110] Where D represents the diagonal matrix of the similarity matrix W.

[0111] Secondly, in one example, the above-mentioned preset feature vector extraction rule can be an extraction rule formulated according to a specific application scenario. Specifically, after the normalized Laplace matrix L is calculated by formula 9, eqzAfter that, L can be calculated by the eigenvalue calculation formula eqz The eigenvalue of L eqz There are multiple eigenvalues, and the eigenvectors corresponding to the first K eigenvalues ​​with the largest eigenvalues ​​can be selected as the target eigenvector. The target eigenvector can be expressed as {x1, x2, ..., x k}∈R k×n , where x k represents any target feature vector.

[0112] In one example, the target text is clustered according to the clustering algorithm and the target feature vector to obtain a clustering result. Specifically, a feature matrix Y can be constructed based on the target feature vector. Y can be calculated using the following formula 10:

[0113] Y=[x1,x2,…,x k ] T =[y1,y2,…,y n ] Formula 10

[0114] Among them, each row in the above matrix Y can be regarded as a K-dimensional space vector, that is, n vectors can be obtained. Then, the feature matrix Y is clustered and analyzed by the K-means clustering algorithm, and finally the clustering result is obtained. The clustering result can be understood as dividing the target text into K classes according to the implicit semantics.

[0115] In addition, after obtaining the above text classification results, the text classification results can also be verified and the clustering accuracy can be calculated. The above clustering accuracy can be used to measure the clustering effect. In one example, the clustering accuracy can be calculated using the clustering accuracy formula, that is, the following formula 11:

[0116]

[0117] Among them, δ is the scale parameter.

[0118] For example, the existing algorithm requires manual selection of the scale parameter δ. First, the optimal value of the scale parameter needs to be determined. Experiments have shown that the clustering accuracy of the existing algorithm is highest when the scale parameter δ is 20. Therefore, the existing algorithm uses this as the standard for clustering analysis; however, in this application, latent factors with higher clustering accuracy are also selected for analysis, that is, the number of latent factors is 35. Here, the latent factor can be understood as the above-mentioned each row in the matrix Y is regarded as a K-dimensional space vector, that is, n vectors can be obtained, and n takes the value of 35.

[0119] Based on the selected scale parameters and the potential factors in the semantic analysis, cluster comparative analysis was performed with dimensions of 600, 800, and 1000. The results are shown in the following table:

[0120] Dimensions Average clustering accuracy of existing algorithms The accuracy of the proposed new clustering method 600 0.6506 0.7098 800 0.6893 0.7212 1000 0.7059 0.7387

[0121] Experimental results show that the clustering accuracy of the text classification method of this application is higher than that of existing algorithms.

[0122] Figure 2 FIG. 1 shows a schematic diagram of the structure of the text classification device provided in an embodiment of the present application. Figure 2 As shown, the potential user terminal determination device 200 includes:

[0123] The word segmentation module 201 is used to perform word segmentation processing on the acquired target text to obtain a co-occurrence matrix and multiple different word segments, wherein the co-occurrence matrix is ​​a matrix composed of the weights of each of the multiple different word segments in the target text;

[0124] A semantic analysis module 202 is configured to input the co-occurrence matrix and the multiple different word segmentations into a pre-built semantic analysis model to obtain conditional probability distribution results of multiple hidden variables on the target text, wherein the conditional probability distribution results are used to indicate implicit semantic information of the target text, and the multiple hidden variables include the multiple different word segmentations;

[0125] A matrix determination module 203 is configured to calculate the text similarity between the target text and each preset text in the preset text library based on the conditional probability distribution result, and determine a similarity matrix;

[0126] The classification module 204 is configured to perform cluster analysis on the target text according to a clustering algorithm and the similarity matrix to obtain a text classification result.

[0127] Optionally, the matrix determination module 203 specifically includes:

[0128] A probability value extraction unit, configured to obtain a probability value corresponding to each of the multiple hidden variables from the conditional probability distribution result;

[0129] a vector construction unit, configured to construct a hidden variable vector of the target text according to the probability value, wherein the hidden variable vector is used to indicate the target text vector, and the target text vector is used to indicate implicit semantic information of the target text;

[0130] A similarity calculation unit is used to calculate the text similarity between the target text and each preset text in the preset text library based on the target text vector to obtain multiple text similarities;

[0131] The similarity matrix determining unit is configured to determine a similarity matrix according to the plurality of text similarities.

[0132] Optionally, the similarity calculation unit is specifically configured to:

[0133] Obtaining a preset text vector for each preset text in the preset text library, wherein the preset text vector is used to indicate a semantic relationship between the plurality of hidden variables and the preset text;

[0134] The similarity between the target text vector and the preset text vector of each preset text is calculated using the angle cosine similarity calculation formula to obtain multiple vector similarities, wherein the vector similarities are used to indicate the text similarity.

[0135] Optionally, the classification module 204 is specifically configured to:

[0136] Performing Laplace transform on the similarity matrix to obtain a Laplace matrix;

[0137] Based on a preset feature vector extraction rule, obtaining a target feature vector from the Laplacian matrix;

[0138] The target text is clustered according to a clustering algorithm and the target feature vector to obtain a clustering result, where the clustering result is used to indicate a plurality of implicit semantic categories of the target text.

[0139] Optionally, the apparatus 200 further includes:

[0140] A sample set acquisition module is used to acquire a sample set, wherein the sample set includes a sample text set, a sample word set, a sample co-occurrence matrix, and a sample hidden variable set. The sample text set includes multiple sample texts, the sample word set includes multiple sample word sets, the sample co-occurrence matrix is ​​a matrix composed of weights of the multiple sample word sets in the multiple sample texts, and the sample hidden variable set includes multiple sample hidden variables.

[0141] A first calculation module is configured to calculate, based on the multiple sample texts, the multiple sample segmentations, the sample co-occurrence matrix, and the multiple sample hidden variables, a first probability distribution of the multiple sample segmentations over the multiple sample texts, a second probability distribution of the multiple sample hidden variables over the multiple sample texts, and a third probability distribution of the multiple sample segmentations over the multiple sample hidden variables;

[0142] A second calculation module is used to iteratively calculate the second probability distribution, the third probability distribution, and the fourth probability distribution using a maximum expectation algorithm for the first probability distribution, the second probability distribution, and the third probability distribution until a preset iteration stop condition is met, thereby obtaining a constructed semantic analysis model, wherein the fourth probability distribution is used to indicate the probability distribution of the multiple sample hidden variables on the multiple sample texts and the multiple sample word segmentations, and the fourth probability distribution is calculated based on the first probability distribution, the second probability distribution, and the third probability distribution.

[0143] Optionally, the word segmentation module 201 is specifically configured to:

[0144] Segmenting the target text according to a preset segmentation rule to obtain multiple segmentations;

[0145] Dimensionality reduction and weighting processing are performed on the multiple segmentations to obtain a co-occurrence matrix and multiple different segmentations.

[0146] Optionally, the apparatus 200 further includes:

[0147] The verification module is used to verify the text classification results and calculate the clustering accuracy.

[0148] Based on this, after the acquired target text is segmented, the co-occurrence matrix and multiple different segmentations are input into a pre-trained semantic analysis model to obtain the conditional probability distribution results of the semantic relationship between multiple hidden variables and the target text. The conditional probability distribution results are the probabilities calculated taking into account semantic issues such as polysemy of the segmentation word, and the accuracy of the calculation results is higher; then, based on the conditional probability distribution results, the similarity between the target text and each preset text in the preset text library is calculated, and the similarity matrix is ​​determined. Finally, the target text is clustered according to the clustering algorithm and the similarity matrix. In this way, through the text classification method of the present application, according to the pre-built semantic analysis model, it is possible to more comprehensively identify the hidden meanings between words and the words themselves, determine the potential semantic information of the segmentation words in the target text, and then calculate the similarity between texts based on the conditional probability distribution results of the hidden vectors in the target text, so that the clustering results of clustering using the clustering algorithm are more accurate, thereby reducing the error rate of text matching.

[0149] The text classification device provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0150] Figure 3 A schematic diagram of the hardware structure of a text classification device provided in an embodiment of the present application is shown.

[0151] The text classification device may include a processor 301 and a memory 302 storing computer program instructions.

[0152] Specifically, the processor 301 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0153] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 702 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 302 may include removable or non-removable (or fixed) media. Where appropriate, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 702 is a non-volatile solid-state memory.

[0154] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.

[0155] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any one of the text classification methods in the above embodiments.

[0156] In one example, the text classification device may further include a communication interface 303 and a bus 310. Figure 3 As shown, the processor 301 , the memory 302 , and the communication interface 303 are connected via a bus 310 and communicate with each other.

[0157] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0158] Bus 310 includes hardware, software or both, and couples the components of text classification device to each other. For example, and not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnect (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. Where appropriate, bus 310 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.

[0159] The text classification device can execute the text classification method in the embodiment of the present application based on the conditional probability distribution result output by the semantic analysis model, thereby realizing the combination of Figure 1 Described text classification method and device.

[0160] In addition, in conjunction with the text classification method in the above embodiments, the present application embodiment may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the text classification methods in the above embodiments is implemented.

[0161] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0162] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0163] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.

[0164] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A text classification method, characterized in that: The method comprises: Performing word segmentation processing on the acquired target text to obtain a co-occurrence matrix and a plurality of different word segments, wherein the co-occurrence matrix is ​​a matrix formed by the weight of each word in the plurality of different word segments in the target text; Inputting the co-occurrence matrix and the multiple different word segmentations into a pre-built semantic analysis model to obtain conditional probability distribution results of multiple hidden variables on the target text, wherein the conditional probability distribution results are used to indicate implicit semantic information of the target text, and the multiple hidden variables include the multiple different word segmentations; According to the conditional probability distribution result, calculating the text similarity between the target text and each preset text in the preset text library, and determining a similarity matrix; Performing cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain a text classification result; The step of calculating the text similarity between the target text and each preset text in the preset text library according to the conditional probability distribution result and determining a similarity matrix specifically includes: Obtaining a probability value corresponding to each hidden variable in the multiple hidden variables from the conditional probability distribution result; According to the probability value, construct a hidden variable vector of the target text, wherein the hidden variable vector is used to indicate the target text vector, and the target text vector is used to indicate implicit semantic information of the target text; Calculating the text similarity between the target text and each preset text in the preset text library according to the target text vector to obtain multiple text similarities; Determining a similarity matrix according to the multiple text similarities; The step of calculating the text similarity between the target text and each preset text in the preset text library according to the target text vector to obtain multiple text similarities specifically includes: Acquire a preset text vector for each preset text in the preset text library, wherein the preset text vector is used to indicate a semantic relationship between the plurality of hidden variables and the preset text; The similarity between the target text vector and the preset text vector of each preset text is calculated using the angle cosine similarity calculation formula to obtain multiple vector similarities, wherein the vector similarities are used to indicate the text similarities.

2. The text classification method according to claim 1, characterized in that: The step of performing cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain a text classification result specifically includes: Performing Laplace transformation on the similarity matrix to obtain a Laplace matrix; Based on a preset feature vector extraction rule, obtaining a target feature vector from the Laplacian matrix; The target text is clustered according to a clustering algorithm and the target feature vector to obtain a clustering result, wherein the clustering result is used to indicate a plurality of implicit semantic categories of the target text.

3. The text classification method according to claim 1, characterized in that: Before performing word segmentation processing on the acquired target text to obtain a co-occurrence matrix and a plurality of different word segments, the method further includes: Acquire a sample set, wherein the sample set includes a sample text set, a sample word set, a sample co-occurrence matrix and a sample hidden variable set, wherein the sample text set includes a plurality of sample texts, the sample word set includes a plurality of sample word sets, the sample co-occurrence matrix is ​​a matrix formed by weights of the plurality of sample word sets in the plurality of sample texts, and the sample hidden variable set includes a plurality of sample hidden variables; According to the multiple sample texts, the multiple sample segmentations, the sample co-occurrence matrix and the multiple sample hidden variables, calculate a first probability distribution of the multiple sample segmentations on the multiple sample texts, a second probability distribution of the multiple sample hidden variables on the multiple sample texts and a third probability distribution of the multiple sample segmentations on the multiple sample hidden variables; For the first probability distribution, the second probability distribution and the third probability distribution, the maximum expectation algorithm is used to iteratively calculate the second probability distribution, the third probability distribution and the fourth probability distribution until a preset iteration stop condition is met, thereby obtaining a constructed semantic analysis model, wherein the fourth probability distribution is used to indicate the probability distribution of the multiple sample hidden variables on the multiple sample texts and the multiple sample word segmentations, and the fourth probability distribution is calculated based on the first probability distribution, the second probability distribution and the third probability distribution.

4. The text classification method according to claim 1, characterized in that: The word segmentation processing of the acquired target text to obtain a co-occurrence matrix and a plurality of different word segments specifically includes: According to a preset word segmentation rule, the target text is segmented to obtain a plurality of word segments; The multiple word segments are subjected to dimensionality reduction and weighting processing to obtain a co-occurrence matrix and multiple different word segments.

5. The text classification method according to claim 1, characterized in that: After performing cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain the text classification result, the method further includes: The text classification result is verified and the clustering accuracy is calculated.

6. A text classification device, characterized in that: The device comprises: A word segmentation module is used to perform word segmentation processing on the acquired target text to obtain a co-occurrence matrix and a plurality of different word segments, wherein the co-occurrence matrix is ​​a matrix composed of the weights of each word in the plurality of different word segments in the target text; A semantic analysis module, used for inputting the co-occurrence matrix and the multiple different word segmentations into a pre-built semantic analysis model to obtain conditional probability distribution results of multiple hidden variables on the target text, wherein the conditional probability distribution results are used to indicate implicit semantic information of the target text, and the multiple hidden variables include the multiple different word segmentations; A matrix determination module, used to calculate the text similarity between the target text and each preset text in the preset text library according to the conditional probability distribution result, and determine a similarity matrix; A classification module, used to perform cluster analysis on the target text according to the clustering algorithm and the similarity matrix to obtain a text classification result; The step of calculating the text similarity between the target text and each preset text in the preset text library according to the conditional probability distribution result and determining a similarity matrix specifically includes: Obtaining a probability value corresponding to each hidden variable in the multiple hidden variables from the conditional probability distribution result; According to the probability value, construct a hidden variable vector of the target text, wherein the hidden variable vector is used to indicate the target text vector, and the target text vector is used to indicate implicit semantic information of the target text; Calculating the text similarity between the target text and each preset text in the preset text library according to the target text vector to obtain multiple text similarities; Determining a similarity matrix according to the multiple text similarities; The step of calculating the text similarity between the target text and each preset text in the preset text library according to the target text vector to obtain multiple text similarities specifically includes: Acquire a preset text vector for each preset text in the preset text library, wherein the preset text vector is used to indicate a semantic relationship between the plurality of hidden variables and the preset text; The similarity between the target text vector and the preset text vector of each preset text is calculated using the angle cosine similarity calculation formula to obtain multiple vector similarities, wherein the vector similarities are used to indicate the text similarities.

7. A text classification device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the text classification method according to any one of claims 1 to 5 is implemented.

8. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the text classification method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Emotion classification method and device, electronic equipment and storage medium

    CN112732915A

  • Determining failure modes of devices based on text analysis

    EP3477487A1