Text clustering method and related device
The text clustering method employs tpBERT for dimensionality reduction and DBSCAN clustering, coupled with a modified PGNet model, to efficiently generate summaries from multiple documents, overcoming the limitations of traditional methods by preserving semantic information and improving cluster coherence.
Patent Information
- Application Number
- CN202111396748.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-23
AI Technical Summary
Traditional text clustering methods cannot effectively capture semantic information when processing long texts, and the cluster center offset during clustering results in large differences in text similarity in the same cluster, making it impossible to generate common abstracts for multiple chapters of documents.
The tpBERT model is used for semantic dimensionality reduction processing, the DBSCAN algorithm is used to perform text clustering without specified cluster numbers, and the MD_PGNet model is generated to generate a summary of cluster clusters, which is improved to process multi-chapter documents.
It realizes effective clustering and common summary generation of multi-section documents, solves the defect of only processing single document summary generation in traditional methods, and improves the flexibility and automated processing capabilities of text clustering.
Smart Images

Figure CN114328910B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of natural language processing applications, and specifically relates to a text clustering method and related devices. Background Art
[0002] With the continuous advancement of the informatization process, the ever-increasing text information brings great troubles to people, and information processing technology can effectively help people mine effective information from massive data. Text category classification is the most basic method of information processing technology. Among them, text category classification mainly includes supervised text classification methods and unsupervised or semi-supervised text clustering methods at present. The supervised text classification method constructs proprietary or domain-oriented texts by organizing artificial annotations of text data through predefined category labels. Once the predefined category labels are determined, it is very difficult to change them. Therefore, the supervised text classification method greatly limits the expansion of text category classification. Based on semi-supervised or unsupervised text clustering methods, the problem of text category classification can be solved, and they have been widely used in text category classification. The text clustering function divides documents with high similarity into the same category by performing clustering analysis on a large number of input texts. The similarity of documents in the same category is relatively large, while the similarity of documents in different categories is relatively small. As an unsupervised machine learning method, clustering has a certain degree of flexibility and high automated processing ability because it does not require a training process and does not require manual annotation of document categories in advance. For example: traditional news clustering methods based on single_pass, and topic text clustering methods based on LDA.
[0003] However, these methods also have some deficiencies. There are mainly several problems with current text vector features: the vector expression based on the bag-of-words model has poor feature expression effect for long texts, and the vector expressions of TF-IDF based on word frequency statistics and LDA based on topics both ignore the context semantic association information of words; in addition, during the clustering process, as the text data increases and the cluster centers shift, the similarity difference between texts within the same cluster will be relatively large. Therefore, the defect that only the summary generation of a single document can be achieved in the traditional solution cannot be solved.
[0004] Therefore, there is an urgent need for a new text clustering method to solve the above problems. Summary of the Invention
[0005] This application provides a text clustering method and related devices to solve the defect that only the summary generation of a single document can be achieved in the traditional solution.
[0006] To solve the above technical problems, a technical solution adopted in this application is: to provide a text clustering method, including: obtaining a plurality of documents; in response to the existence of a to-be-processed document with a character length exceeding a threshold among the plurality of documents, performing dimensionality reduction processing on the to-be-processed document so that the character length of the to-be-processed document is less than or equal to the threshold; clustering all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster; and generating a corresponding summary for each clustering cluster.
[0007] Among them, the to-be-processed document includes a body and a title. The step of performing dimensionality reduction processing on the to-be-processed document so that the character length of the to-be-processed document is less than or equal to the threshold includes: segmenting the body of the to-be-processed document to obtain a plurality of sentences; obtaining a sentence feature vector for each sentence in the body and obtaining a title feature vector for the title; obtaining a similarity value between the title feature vector and each sentence feature vector; and splicing a plurality of sentences with higher similarity values to form the to-be-processed document after dimensionality reduction.
[0008] Among them, the sentence feature vector and the title feature vector are obtained based on a trained tpBERT model; among them, the step of training the tpBERT model includes: constructing a plurality of training text pairs, each training text pair includes a first training text and a second training text, and each training text pair is labeled with a similarity label, and the similarity label is 0 or 1; respectively performing feature extraction on the first training text and the second training text to obtain corresponding first output vectors and second output vectors; splicing the first output vector and the second output vector to obtain a first spliced vector; using a first activation function and the first spliced vector to obtain a similarity prediction value; and updating the parameters in the tpBERT model based on the similarity prediction value and the corresponding similarity label.
[0009] Among them, the step of splicing the first output vector and the second output vector to obtain a first spliced vector includes: subtracting the first output vector from the second output vector to obtain a first difference vector and subtracting the second output vector from the first output vector to obtain a second difference vector; performing an exclusive OR operation on the first output vector, the second output vector, the first difference vector, and the second difference vector to obtain the first spliced vector.
[0010] Among them, the step of clustering all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster includes: using the DBSCAN clustering algorithm to cluster all the documents to obtain at least one clustering cluster.
[0011] Among them, the step of generating a corresponding summary for each of the clustering clusters includes: for each sentence in each document in the current clustering cluster, obtaining the position vector of each word in the current sentence, and obtaining a semantic feature vector based on the position vectors of all the words in the sentence and the sentence feature vector of the sentence; encoding all the semantic feature vectors of each document in the current clustering cluster to obtain a hidden state vector of the middle layer; and decoding the hidden state vectors of all the documents in the current clustering cluster to obtain the summary.
[0012] Among them, the summary is obtained based on the trained MD_PGNet model; among them, the steps of training the MD_PGNet model include: constructing a plurality of training text clusters, each of the training text clusters including a plurality of training documents with a similarity exceeding a preset value, and each text cluster being set with a corresponding summary label; sequentially and parallelly inputting the words in a plurality of training documents in the same training text cluster into the MD_PGNet model; obtaining the decoding state vectors of all the words decoded in the MD_PGNet model at the current time step, and the encoding state vectors corresponding to all the words decoded; obtaining the average attention of the training text cluster based on the decoding state vector and the encoding state vector; obtaining the copy weight probability based on the average attention, the decoding state vector and the summary label, and obtaining the generation word probability at the current time step based on the decoding state vector and the encoding state vector; obtaining the predicted word probability at the current time step based on the copy weight probability and the generation word probability; obtaining the total loss based on the predicted word probability and the loss of the coverage vectors of all the training documents in the training text cluster, and adjusting the parameters in the MD_PGNet model according to the total loss.
[0013] Among them, the step of obtaining the average attention of the training text cluster based on the decoding state vector and the encoding state vector includes: obtaining the attention weight of each word based on the decoding state vector and the encoding state vector of each word; for non-first words, updating the attention weight of the non-first words based on the attention weight and the coverage vector; obtaining the sum value of the attention weights of all the words decoded in each training text, and multiplying the sum value of each training text by the corresponding passage weight coefficient to obtain a first value; and taking the average value of the first values of all the training texts in the training text cluster as the average attention.
[0014] Among them, the step of obtaining the total loss based on the predicted word probability and the coverage vector of all training documents in the training text cluster includes: obtaining the sum of the first losses of the coverage vectors of each training text in the training text cluster at the current time step; obtaining the product of the second coefficient and the sum of the first losses, and taking the difference between the product and the logarithm value of the predicted word probability as the total loss.
[0015] To solve the above technical problems, another technical solution adopted by this application is: to provide a text clustering device, including: an obtaining module, configured to obtain destination information; a processing module, coupled to the obtaining module, configured to, in response to the existence of a to-be-processed document with a character length exceeding a threshold among a plurality of the documents, perform dimensionality reduction processing on the to-be-processed document so that the character length of the to-be-processed document is less than or equal to the threshold; a clustering module, coupled to the processing module, configured to cluster all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster; a generating module, coupled to the clustering module, configured to generate a corresponding summary for each of the clustering clusters.
[0016] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is configured to execute the program instructions to implement the text clustering method mentioned in any of the above embodiments.
[0017] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is used to implement the text clustering method mentioned in any of the above embodiments.
[0018] Different from the prior art, the beneficial effect of this application is: the text clustering method provided by this application includes: obtaining a plurality of documents; in response to the existence of a to-be-processed document with a character length exceeding a threshold among the plurality of documents, performing dimensionality reduction processing on the to-be-processed document so that the character length of the to-be-processed document is less than or equal to the threshold; clustering all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster; generating a corresponding summary for each of the clustering clusters. Through this design method, the PGNet model is improved, enabling it to process multiple documents simultaneously and obtain the common summary of multi-chapter documents. The multi-chapter text short description generation method based on the PGNet model solves the defect in the traditional solution that can only solve the summary generation of a single document. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0020] Figure 1 is the overall framework diagram of the text clustering method provided by the present application;
[0021] Figure 2 is the flowchart of an implementation manner of the text clustering method of the present application;
[0022] Figure 3 is Figure 2 the flowchart of an implementation manner corresponding to step S2 in;
[0023] Figure 4 is the flowchart of training the tpBERT model;
[0024] Figure 5 is the structural diagram of the tpBERT model;
[0025] Figure 6 is Figure 4 the flowchart of an implementation manner of step S22 in;
[0026] Figure 7 is Figure 2 the flowchart of an implementation manner of step S4 in;
[0027] Figure 8 is the flowchart of training the MD_PGNet model;
[0028] Figure 9 is the network structure diagram of the MD_PGNet model;
[0029] Figure 10 is Figure 8 the flowchart of an implementation manner of step S43 in;
[0030] Figure 11 is Figure 8 the flowchart of an implementation manner of step S46 in;
[0031] Figure 12 is the framework diagram of an implementation manner of the text clustering device of the present application;
[0032] Figure 13 is the framework diagram of an implementation manner of the electronic device of the present application;
[0033] Figure 14It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0035] Please refer to Figure 1 and Figure 2 , Figure 1 is the overall framework diagram of the text clustering method provided by the present application. Figure 2 is the schematic flowchart of an embodiment of the text clustering method of the present application. The overall solution provided by the present application is as follows: input multiple documents. When there are ultra-long texts with character lengths exceeding the threshold in the multiple documents, use the tpBERT model based on sentence pairs to perform semantic dimensionality reduction processing on the above ultra-long texts, and complete text clustering without a specified number of clusters (for example, clustering into politics, entertainment, etc.). When there are no ultra-long texts with character lengths exceeding the threshold in the documents, directly enter the text clustering step, and finally generate short description topics for multiple chapters and output. Specifically, the text clustering method includes:
[0036] S1: Obtain multiple documents.
[0037] Specifically, the above documents include multiple documents to be processed.
[0038] S2: Determine whether there are documents to be processed with character lengths exceeding the threshold in the multiple documents.
[0039] S3: If so, perform dimensionality reduction processing on the documents to be processed so that the character lengths of the documents to be processed are less than or equal to the threshold.
[0040] Specifically, if there are documents to be processed with character lengths exceeding the threshold in the multiple documents, perform dimensionality reduction processing on the documents to be processed so that the character lengths of the documents to be processed are less than or equal to the threshold. Due to the limitation of the input length of 512 characters of the tpBERT model, and due to the neural network structure of BERT itself, experimental analysis shows that exceeding 300 characters will cause failure to capture longer semantic information. Therefore, the tpBERT model cannot completely obtain the semantic information of the long chapter content when processing long chapter text data. Therefore, if there are documents to be processed with character lengths exceeding the threshold in the multiple documents, it is necessary to first perform semantic dimensionality reduction processing on the ultra-long text.
[0041] Specifically, in this embodiment, the documents to be processed include the main text and the title. Please refer toFigure 3 , Figure 3 is Figure 2 The flowchart corresponding to step S3 in Figure 2 is a schematic diagram of a process of an embodiment. Specifically, the step of dimensionality reduction processing on the document to be processed in step S2 to make the character length of the document to be processed less than or equal to the threshold includes:
[0042] S10: Segment the body text of the document to be processed to obtain a plurality of sentences.
[0043] Specifically, segment the body text of the document to be processed according to punctuation marks, that is, convert a long article into a list of sentences, and the above list of sentences includes a plurality of sentences in the body text. Of course, in other embodiments, the body text of the document to be processed can also be segmented in other ways, which is not limited in this application.
[0044] S11: Obtain the sentence feature vector of each sentence in the body text and the title feature vector of the title.
[0045] Specifically, in this embodiment, the sentence feature vector and the title feature vector are obtained based on the trained tpBERT model. The sentence feature vector representation of each sentence is obtained based on the trained tpBERT model, and an article is represented as a plurality of vectors. In addition, in this embodiment, when actually applied, a single sentence is input, and the sentence feature vector corresponding to each sentence is obtained based on the tpBERT model, and it will not enter the step of the activation function softmax. Through the above design method, although the input texts are different, if the meanings of two sentences are the same, the sentence feature vectors output by these two sentences are the same. Further, for an article, the importance of the title is also very high. The title can be vectorized based on the optimized trained tpBERT model to obtain the title feature vector of the title.
[0046] S12: Obtain the similarity value between the title feature vector and each sentence feature vector.
[0047] Specifically, use cosine similarity to calculate the similarity between the title feature vector of the above title and the sentence feature vector of each sentence in the article list. In this embodiment, based on the topN analysis method, the top N similar sentences are dynamically selected so that the total length of the top N similar sentences is less than 512 bytes. The practical significance of this step is that the top N sentences with high similarity to the title can express most of the semantic information of the entire article. Among them, the number of N is determined by the total length of the sentences.
[0048] S13: Concatenate multiple sentences with higher similarity values to form the document to be processed after dimensionality reduction.
[0049] Specifically, multiple statements with relatively high similarity values (the first N similar statements) are concatenated into a new long passage content (new_discourse). At this time, the total length of the text is reduced to less than 512 characters, thereby achieving semantic dimensionality reduction of long texts.
[0050] In the process of processing long texts, a semantic dimensionality reduction method is used to perform clause retrieval on the longer texts, retaining the semantic feature information in the original texts, so that the obtained text vector representation eliminates the useless information in the long texts, reduces the complexity of subsequent processing, and facilitates subsequent pre-training language model processing.
[0051] The current state-of-the-art natural language processing method is nothing more than the pre-training model + fine-tuning approach. This approach proposes an improved pre-training language model: tpBERT model specifically designed for the characteristics of text pairs (text pair). The standard BERT model is a simple semantic modeling based on text context information and cannot explicitly capture the key information in sentences. Therefore, when facing similar text pairs, the content between the two statements is highly similar, and only the information of a small number of keywords is different. It is difficult for the BERT model to effectively distinguish, thus affecting the deep semantics between the entire statement pairs, and even the semantics are completely opposite. Therefore, the improved method based on the BERT model is to use a Keyword-BERT model as the core architecture of this method, and perform optimization and improvement of text pairs on this basis, realizing a new pre-training language model, the tpBERT model. Specifically, please refer to Figure 4 and Figure 5 , Figure 4 is a schematic diagram of the process of training the tpBERT model, Figure 5 is a schematic diagram of the structure of the tpBERT model. Specifically, the steps of training the tpBERT model include:
[0052] S20: Construct multiple training text pairs.
[0053] Specifically, when training the tpBERT model, it is based on text pairs. Each training text pair includes a first training text Sentence1 and a second training text Sentence2, and each training text pair is labeled with a similarity label. In this embodiment, the similarity label of each training text pair can be manually labeled as 0 or 1. Specifically, when the first training text Sentence1 and the second training text Sentence2 are similar, their similarity label is 1; when the first training text Sentence1 and the second training text Sentence2 are not similar, their similarity label is 0.
[0054] S21: Extract features from the first training text and the second training text respectively to obtain corresponding first output vector and second output vector.
[0055] Specifically, input the first training text Sentence1 and the second training text Sentence2 into the keyword-BERT model respectively, and obtain the first output vector Emb1 and the second output vector Emb2 through the feature extraction layer Transformer Layer respectively.
[0056] S22: Concatenate the first output vector and the second output vector to obtain a first concatenated vector.
[0057] Specifically, in this embodiment, please refer to Figure 6 , Figure 6 is Figure 4 a schematic flowchart of an implementation manner of step S22 in
[0058] S220: Subtract the first output vector from the second output vector to obtain a first difference vector, and subtract the first output vector from the second output vector to obtain a second difference vector.
[0059] Specifically, subtract the first output vector Emb1 from the second output vector Emb2 to obtain a first difference vector Emb1 - Emb2, and subtract the first output vector Emb1 from the second output vector Emb2 to obtain a second difference vector Emb2 - Emb1. Among them, the first difference vector Emb1 - Emb2 and the second difference vector Emb2 - Emb1 are used to enhance mutual information.
[0060] S221: Perform exclusive OR processing on the first output vector, the second output vector, the first difference vector, and the second difference vector to obtain a first concatenated vector.
[0061] Specifically, in this embodiment, perform exclusive OR processing on the first output vector Emb1, the second output vector Emb2, the first difference vector Emb1 - Emb2, and the second difference vector Emb2 - Emb1 to obtain a first concatenated vector Output (i.e., the new output feature), and its calculation formula is: Output = Emb1 ⊕ Emb2 ⊕ (Emb2 - Emb1) ⊕ (Emb1 - Emb2), where ⊕ is the exclusive OR operator. Of course, in other embodiments, the first difference vector Emb1 - Emb2 or the second difference vector Emb2 - Emb1 in the formula can also be replaced by the product Emb1 * Emb2 of the first output vector Emb1 and the second output vector Emb2, which is not limited in this application.
[0062] S23: Obtain a similarity prediction value by using the first activation function and the first concatenated vector.
[0063] Specifically, the first splicing vector Output is output to the first activation function softmax, and the probabilities of the similarity label of the text pair being 1 and the similarity label of the text pair being 0 are output respectively. Finally, the maximum probability is output as the similarity prediction value.
[0064] S24: Update the parameters in the tpBERT model based on the similarity prediction value and the corresponding similarity label.
[0065] Specifically, perform fine-tuning training on a large number of similar text pair datasets, and update the parameters in the tpBERT model and the keyword-BERT model based on the similarity prediction value and the corresponding similarity label until convergence is achieved, which means the training is completed. The method for determining whether the convergence state is reached can be to determine whether the number of training times reaches the set number, or other methods, which are not limited in this application. The improved keybert-BERT model is used to calculate the text feature vector, and then a text clustering model is constructed through the feature vector. The improved keybert-BERT model modifies the loss function of the text pre-training language model, thereby optimizing the text feature vector expression method.
[0066] S4: Otherwise, directly proceed to step S5.
[0067] Specifically, if there is no document to be processed with a character length exceeding the threshold in multiple documents, directly proceed to the step of clustering all documents with a character length less than or equal to the threshold to obtain at least one clustering cluster.
[0068] S5: Cluster all documents with a character length less than or equal to the threshold to obtain at least one clustering cluster.
[0069] Furthermore, in this embodiment, step S5 specifically includes: using the DBSCAN clustering algorithm to cluster all documents to obtain at least one clustering cluster. Specifically, for a large number of document sets composed of the dimensionality-reduced long passage content (new_discourse), the DBSCAN clustering algorithm is used to implement text clustering without specifying the number of clusters. DBSCAN is a density-based spatial clustering algorithm that infers the number of clusters based on the data and can generate clusters for any shape. Step S5 adopts a publicly common method. Of course, other similar methods can also be used to cluster all documents, which are not limited in this application. The specific process of the DBSCAN clustering algorithm is as follows:
[0070] Input: Sample set D = (x1, x2,..., x m), neighborhood parameters (∈, MinPts), and sample distance measurement methods. Specifically, the neighborhood parameters (∈, MinPts) are used to describe the tightness of the sample distribution in the neighborhood. Among them, ∈ describes the neighborhood distance threshold of a certain sample, and MinPts describes the threshold of the number of samples in the neighborhood with a distance of ∈ from a certain sample.
[0071] Output: Cluster partition C.
[0072] 1) Initialize the set of core objects Initialize the number of clustering clusters k = 0, initialize the set of unvisited samples Γ = D, and the cluster partition
[0073] 2) For j = 1, 2,... m, find all core objects according to the following steps:
[0074] a) Through the distance measurement method, find the sample x j 's ∈-neighborhood subsample set N∈(x j );
[0075] b) If the number of samples in the subsample set satisfies |N∈(x j )| ≥ MinPts, add the sample xj to the set of core object samples: Ω = Ω ∪ {x j};
[0076] 3) If the set of core objects then the algorithm ends, otherwise go to step 4;
[0077] 4) In the set of core objects Ω, randomly select a core object o, initialize the current cluster core object queue Ωcur = {o}, initialize the category serial number k = k + 1, initialize the current cluster sample set Ck = {o}, and update the set of unvisited samples Γ = Γ - {o};
[0078] 5) If the current cluster core object queue then the current clustering cluster C k is generated, update the cluster partition C = {C1, C2,..., C k}, update the set of core objects Ω = Ω - C k , and go to step 3. Otherwise, update the set of core objects Ω = Ω - C k ;
[0079] 6) Take out a core object o' from the current cluster core object queue Ωcur, find all ∈-neighborhood subsample sets N∈(o') through the neighborhood distance threshold ∈, let Δ = N∈(o') ∩ Γ, and update the current cluster sample set C k = C k∪Δ, update the unvisited sample set Γ = Γ - Δ, update Ωcur = Ωcur ∪ (Δ ∩ Ω) - o′, and go to step 5;
[0080] The output result is: the cluster partition C = {C1, C2,..., C k}.
[0081] S6: Generate corresponding summaries for each clustering cluster.
[0082] For the multiple clustering clusters obtained in the above step S5, each clustering cluster contains multiple documents. Take each document as a discourse and input it into the trained MD_PGNet model to obtain the corresponding short summary of the clustering cluster. Specifically, in this embodiment, please refer to Figure 7 , Figure 7 is Figure 2 a schematic flowchart of an implementation manner of step S6 in
[0083] S30: For each sentence in each document in the current clustering cluster, obtain the position vector of each word in the current sentence, and obtain the semantic feature vector based on the position vectors of all words in the sentence and the sentence feature vector of the sentence.
[0084] Specifically, in this embodiment, a function of obtaining the position vector of each word in the current sentence is newly added to the text input. Specifically, for the input form of the encoder Encoder, this application adopts a new combined form: input embedding + position embedding (PE), and at the same time utilizes the absolute position and relative position of the words. The position vector can be expressed as a linear transformation of the feature vector of the position, which provides great convenience for the model to capture the relative position relationship between words. The position vector can be encoded in the form of a sine-cosine function, as follows:
[0085]
[0086]
[0087] input = input + PE) (3)
[0088] Among them, i represents the dimension, pos represents the word vector at each position, dim represents the dimension size of the feature vector after dimensionality reduction of the input document, which can be 512 here. If the dimension in a certain document is less than 512, it can be supplemented to 512. Specifically, in this embodiment, if a certain word in a sentence is the nth (n is an even number) word, the sine function (Formula 1) is used for position encoding. If a certain word in a sentence is the nth (n is an odd number) word, the cosine function (Formula 2) is used for position encoding. After obtaining the position vector of each word using the above sine function and cosine function, the position vector of each word is superimposed with the sentence feature vector of the sentence where the word is located to obtain the semantic feature vector of the sentence. This can significantly enhance the position-related information between text characters, thereby enhancing the expression of semantic features.
[0089] S31: Encode all semantic feature vectors of each document in the current clustering cluster to obtain the hidden state vector of the intermediate layer.
[0090] Specifically, in this embodiment, the encoder Encoder is used to encode all semantic feature vectors of each document in the current clustering cluster to obtain a hidden state vector h of the intermediate layer i . The above encoder Encoder can adopt a bidirectional LSTM network structure, etc., and the present application does not limit this here.
[0091] S32: Decode the hidden state vectors of all documents in the current clustering cluster to obtain the summary.
[0092] Specifically, after obtaining the hidden state vector h of the intermediate layer in step S31 i , the decoder Decoder is used to decode the hidden state vectors h of all documents in the current clustering cluster i . Finally, the summary short summary corresponding to the current clustering cluster is obtained. Here, the summary short summary refers to the short description topic text that can express the meaning of all documents in the current clustering cluster. The above decoder Dncoder can adopt a unidirectional LSTM network structure, etc., and the present application does not limit this here.
[0093] Preferably, in this embodiment, the summary in step S6 is obtained based on the trained MD_PGNet model. Since the networks of the previous ordinary encoder architectures mainly have two problems: (1) the text to be generated cannot obtain the important semantic vocabulary that can express the meaning from the original input text; (2) there is a situation of duplicate content in the generated text. Therefore, in the present application, the PGNet model is referred to and two mechanisms are introduced to solve the corresponding problems. These two mechanisms are the generation-copy selection mechanism and the duplicate text coverage mechanism. Please refer to Figure 8 andFigure 9 , Figure 8 is a schematic flow chart for training the MD_PGNet model, Figure 9 and is a network structure diagram of the MD_PGNet model. Specifically, the steps for training the MD_PGNet model include:
[0094] S40: Construct multiple training text clusters.
[0095] Specifically, in this embodiment, each training text cluster includes multiple training documents with a similarity exceeding a preset value, and each text cluster is set with a corresponding abstract label. The above preset value can be artificially set according to actual needs, and this application does not make a limitation here.
[0096] S41: Sequentially and in parallel input the words in multiple training documents in the same training text cluster into the MD_PGNet model.
[0097] Specifically, during the training phase of the MD_PGNet model, the words in the training documents in each training text cluster are input into the MD_PGNet model. In this embodiment, the words in the training documents in the same training text cluster are sequentially and in parallel input into the MD_PGNet model.
[0098] S42: Obtain the decoding state vectors of all the words decoded in the MD_PGNet model at the current time step, and the encoding state vectors corresponding to all the decoded words.
[0099] Specifically, in this embodiment, in the MD_PGNet model, the decoding state vector of the i-th word at time step t after decoding is s t , and the encoding state vector of the i-th word after decoding is h i .
[0100] S43: Obtain the average attention of the training text cluster based on the decoding state vector and the encoding state vector.
[0101] Specifically, in this embodiment, please refer to Figure 10 , Figure 10 which Figure 8 is a schematic flow chart of an implementation manner of step S43 in
[0102] S430: Obtain the attention weights of the words based on the decoding state vectors and the encoding state vectors of each word.
[0103] Specifically, use the decoding state vector s t of each word obtained in step S43 above at time step t and the encoding state vector h i to obtain the attention weight a of the word at this time step tt , and its specific calculation formula is:
[0104] q = W h h i + W s s t + b1(4)
[0105] a t = softmax(v T tanh(q)) (5)
[0106] Among them, W h , W s , and b are all parameters of the neural network, q is the calculated input feature, v T is the attention coefficient, and formula 5 is the calculation method of the standard attention weight.
[0107] S431: For non-first words, update the attention weights of non-first words based on word attention weights and the coverage vector.
[0108] In this embodiment, the words decoded at the next time step include the words decoded at the previous time step. Then, the attention weight a t of the words that have been decoded at the previous time step will also change. To solve the problem of repetitive text generation, it is necessary to establish a coverage vector c i , and add it to the attention weight a t . Therefore, during the training process, the input feature q is recalculated starting from the second word. The purpose of this is to prevent the repetition of the generated text. Specifically, in this embodiment, for non-first words, the input feature q of non-first words is updated using the attention weight a t and the coverage vector c i . Its specific calculation methods are shown in formulas 6 and 7, and finally it is returned to formula 5 to recalculate the attention weight a t .
[0109]
[0110]
[0111] S432: Obtain the sum value of the attention weights of all the words that have been decoded in each training text, and multiply the sum value of each training text by the corresponding passage weight coefficient to obtain the first value.
[0112] Specifically, by calculating, obtain the attention weight a tThe sum value is obtained, and the sum value is multiplied by the corresponding passage weight coefficient γ to obtain the first value A. The calculation formula for the first value A is:
[0113] A = γ(∑ i a t ) (8)
[0114] Among them, the passage weight coefficient γ can be set according to the actual situation, and this application does not limit it here.
[0115] S433: Use the average value of the first values of all training texts in the training text cluster as the average attention.
[0116] Specifically, calculate the average value of the first value A in the above step S432, and this average value is the average attention a n_disc , and the average attention a n_disc here represents the cluster attention weight vector of each training text cluster. Generally speaking, the calculation formula for the average attention a n_disc is as shown in Formula 9:
[0117]
[0118] Among them, j is the article number, j = (1, 2... n), and γ is the passage weight coefficient.
[0119] S44: Obtain the copy weight probability based on the average attention, the decoding state vector, and the summary label, and obtain the generation word probability at the current time step based on the decoding state vector and the encoding state vector.
[0120] Specifically, in this embodiment, a selection threshold needs to be set to determine whether the content generated at time step t comes from the vocabulary or the input text, that is, to generate the copy weight probability p gen . Specifically, use the average attention a n_disc , the decoding state vector s t and the summary label to obtain the copy weight probability p gen , and the specific calculation formula is:
[0121] p gen = α * (W h a n_disc + W s s t + W x x t + b2) (10)
[0122] Among them, α is the weight coefficient of each cluster, and x t is the input of the decoder, that is, the summary label corresponding to each text cluster.
[0123] In addition, in this embodiment, the decoding state vector s t and the encoding state vector h i are mapped through a linear layer to obtain a word probability distribution, and then the probability p of the generated word at the current time step t is obtained through the activation function softmax vocab , and the word to be generated can be predicted. The calculation formula is:
[0124] p vocab = softmax(V([s t , h i ) + b3) (11)
[0125] where [s t , h i is concatenation.
[0126] S45: Obtain the predicted word probability at the current time step based on the copy weight probability and the generated word probability.
[0127] Specifically, use the copy weight probability p gen and the generated word probability p vocab to obtain the predicted word probability P of the word (character) predicted to be generated at the current time step t. At the same time, the vocabulary in the input data will be added to expand the existing word list to form a larger word list to expand the word list.
[0128] P = p gen * p vocab + (1 - p gen ) * a n_disc (12)
[0129] S46: Obtain the total loss based on the predicted word probability and the loss of the coverage vector of all training documents in the training text cluster, and adjust the parameters in the MD_PGNet model according to the total loss.
[0130] Specifically, in this embodiment, please refer to Figure 11 , Figure 11 is Figure 8 a schematic flowchart of an implementation manner of step S46 in
[0131] S460: Obtain the sum of the first losses of the coverage vectors of each training text in the training text cluster at the current time step.
[0132] Specifically, use the loss of the general coverage mechanism of the PGNet model, combined with the loss of the predicted word probability P, on the newly improved model to complete the training of the summary generation model for multiple documents. First, calculate the sum of the first losses cov_loss of the coverage vectors of all documents in the training text cluster at the current time step t:
[0133]
[0134] Among them, are the attention weight and the coverage vector of the i-th word in a single chapter respectively, and j is the article number.
[0135] S461: Obtain the product of the sum of the second coefficient and the first loss, and use the difference between the product and the logarithm of the predicted word probability as the total loss.
[0136] Specifically, after obtaining the sum of the first losses cov_loss at the current time step t in step S460, the loss functions of the two processes are weighted and calculated, and finally the total loss loss is obtained. The specific calculation formula is:
[0137] loss = -logP(w t ) + θcov_loss t (14)
[0138] Among them, θ is the weighting coefficient.
[0139] In summary, by improving the PGNet model, it can process multiple documents simultaneously and obtain the common summary of multiple documents. The multi-chapter text short description generation method based on the PGNet model solves the defect that only the summary generation of a single document can be solved in the traditional solution.
[0140] Please refer to Figure 12 , Figure 12 is a schematic framework diagram of an embodiment of the text clustering device of the present application. The parking lot recommendation device specifically includes:
[0141] An obtaining module 10, configured to obtain multiple documents.
[0142] A processing module 11, coupled to the obtaining module 10, configured to perform dimensionality reduction processing on a to-be-processed document in response to the to-be-processed document having a character length exceeding a threshold among the multiple documents, so that the character length of the to-be-processed document is less than or equal to the threshold.
[0143] A clustering module 12, coupled to the processing module 11, configured to cluster all documents with a character length less than or equal to the threshold to obtain at least one clustering cluster.
[0144] A generating module 13, coupled to the clustering module 12, configured to generate a corresponding summary for each clustering cluster.
[0145] Optionally, in this embodiment, the processing module 11 includes a segmentation module, a feature vector module, a similarity module, and a first splicing module that are sequentially connected to each other. Specifically, the segmentation module is configured to segment the body text of the document to be processed to obtain a plurality of sentences; the feature vector module is configured to obtain the sentence feature vectors of each sentence in the body text and the title feature vector of the title; the similarity module is configured to obtain the similarity values between the title feature vector and each sentence feature vector; the first splicing module is configured to splice a plurality of sentences with higher similarity values to form the document to be processed after dimensionality reduction.
[0146] In one embodiment, the text clustering device provided by the present application further includes a first training module 14 for training the tpBERT model. Optionally, in this embodiment, the first training module 14 is connected to the obtaining module 10. Of course, in other embodiments, both ends of the first training module 14 may also be connected to the obtaining module 10 and the processing module 11, which is not limited herein. Among them, the first training module 14 includes a first construction module, a feature extraction module, a second splicing module, a similarity prediction value module, and a first update module that are sequentially connected to each other. Specifically, the first construction module is configured to construct a plurality of training text pairs, where each training text pair includes a first training text and a second training text, and each training text pair is labeled with a similarity label, and the similarity label is 0 or 1; the feature extraction module is configured to perform feature extraction on the first training text and the second training text respectively to obtain corresponding first output vectors and second output vectors; the second splicing module is configured to splice the first output vector and the second output vector to obtain a first spliced vector; the similarity prediction value module is configured to obtain a similarity prediction value by using a first activation function and the first spliced vector; the first update module is configured to update the parameters in the tpBERT model based on the similarity prediction value and the corresponding similarity label.
[0147] Optionally, in this embodiment, the second splicing module includes a difference vector module and an exclusive OR module that are sequentially connected to each other. The difference vector module is configured to subtract the first output vector from the second output vector to obtain a first difference vector and subtract the second output vector from the first output vector to obtain a second difference vector; the exclusive OR module is configured to perform an exclusive OR operation on the first output vector, the second output vector, the first difference vector, and the second difference vector to obtain a first spliced vector.
[0148] Further, the clustering module is specifically configured to cluster all documents by using the DBSCAN clustering algorithm to obtain at least one clustering cluster.
[0149] In yet another embodiment, the generation module 13 specifically includes a semantics module, an encoding module, and a decoding module that are sequentially connected to each other. Specifically, the semantics module is used to obtain the position vector of each word in the current statement for each statement in each document in the current clustering cluster, and obtain a semantic feature vector based on the position vectors of all words in the statement and the statement feature vector of the statement; the encoding module is used to encode all semantic feature vectors of each document in the current clustering cluster to obtain a hidden state vector in the middle layer; the decoding module is used to decode the hidden state vectors of all documents in the current clustering cluster to obtain a summary.
[0150] In yet another embodiment, the text clustering device provided in this application further includes a second training module 15 for training the MD_PGNet model. Optionally, in this embodiment, the second training module 15 is connected to the obtaining module 10. Of course, in other embodiments, the two ends of the second training module 15 can also be respectively connected to the clustering module 12 and the generation module 13, which is not limited in this application. Among them, the second training module 15 includes a second construction module, an input module, an encoding and decoding module, an attention module, a first probability module, a second probability module, and a loss module that are sequentially connected to each other. The second construction module is used to construct a plurality of training text clusters, where each training text cluster includes a plurality of training documents with a similarity exceeding a preset value, and each text cluster is set with a corresponding summary label; the input module is used to sequentially and parallelly input the words in a plurality of training documents in the same training text cluster into the MD_PGNet model; the encoding and decoding module is used to obtain the decoding state vectors of all words decoded in the MD_PGNet model at the current time step, and the encoding state vectors corresponding to all words decoded; the attention module is used to obtain the average attention of the training text cluster based on the decoding state vector and the encoding state vector; the first probability module is used to obtain a copy weight probability based on the average attention, the decoding state vector, and the summary label, and obtain a generation word probability at the current time step based on the decoding state vector and the encoding state vector; in addition, the second probability module is further used to obtain a predicted word probability at the current time step based on the copy weight probability and the generation word probability; the loss module is used to obtain a total loss based on the loss of the predicted word probability and the coverage vectors of all training documents in the training text cluster, and adjust the parameters in the MD_PGNet model according to the total loss.
[0151] Optionally, in this embodiment, the attention module further includes a weight module, a second update module, a first numerical value module, and an average value module that are sequentially connected to each other. Specifically, the weight module is used to obtain the attention weight of a word based on the decoding state vector and the encoding state vector of each word; the second update module is used to update the attention weight of a non-first word based on the attention weight and the coverage vector for a non-first word; the first numerical value module is used to obtain the sum value of the attention weights of all the words that have been decoded in each training text, and multiply the sum value of each training text by the corresponding passage weight coefficient to obtain a first numerical value; the average value module is used to take the average value of the first numerical values of all the training texts in the training text cluster as the average attention.
[0152] Optionally, in this embodiment, the loss module further includes a first loss sum module and a total loss module that are connected to each other. Specifically, the first loss sum module is used to obtain the sum of the first losses of the coverage vectors of each training text in the training text cluster at the current time step; the total loss module is used to obtain the product of the second coefficient and the sum of the first losses, and take the difference between the product and the logarithm value of the predicted word probability as the total loss.
[0153] Please refer to Figure 13 , Figure 13 which is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device includes a memory 20 and a processor 22 that are coupled to each other. Specifically, in this embodiment, program instructions are stored in the memory 20, and the processor 22 is configured to execute the program instructions to implement the text clustering method mentioned in any of the above embodiments.
[0154] Specifically, the processor 22 can also be referred to as a CPU (Central Processing Unit). The processor 22 may be an integrated circuit chip with signal processing capabilities. The processor 22 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 22 may be implemented jointly by multiple integrated circuit chips.
[0155] Please refer to Figure 14 , Figure 14It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 30 stores a computer program 300 that can be read by a computer. The computer program 300 can be executed by a processor to implement the text clustering method mentioned in any of the above embodiments. Among them, the computer program 300 can be stored in the above computer-readable storage medium 30 in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The computer-readable storage medium 30 with a storage function can be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.
[0156] In summary, different from the prior art, the text clustering method provided by the present application includes: obtaining a plurality of documents; in response to the existence of a document to be processed with a character length exceeding a threshold among the plurality of documents, performing dimensionality reduction processing on the document to be processed so that the character length of the document to be processed is less than or equal to the threshold; clustering all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster; generating a corresponding abstract for each clustering cluster. Through this design method, the PGNet model is improved so that it can process multiple documents simultaneously and obtain a common abstract of multiple multi-chapter documents. The multi-chapter text short description generation method based on the PGNet model solves the defect that only the abstract generation of a single document can be solved in the traditional solution.
[0157] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.
Claims
1. A text clustering method, characterized in that, Including: Obtaining a plurality of documents; In response to the presence of a to-be-processed document in the plurality of documents whose character length exceeds a threshold, performing dimensionality reduction processing on the to-be-processed document so that the character length of the to-be-processed document is less than or equal to the threshold; Clustering all the documents whose character length is less than or equal to the threshold to obtain at least one clustering cluster; Generating a corresponding summary for each of the clustering clusters; The summary is obtained based on the trained MD_PGNet model; wherein, the steps of training the MD_PGNet model include: constructing a plurality of training text clusters, each training text cluster includes a plurality of training documents with a similarity exceeding a preset value, and each text cluster is set with a corresponding summary label; sequentially and parallelly inputting the words in the plurality of training documents in the same training text cluster into the MD_PGNet model; obtaining the decoding state vectors of all the words decoded in the MD_PGNet model at the current time step, and the encoding state vectors corresponding to all the words decoded; obtaining the average attention of the training text cluster based on the decoding state vector and the encoding state vector; obtaining the copy weight probability based on the average attention, the decoding state vector and the summary label, and obtaining the generation word probability at the current time step based on the decoding state vector and the encoding state vector; obtaining the predicted word probability at the current time step based on the copy weight probability and the generation word probability; obtaining the total loss based on the predicted word probability and the loss of the coverage vectors of all the training documents in the training text cluster, and adjusting the parameters in the MD_PGNet model according to the total loss; Wherein, the step of obtaining the average attention of the training text cluster based on the decoding state vector and the encoding state vector includes: obtaining the attention weight of each word based on the decoding state vector and the encoding state vector of the word; for non-first words, updating the attention weight of the non-first words based on the attention weight and the coverage vector; obtaining the sum value of the attention weights of all the words decoded in each training text, and multiplying the sum value of each training text by the corresponding passage weight coefficient to obtain a first value; taking the average value of the first values of all the training texts in the training text cluster as the average attention.
2. The text clustering method according to claim 1, characterized in that The to-be-processed document includes a main body and a title, and the step of performing dimensionality reduction processing on the to-be-processed document so that the character length of the to-be-processed document is less than or equal to the threshold includes: Segmenting the main body of the to-be-processed document to obtain a plurality of sentences; Obtaining the sentence feature vector of each sentence in the main body, and obtaining the title feature vector of the title; Obtaining the similarity value between the title feature vector and each sentence feature vector; Concatenating the first N similar sentences to form the dimensionality-reduced to-be-processed document.
3. The text clustering method according to claim 2, characterized in that, The sentence feature vector and the title feature vector are obtained based on the feature extraction layer of the trained tpBERT model; wherein, the steps of training the tpBERT model include: Construct multiple training text pairs, where each training text pair includes a first training text and a second training text, and each training text pair is labeled with a similarity label, and the similarity label is 0 or 1; Extract features from the first training text and the second training text respectively to obtain corresponding first output vectors and second output vectors; Concatenate the first output vector and the second output vector to obtain a first concatenated vector; Use a first activation function and the first concatenated vector to obtain a similarity prediction value; Update the parameters in the tpBERT model based on the similarity prediction value and the corresponding similarity label.
4. The text clustering method according to claim 3, wherein The step of concatenating the first output vector and the second output vector to obtain a first concatenated vector includes: Subtract the second output vector from the first output vector to obtain a first difference vector, and subtract the first output vector from the second output vector to obtain a second difference vector; Perform an exclusive OR operation on the first output vector, the second output vector, the first difference vector, and the second difference vector to obtain the first concatenated vector.
5. The text clustering method according to claim 1, wherein The step of clustering all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster includes: Use the DBSCAN clustering algorithm to cluster all the documents to obtain at least one clustering cluster.
6. The text clustering method according to claim 1, wherein The step of generating a corresponding summary for each clustering cluster includes: For each sentence in each document in the current clustering cluster, obtain the position vector of each word in the current sentence, and obtain a semantic feature vector based on the position vectors of all the words in the sentence and the sentence feature vector of the sentence; Encode all the semantic feature vectors of each document in the current clustering cluster to obtain a hidden state vector in the middle layer; Decode the hidden state vectors of all the documents in the current clustering cluster to obtain the summary.
7. The text clustering method according to claim 1, wherein The step of obtaining the total loss based on the predicted word probability and the loss of the coverage vectors of all the training documents in the training text cluster includes: Obtain the sum of the first losses of the coverage vectors of each training text in the training text cluster at the current time step; Obtain the product of a second coefficient and the sum of the first losses, and use the difference between the product and the logarithm of the predicted word probability as the total loss.
8. A text clustering device, characterized in that, Includes: An obtaining module for obtaining multiple documents; A processing module coupled to the obtaining module, configured to perform dimensionality reduction processing on a to-be-processed document in response to the existence of a to-be-processed document with a character length exceeding the threshold among the multiple documents, so that the character length of the to-be-processed document is less than or equal to the threshold; A clustering module coupled to the processing module, configured to cluster all the documents with a character length less than or equal to the threshold to obtain at least one clustering cluster; A generating module coupled to the clustering module, configured to generate a corresponding summary for each clustering cluster; The abstract is obtained based on the trained MD_PGNet model. The text clustering device further includes a second training module, and the second training module includes a second construction module, an input module, an encoding and decoding module, an attention module, a first probability module, a second probability module, and a loss module. The second construction module is configured to construct a plurality of training text clusters, each of the training text clusters includes a plurality of training documents with a similarity exceeding a preset value, and each text cluster is provided with a corresponding abstract label. The input module is configured to sequentially and parallelly input words in a plurality of training documents in the same training text cluster into the MD_PGNet model. The encoding and decoding module is configured to obtain the decoding state vectors of all words decoded in the MD_PGNet model at the current time step, and the encoding state vectors corresponding to all words decoded. The attention module is configured to obtain the average attention of the training text cluster based on the decoding state vector and the encoding state vector. The first probability module is configured to obtain a copy weight probability based on the average attention, the decoding state vector, and the abstract label, and obtain a generated word probability at the current time step based on the decoding state vector and the encoding state vector. The first probability module is configured to obtain a predicted word probability at the current time step based on the copy weight probability and the generated word probability. The loss module is configured to obtain a total loss based on the predicted word probability and the loss of the coverage vectors of all training documents in the training text cluster, and adjust the parameters in the MD_PGNet model according to the total loss. Wherein, the attention module includes a weight module, a second update module, a first numerical module, and an average value module. The weight module is configured to obtain the attention weight of each word based on the decoding state vector and the encoding state vector of each word. The second update module is configured to, for non-first words, update the attention weight of non-first words based on the attention weight and the coverage vector. The first numerical module is configured to obtain the sum value of the attention weights of all words decoded in each training text, and multiply the sum value of each training text by the corresponding passage weight coefficient to obtain a first numerical value. The average value module is configured to use the average value of the first numerical values of all training texts in the training text cluster as the average attention.
9. An electronic device, characterized in that, It includes a memory and a processor coupled to each other. The memory stores program instructions, and the processor is configured to execute the program instructions to implement the text clustering method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to implement the text clustering method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A method for automatic abstracting for electronic official documents of enterprises
CN106407182A
Text processing method and device, electronic equipment and computer readable storage medium
CN111737461A
Long text clustering method and device based on pre-training language model
CN112836043A
Text abstract generation method and device, equipment and storage medium
CN113268586A