Document blood relationship analysis method based on BERT model
The embedded vectors are generated through the BERT model and combined with the hierarchical density clustering algorithm and multi-stage screening mechanism, the problem of inflexible semantic modeling and clustering in document blood relationship mining is solved, and the accuracy and efficiency of document blood relationship recognition is improved.
Patent Information
- Application Number
- CN202510635445.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-22
AI Technical Summary
The prior art has insufficient semantic modeling capabilities and parameter rigidity in document kinship mining, resulting in low accuracy and practicality. Traditional methods cannot effectively capture text context semantics and inflexible clustering results.
The BERT model is used to generate embedded vectors, combine the hierarchical density clustering algorithm and multi-stage screening mechanism, and identify the blood relationship between documents through dimensionality reduction processing, hierarchical density clustering and statistical distance methods.
It improves the accuracy, recall and calculation efficiency of document blood relationship mining, solves the shortcomings caused by semantic loss and parameter rigidity in traditional methods, and achieves higher recognition accuracy and adaptability.
Smart Images

Figure CN120523943A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of document lineage relationship analysis, and specifically relates to a document lineage relationship analysis method based on the BERT model. Background Art
[0002] With the acceleration of informatization and digitalization, enterprises and organizations generate massive amounts of document data (such as contracts, reports, and technical documents) in their daily operations. These documents are not only the core carriers of business processes but also carry critical knowledge assets. However, traditional document management methods face many challenges when dealing with the relevance between documents: the need to mine document lineage relationships (such as version iterations, revision relationships, and content similarity) is increasingly urgent, but existing technologies have significant deficiencies in semantic modeling, parameter flexibility, and clustering accuracy. As a result, the accuracy and coverage of lineage relationship identification are difficult to meet actual needs.
[0003] The current mainstream document kinship mining methods mainly rely on two types of technologies: a dual screening mechanism based on the LDA topic model and a rule-based method based on keyword matching. The LDA topic model generates a topic distribution vector by counting the co-occurrence frequency of words in the document, and then performs a secondary screening based on the edit distance of the document title; the rule-based method identifies document associations through manually defined keywords (such as "revise" and "update"). Although these methods are effective in specific scenarios, LDA ignores contextual information based on the bag-of-words model, resulting in loss of semantic associations, and the number of topics needs to be manually set; the edit distance is less sensitive to titles with the same semantics but different expressions, making it easy to mistakenly delete related documents.
[0004] In summary, the core flaws of existing technologies lie in insufficient semantic modeling capabilities and rigid parameters. LDA topic models fail to capture the contextual semantics of text (for example, "dislike" and "like" may be classified as the same topic), and the pre-set number of topics can easily lead to overfitting or underfitting. Edit distance only measures character-level differences, limiting its ability to identify documents with similar semantics but different expressions.
[0005] Furthermore, traditional clustering algorithms (such as K-Means) require a preset number of clusters, making them difficult to adapt to the dynamic changes in document topic distribution, resulting in poor interpretability of clustering results. These issues collectively restrict the accuracy and practicality of document kinship mining, necessitating the introduction of more advanced semantic modeling and dynamic clustering methods to overcome these technical bottlenecks. Summary of the Invention
[0006] The present invention provides a document kinship analysis method based on the BERT model to solve the problems of low accuracy and practicality in document kinship mining caused by insufficient semantic modeling capabilities and parameter rigidity.
[0007] The technical solution adopted in the present invention is:
[0008] A document lineage relationship analysis method based on the BERT model.
[0009] According to a document set including multiple documents, an embedding vector set including multiple embedding vectors is obtained by using a BERT model, and a dimensionality reduction process is performed to reduce the dimensions of the embedding vectors;
[0010] According to the embedding vector corresponding to the target document and the embedding vector set, a target topic cluster containing multiple candidate documents is obtained through a hierarchical density clustering algorithm;
[0011] Obtaining documents having the same blood relationship with the target document by using a statistical distance method based on the target document and the candidate documents;
[0012] When the number of documents with the same blood relationship is less than a preset threshold, adjacent topic clusters are obtained through the similarity of probability distribution between topic clusters, and the adjacent topic clusters are merged into the target topic cluster. The statistical distance method is then used again to obtain documents with the same blood relationship as the target document.
[0013] The document kinship analysis method based on the BERT model provided in the present invention also has the following additional technical features:
[0014] According to the document set containing multiple documents, the BERT model is used to obtain an embedding vector set containing multiple embedding vectors, specifically:
[0015] According to the document, the query vector, key vector, and value vector are obtained one by one through the BERT model;
[0016] The embedding vector is obtained according to the query vector, the key vector, and the value vector.
[0017] Perform dimensionality reduction to reduce the embedding vector dimension, specifically:
[0018] The dimensionality reduction process reduces the dimension of the embedding vector by controlling the neighborhood size according to the distance between points.
[0019] The embedding vector is obtained according to the query vector, the key vector, and the value vector, specifically:
[0020] Obtaining a similarity score based on the query vector and the key vector, and obtaining a probability value based on the similarity score through a softmax function;
[0021] The value vector is weighted according to the probability value to obtain the embedding vector.
[0022] The hierarchical density clustering algorithm is specifically:
[0023] The hierarchical density clustering algorithm sets a cluster internal point density requirement, and when the point density obtained by the hierarchical density clustering is greater than or equal to the cluster internal point density requirement, a topic cluster is obtained;
[0024] The hierarchical density clustering algorithm sets a minimum number of documents for a valid cluster. When the number of documents in a subject cluster is greater than or equal to the minimum number of documents in a valid cluster, the subject cluster is determined to be valid.
[0025] According to the target document and the candidate documents, a document having the same blood relationship with the target document is obtained by a statistical distance method, specifically:
[0026] Obtaining multiple statistical distances based on the embedding vector of the target document and the embedding vectors of the multiple candidate documents by using a statistical distance method;
[0027] The plurality of candidate documents are sorted according to the statistical distance, and when the statistical distance is greater than a preset distance threshold, the corresponding candidate document is deleted to obtain a document having the same blood relationship with the target document.
[0028] Through the similarity of probability distribution between topic clusters, adjacent topic clusters are obtained, specifically:
[0029] According to each topic cluster, the probability distribution of the topic cluster is obtained through averaging processing;
[0030] Obtaining divergence values between the probability distributions of the plurality of topic clusters and the probability distribution of the target topic cluster by using an information entropy method;
[0031] According to the divergence value sorting, when the divergence value is less than a preset divergence value threshold, the corresponding topic cluster is determined to be an adjacent topic cluster.
[0032] The information entropy method is specifically as follows:
[0033] The information entropy method is the JS divergence method.
[0034]
[0035] Where p is the probability distribution of the target topic cluster, q is the probability distribution of other topic clusters, and KL represents the KL divergence.
[0036] The present invention also provides a storage medium,
[0037] The storage medium stores a computer program, which, when executed, implements the steps of any one of the document kinship analysis methods based on the BERT model.
[0038] The present invention again provides a processing device, comprising:
[0039] memory for storing computer programs;
[0040] A processor is used to implement the steps of any one of the document lineage relationship analysis methods based on the BERT model when executing the computer program.
[0041] Due to the adoption of the above technical solution, the beneficial effects achieved by the present invention are as follows:
[0042] 1. In the present invention, based on a document set containing multiple documents, a BERT model is used to obtain an embedding vector set containing multiple embedding vectors. Generating document embedding vectors through the BERT model can capture the contextual semantics of the text (such as the opposition between "dislike" and "like"), and solve the semantic loss problem caused by the traditional LDA topic model based on bag-of-words statistics. The same word generates different embedding vectors in different contexts (such as the difference between "apple" in fruit and technology scenarios), avoiding the semantic confusion caused by static word vectors in traditional methods.
[0043] Dimensionality reduction is performed to reduce the dimension of the embedding vector. By controlling the neighborhood size and distance metric, the high-dimensional embedding vector is compressed to 5 dimensions while preserving the semantic differences of the documents, reducing the computational effort of subsequent clustering and distance calculations.
[0044] A hierarchical density clustering algorithm is used to obtain target topic clusters containing multiple candidate documents. This algorithm does not rely on a preset number of clusters and automatically adapts to density differences in document topic distribution (e.g., sparse versus dense regions), avoiding the overfitting or underfitting issues often encountered by algorithms like K-Means due to a fixed number of clusters. It strikes a balance between local details (e.g., document revisions) and global structure (e.g., cross-topic associations), ensuring that the target topic clusters encompass both fine-grained version iteration relationships and the ability to distinguish between different topic clusters.
[0045] Based on hierarchical density clustering, a multi-stage screening mechanism is implemented, using statistical distance methods to achieve refined screening. This method combines the similarity of embedding vectors with the matching degree of topic probability distribution to verify kinship from both local (document content) and global (topic structure) dimensions. Furthermore, when initial screening results are insufficient, adjacent clusters with similar topic probability distributions are merged based on similarity, expanding the search scope and avoiding omissions caused by blurred topic boundaries.
[0046] In summary, this invention systematically solves the core problems of semantic loss, parameter rigidity and clustering inflexibility in the existing technology through the semantic modeling capability of the BERT model, the adaptive clustering strategy of the hierarchical density clustering algorithm and the multi-stage screening mechanism, thereby improving the accuracy, recall rate and computational efficiency of document kinship mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0048] Figure 1 The figure is a flowchart of the document lineage relationship analysis method based on the BERT model under one embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to more clearly illustrate the overall concept of the present invention, a detailed description is given below in an exemplary manner in conjunction with the accompanying drawings.
[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0051] like Figure 1 As shown in FIG, a document lineage relationship analysis method based on the BERT model includes:
[0052] S100: Based on a document set including multiple documents, an embedding vector set including multiple embedding vectors is obtained through a BERT model, and dimensionality reduction processing is performed to reduce the dimensions of the embedding vectors.
[0053] The core goal of this step is to convert the original document collection into an embedded vector representation and perform dimensionality reduction on this basis to reduce computational complexity, retain key semantic information, and provide high-quality input data for subsequent clustering and kinship analysis.
[0054] Embedding vectors are generated through the BERT model, using a lightweight model based on the BERT architecture (such as paraphrase-multilingual-MiniLM-L12-v2). This model significantly reduces computing resource consumption (384-dimensional output) while maintaining BERT's contextual awareness capabilities, making it suitable for multilingual scenarios (supporting more than 50 languages).
[0055] The document is represented by d, and the document set can be represented as D = {d1, d2, ..., d n}, where n represents the number of documents, d1 to d n Represent the documents in the document collection. Each document (d) is divided into sentences or paragraphs, input into the BERT model, and the final document embedding vector (e) is generated. Repeat the above process for all documents in the document collection to generate the embedding vector set E = {e1, e2, ..., e n}.
[0056] BERT's contextual awareness addresses the shortcomings of traditional LDA models. For example, LDA might classify "dislike" and "like" as the same topic, but BERT, through its self-attention mechanism, can distinguish the opposing semantics of the two in different contexts, thereby improving the semantic accuracy of the embedding vector.
[0057] The UMAP algorithm is used for nonlinear dimensionality reduction to obtain the vector set V={V1,V2,…,V n Compared with linear methods such as PCA, UMAP is better at preserving the local and global structures of high-dimensional data.
[0058] UMAP preserves local neighborhood relationships (such as the closeness of revised documents) through manifold learning while maintaining global topic distribution (such as the separation of different topic clusters). This allows the low-dimensional vectors after dimensionality reduction to reflect subtle differences between documents while supporting the effectiveness of subsequent clustering algorithms.
[0059] Understandably, the computational complexity of high-dimensional vectors (such as clustering and distance calculations) is much higher than that of 5-dimensional vectors. Dimensionality reduction significantly reduces computational overhead, making it suitable for large document collections (e.g., millions of documents). Furthermore, data points are sparsely distributed in high-dimensional space, making it difficult for traditional clustering algorithms to distinguish valid clusters from noise. Dimensionality reduction, through UMAP, preserves local density information, making it easier for density clustering algorithms like HDBSCAN to identify topic clusters.
[0060] This step systematically addresses the shortcomings of traditional methods in semantic modeling and computational efficiency by combining BERT embedding and UMAP dimensionality reduction. BERT's contextual awareness ensures the semantic accuracy of the embedding vectors, while UMAP's parameterized dimensionality reduction strategy further optimizes the quality of the low-dimensional representation, laying the data foundation for subsequent clustering and kinship mining.
[0061] S200: Obtaining a target topic cluster including a plurality of candidate documents through a hierarchical density clustering algorithm according to the embedding vector corresponding to the target document and the embedding vector set.
[0062] The core goal of this step is to automatically identify candidate documents that are semantically similar to the target document from the set of embedding vectors after dimensionality reduction through the hierarchical density clustering algorithm (HDBSCAN), and form a target topic cluster containing documents with potential same-blood relationship.
[0063] It should be noted that the hierarchical density clustering algorithm (HDBSCAN) combines the density clustering idea of DBSCAN and the structural advantage of hierarchical clustering, and reflects the density connectivity of data points by constructing a single-link hierarchical clustering tree.
[0064] According to the embedding vector (Vt) of the target document, that is, the low-dimensional representation after dimensionality reduction (such as 5-dimensional UMAP vector). The embedding vector set of the document set (V={V1,V2,…,V n}), which is a low-dimensional representation of all documents. The target topic cluster (S) is obtained, which contains a set of candidate documents with semantic similarities to the target document. This process not only identifies documents with semantic similarities to the target document but also reveals potential relationships between documents, providing an important basis for subsequent lineage document analysis.
[0065] HDBSCAN identifies clusters based on density differences, effectively handling the multi-density nature of document topic distribution. For example, revised documents under the same topic may be densely distributed (high-density clusters), while documents across different topics may be dispersed (low-density areas). The algorithm automatically distinguishes between the two. For isolated documents (e.g., documents with irrelevant content), HDBSCAN marks them as noise points and excludes them from topic clusters to prevent interference with subsequent kinship determination.
[0066] Understandably, K-Means requires manual setting of the number of clusters (k). Improper settings (e.g., too few or too many) can lead to overfitting or underfitting. HDBSCAN, on the other hand, eliminates the need for a preset number of clusters and dynamically determines the number of clusters based on the data distribution, adapting to the dynamic changes in document topics. Furthermore, document topic distributions may exhibit multi-scale density (e.g., differences in density between long and short document clusters). HDBSCAN adapts to these differences through density-driven clustering, whereas K-Means assumes spherical clusters with uniform density, which can easily lead to incorrect partitioning.
[0067] This step systematically addresses the shortcomings of traditional methods in terms of cluster number presetting and density adaptability through the HDBSCAN algorithm's hierarchical density clustering. The algorithm effectively filters out noise and outliers while preserving document semantic connections, providing a high-quality set of candidate documents for subsequent statistical distance screening. This design not only improves the accuracy of kinship identification but also enhances the method's applicability in complex document distribution scenarios.
[0068] S300: Based on the target document and the candidate documents, a document having the same blood relationship with the target document is obtained by a statistical distance method.
[0069] The core goal of this step is to quantify the semantic similarity between the target document and the candidate documents through statistical distance methods (such as Hellinger distance) and to screen out documents that have the same blood relationship with the target document.
[0070] The Hellinger distance is a statistical method based on the difference in probability distributions and is defined as:
[0071]
[0072] Among them, p and q are the embedding vectors of the target document and the candidate document respectively.
[0073] It should be noted that the Hellinger distance not only focuses on the matching degree of the main peak position, but also captures small deviations in low-probability areas (such as subtle differences in specific terms in the document). For example, when the target document and candidate documents are similar in core content but have individual word frequency differences, the Hellinger distance can more accurately reflect this difference.
[0074] Compared with cosine similarity or Euclidean distance, Hellinger distance is insensitive to vector length and is more suitable for processing the normalized characteristics of document embedding vectors (such as the 5-dimensional vector after UMAP dimensionality reduction).
[0075] In this step, according to the embedding vector V of the target document t (5-dimensional UMAP vector), the embedding vector set of candidate documents V = {V1, V2, ..., V n}(5-dimensional UMAP vector). For each candidate document V i , calculate its value with V t The Hellinger distance H(V t ,V i ). Arrange all candidate documents in ascending order of Hellinger distance (the smaller the distance, the higher the similarity).
[0076] Understandably, HDBSCAN clustering has identified candidate documents with similar semantics, but there may be noisy or marginal documents. Hellinger distance further verifies the similarity of these documents, reducing misclassifications due to the density-driven nature of clustering algorithms (such as misclassifying documents in densely populated areas across topics). Candidate documents that haven't been filtered by statistical distance may contain documents with similar content but significant differences in form (such as different versions of the same topic). Statistical distance filtering ensures that only highly matched documents are retained, avoiding misclassifications of kinship relationships.
[0077] This step systematically addresses the semantic verification of candidate documents after HDBSCAN clustering using the Hellinger distance method. By properly selecting a statistical distance method (Hellinger distance), we can further filter out noisy documents while retaining the clustering results, ensuring the accuracy and practicality of the final kinship relationship.
[0078] S400: When the number of documents with the same blood relationship is less than a preset threshold, adjacent topic clusters are obtained through the similarity of probability distribution between topic clusters, and the adjacent topic clusters are merged into the target topic cluster. The statistical distance method is used again to obtain documents with the same blood relationship as the target document.
[0079] The core goal of this step is to further mine potential blood relationship documents by expanding the search scope (merging adjacent topic clusters) when the number of blood relationship documents initially screened out is insufficient.
[0080] For all documents in each topic cluster, the mean of its topic probability distribution (e.g., the distribution vector generated by the LDA or BERT topic model) is calculated as the representative distribution of the topic cluster. The JS divergence (Jensen-Shannon Divergence) is used to measure the distribution similarity between the target topic cluster and other topic clusters.
[0081] It should be noted that JS divergence is more sensitive to differences in high-probability areas and can effectively capture the main peak position shift of topic distribution (for example, the similarity between the target topic cluster and the adjacent clusters in the core topic), avoiding misjudgment due to slight differences in low-probability areas.
[0082] By dynamically identifying adjacent clusters that are highly similar to the target topic cluster using JS divergence, we address the omission problem often encountered by traditional methods due to blurred topic boundaries. For example, in documents with similar content across different topics (such as the "machine learning" and "deep learning" topic clusters), the merge operation can capture potential kinship relationships.
[0083] Set a JS divergence threshold (e.g., 0.3). When the JS divergence between other topic clusters and the target topic cluster is less than the threshold, they are considered adjacent topic clusters. The divergence threshold is determined based on recall accuracy. For high recall scenarios, if more potential lineage documents need to be covered (such as historical document mining), the threshold can be lowered to 0.25. For high precision scenarios, if strict control of false positives is required (such as legal document compliance audits), the threshold can be raised to 0.35.
[0084] Merge documents from adjacent topic clusters into the target topic cluster to form a new set of candidate documents. For example, after the target topic cluster S is merged with adjacent clusters C1 and C2, the candidate document set is expanded to S∪C1∪C2. In the merged candidate document set, the Hellinger distance between the target document and each candidate document is calculated again, and a threshold (such as 0.2) is set for filtering. Merging adjacent topic clusters may introduce documents that partially overlap with the target topic cluster but do not fully match it. Secondary filtering uses the Hellinger distance to ensure that only highly matching documents are retained in the final result, avoiding noise interference caused by the merge operation.
[0085] This step systematically addresses the issue of insufficient kinship documents by calculating the similarity of probability distributions between topic clusters (JS divergence) and performing secondary statistical distance screening (Hellinger distance). This method expands the search scope while retaining the core documents of the target topic cluster, significantly improving the recall rate of kinship mining while maintaining a high precision through secondary screening.
[0086] As a preferred embodiment of the present invention, based on a document set containing multiple documents, an embedding vector set containing multiple embedding vectors is obtained through a BERT model, specifically:
[0087] According to the document, the query vector, key vector, and value vector are obtained one by one through the BERT model;
[0088] The embedding vector is obtained according to the query vector, the key vector, and the value vector.
[0089] The core goal of this implementation is to generate high-quality document embedding vectors through the BERT model, providing semantically rich embedding vector representations for subsequent kinship analysis.
[0090] The document is represented by d, and the document set can be represented as D = {d1, d2, ..., d n}, where n represents the number of documents, d1 to d n Represents a document in a document collection.
[0091] For each document d in the document set, the embedding model based on the BERT architecture generates a corresponding embedding vector e to generate the embedding vector set E of all documents in the document set D, E = {e1, e2, ..., e n}, where e1 to e n Corresponding to documents d1 to d n The embedding vector of .
[0092] The BERT model used is the pre-trained paraphrase-multilingual-MiniLM-L12-v2 model from the Sentence Transformers library. This model maps sentences and paragraphs in a document into a 384-dimensional vector. This not only enables the model to excel in tasks such as clustering and semantic search, but also ensures its applicability across multiple languages, including but not limited to English, Chinese, French, and over 50 others. With its lightweight MiniLM architecture, the BERT model enables rapid deployment across different devices while maintaining high performance without sacrificing accuracy or performance.
[0093] Specifically, the BERT model utilizes the Transformer architecture and the Self-Attention mechanism to capture complex semantic relationships within text. In the Self-Attention mechanism, each sentence or word in the input model is first converted into three different vectors: a query vector (Q), a key vector (K), and a value vector (V). The query vector (Q) is used to find associations with other words; the key vector (K) is the object being queried and similarity is calculated with Q; and the value vector (V) contains the actual semantic information, which is weighted and summed using attention weights to generate the final embedding vector.
[0094] Through this mechanism, the BERT model can effectively and selectively incorporate information from the entire sentence into the vector representation of each word, providing a richer and more accurate semantic expression for each word. This design enables the model to better understand the meaning of each word in a sentence within a specific context, thereby improving its performance in document kinship mining tasks.
[0095] The embedding vector is obtained according to the query vector, the key vector, and the value vector, specifically:
[0096] Obtaining a similarity score based on the query vector and the key vector, and obtaining a probability value based on the similarity score through a softmax function;
[0097] The value vector is weighted according to the probability value to obtain the embedding vector.
[0098]
[0099] Among them, d k is the dimension of the key vector, used to scale the dot product result to avoid saturation of the softmax function.
[0100] By calculating the dot product of the query vector (Q) and the key vector (K), QK T Get the similarity score of each word with other words. This process captures the contextual dependencies of words in the document (such as the opposite semantics of "dislike" and "like"). k The similarity score is scaled and the probability value is obtained through the softmax function, that is, These probability values reflect the importance of each word in the document. For example, core terms will have higher weights.
[0101] According to the probability value, the value vector is weighted to obtain the embedding vector e. This step integrates the information of the local word vector into the global document vector, preserving the contextual semantics.
[0102] This implementation systematically addresses the shortcomings of traditional methods in semantic modeling through the BERT model's self-attention mechanism and Softmax normalization. By dynamically generating query vectors, key vectors, and value vectors, and combining them with a multi-head attention mechanism, the algorithm is able to capture global topic structures (such as similarities across documents) while retaining local document details (such as minor differences in revisions). This design not only improves the semantic accuracy of the embedding vector, but also provides a high-quality input foundation for subsequent clustering and distance calculations, significantly improving the accuracy and efficiency of kinship mining.
[0103] As a preferred embodiment of this implementation, dimensionality reduction processing is performed to reduce the dimension of the embedding vector, specifically:
[0104] The dimensionality reduction process reduces the dimension of the embedding vector by controlling the neighborhood size according to the distance between points.
[0105] The core goal of this embodiment is to compress high-dimensional embedding vectors (such as 384 dimensions) into a low-dimensional space (such as 5 dimensions) through a nonlinear dimensionality reduction algorithm (such as UMAP) to solve the "curse of dimensionality" problem and provide high-quality low-dimensional representation for subsequent clustering and statistical distance calculations.
[0106] UMAP is a nonlinear dimensionality reduction method based on manifold learning. It preserves the local and global structure of the data by constructing a local adjacency graph of high-dimensional data and optimizing its low-dimensional embedding. n}, calculate the k nearest neighbors of each point (controlled by the n_neighbors parameter) and construct a high-dimensional adjacency graph. The high-dimensional adjacency graph is mapped to a low-dimensional space, and the adjacency relationship of the low-dimensional embedding is made consistent with the high-dimensional one by optimizing the objective function (minimizing the KL divergence).
[0107] The UMAP algorithm uses n_neighbors = 15 to control the neighborhood size and balance local detail with global distribution. n_components = 5 reduces the dimensionality, preserving sufficient semantic information while reducing computational overhead. Metric = 'cosine' uses cosine similarity to calculate point-to-point distances, which is more suitable for semantic matching of text data.
[0108] Regarding neighborhood size, small neighborhoods focus more on local clustering details, while large neighborhoods emphasize the overall distribution pattern. In one embodiment, a neighborhood size of 15 is used. In document lineage mining, this achieves a good balance between identifying different versions or revisions of the same topic (requiring a certain level of local detail) and clearly distinguishing different topic clusters (avoiding excessive fragmentation). This balance avoids over-segmentation due to a too small neighborhood, and avoids losing important local information due to an overly large neighborhood.
[0109] Cosine similarity is selected as the distance metric between points. In document lineage mining, cosine similarity can effectively identify document versions with similar content but slightly different forms, thereby improving the model's sensitivity to subtle differences between documents.
[0110] The dimension of the output vector after dimensionality reduction is selected to be 5-dimensional, which can not only effectively preserve the semantic similarities and differences between documents, but also avoid the computational burden and sparsity problems caused by too high a dimension.
[0111] This embodiment controls the neighborhood size (n_neighbors = 15) and the dimensionality reduction (n_components = 5), and uses cosine similarity as the distance metric. The UMAP algorithm significantly reduces computational complexity while preserving the semantic relevance of documents, thereby improving the low-dimensional representation quality of the embedding vector.
[0112] As a preferred embodiment of the present invention, the hierarchical density clustering algorithm is specifically:
[0113] The hierarchical density clustering algorithm sets a cluster internal point density requirement, and when the point density obtained by the hierarchical density clustering is greater than or equal to the cluster internal point density requirement, a topic cluster is obtained;
[0114] The hierarchical density clustering algorithm sets a minimum number of documents for a valid cluster. When the number of documents in a subject cluster is greater than or equal to the minimum number of documents in a valid cluster, the subject cluster is determined to be valid.
[0115] The core goal of this implementation is to dynamically identify high-quality topic clusters and filter out noisy documents through two key parameters of the hierarchical density clustering algorithm (HDBSCAN) - the internal point density requirement of the cluster and the minimum number of documents in the valid cluster, thereby improving the accuracy and robustness of document kinship mining.
[0116] The embedding vector V of the target document t And the word embedding vector V of each document in the document collection = {V1, V2, ..., V n After clustering using the Hierarchical Density Scale Clustering (HDBSCAN) algorithm, multiple topic clusters S are obtained. All documents within topic cluster S are considered to be related documents. The topic cluster containing the target document is the target topic cluster, and all documents within the target topic cluster are considered candidate documents. This process not only identifies documents with semantic similarities to the target document but also reveals potential relationships between documents, providing an important basis for subsequent analysis of related documents.
[0117] The Hierarchical Density Based Clustering (HDBSCAN) algorithm is a density-based hierarchical clustering algorithm that combines the core concepts of DBSCAN and hierarchical clustering to provide powerful data exploration capabilities. Unlike traditional clustering algorithms that require a predefined number of clusters, HDBSCAN automatically identifies clusters of varying densities and effectively handles noise points. This means that when faced with complex data distributions, HDBSCAN eliminates the need for manual parameterization to determine the number of clusters; instead, it automatically determines the number based on the characteristics of the data. This flexibility makes HDBSCAN well-suited for text mining tasks, particularly document kinship mining, where the similarities and differences between documents can be highly diverse.
[0118] More specifically, adjusting parameters like min_cluster_size and min_samples in the HDBSCAN algorithm can significantly improve performance in document kinship mining. These parameters not only affect the quality of clustering results but also directly impact the ability to accurately identify kinship relationships between documents.
[0119] The min_samples parameter determines the point density requirement within the cluster, that is, how many other points a point needs to have in its neighborhood in order to be considered a valid cluster member. A larger min_samples value will result in only the densest areas being identified as clusters, thereby filtering out many low-density areas. This helps reduce the impact of noise and makes the final clusters more compact and clear. However, this also means that some actual but sparse topic clusters may be ignored.
[0120] In a specific embodiment, min_samples is set to 5 to better balance the relationship between cluster purity and recall rate. When the point density calculated by HDBSCAN is greater than or equal to min_samples, the region is determined to be a valid topic cluster.
[0121] The min_cluster_size parameter defines the minimum number of documents to be considered a valid cluster. This means that any group containing fewer than min_cluster_size documents will not be considered a separate cluster, but may be treated as noise.
[0122] In a specific embodiment, min_cluster_size is selected as 10 because it directly affects the granularity of the clusters. If the value is set too small, it may result in too many clusters and too fine-grained clusters, making it difficult to identify meaningful document sets; conversely, if it is set too large, some small but important topic clusters may be ignored.
[0123] This implementation sets the point density requirement within the cluster and the minimum number of documents in a valid cluster. HDBSCAN effectively filters out noise and outliers while preserving the semantic association of documents, providing a high-quality candidate document set for subsequent statistical distance screening.
[0124] As a preferred embodiment of the present invention, based on the target document and the candidate documents, a document having the same blood relationship with the target document is obtained by a statistical distance method, specifically:
[0125] Obtaining multiple statistical distances based on the embedding vector of the target document and the embedding vectors of the multiple candidate documents by using a statistical distance method;
[0126] The plurality of candidate documents are sorted according to the statistical distance, and when the statistical distance is greater than a preset distance threshold, the corresponding candidate document is deleted to obtain a document having the same blood relationship with the target document.
[0127] The core goal of this embodiment is to quantify the semantic similarity between the target document and candidate documents through statistical distance methods (such as Hellinger distance) and screen out documents that have the same blood relationship with the target document.
[0128] As a method for measuring the difference between two probability distributions, the Hellinger distance is based on geometric distance and can effectively reflect the shape differences between the distributions. It is particularly sensitive to changes in low-probability areas. This means that even when the main peaks of the two distributions are the same but there is a slight shift, the Hellinger distance can more accurately capture this subtle difference. Specifically, the Hellinger distance can be calculated using the following expression:
[0129]
[0130] Among them, p and q represent the embedding vectors of the candidate document and the target document respectively.
[0131] Next, the candidate documents are sorted in ascending order by their Hellinger distances to the target document, and a threshold (e.g. 0.2) is set to remove documents with a Hellinger distance greater than this threshold from the candidate documents. This process essentially filters out documents that are closer to the target document, meaning they not only have similar topic structures at the macro level but also show a high degree of consistency in details. The remaining candidate documents are considered to be similar to the target document d t A collection of documents with the same lineage.
[0132] HDBSCAN clustering has identified candidate documents with similar semantics, but there may be noisy or marginal documents. Hellinger distance further verifies the similarity of these documents, reducing false positives caused by the density-driven nature of clustering algorithms (such as misclassifying documents in densely populated areas across topics). Unfiltered candidate documents may include documents with similar content but significant differences in form (e.g., different versions of the same topic). Threshold filtering ensures that only highly matched documents are retained, avoiding false positives in consanguinity.
[0133] This implementation systematically addresses the semantic verification problem of candidate documents after HDBSCAN clustering using the Hellinger distance method. By properly selecting the statistical distance method (Hellinger distance) and threshold (T = 0.2), the algorithm can further filter out noisy documents while retaining the clustering results, ensuring the accuracy and practicality of the final kinship relationship.
[0134] As a preferred embodiment of the present invention, adjacent topic clusters are obtained by similarity of probability distribution between topic clusters, specifically:
[0135] According to each topic cluster, the probability distribution of the topic cluster is obtained through averaging processing;
[0136] Obtaining divergence values between the probability distributions of the plurality of topic clusters and the probability distribution of the target topic cluster by using an information entropy method;
[0137] According to the divergence value sorting, when the divergence value is less than a preset divergence value threshold, the corresponding topic cluster is determined to be an adjacent topic cluster.
[0138] The core goal of this implementation is to identify adjacent topic clusters that are highly relevant to the target topic cluster through the similarity of probability distribution between topic clusters, thereby expanding the search scope of bloodline documents and improving the recall rate of bloodline relationship mining.
[0139] If the number of documents with the same blood relationship is less than a certain threshold (e.g. 5), it indicates that the current target topic cluster S may not cover enough documents related to the target document d t Similar or related documents. To ensure that potential document relationships are fully explored and to expand the scope of candidate documents, adjacent topic clusters can be retrieved. The so-called "adjacent" refers to those topic clusters that are close to the target topic cluster S in terms of topic probability distribution. This means that although they are not exactly the same, they are highly similar in some aspects.
[0140] First, we need to calculate the topic probability distribution for each topic cluster. Specifically, we can average the topic probability distributions of all documents within each topic cluster and use this as the topic probability distribution of the topic cluster. This not only effectively summarizes the core characteristics of a topic cluster, but also, to a certain extent, smooths out the abnormal effects that may be caused by individual documents. Specifically, the topic probability mean of a topic cluster can be calculated using the following expression:
[0141]
[0142] Among them, m represents the number of documents in each topic cluster, p i Represents the topic probability distribution of each document in the topic cluster.
[0143] The information entropy method is specifically as follows:
[0144] According to the information entropy method, it is the JS divergence method.
[0145]
[0146] Where p is the probability distribution of the target topic cluster, q is the probability distribution of other topic clusters, and KL represents the KL divergence.
[0147] JS divergence is used to measure the similarity of topic probability distributions between different topic clusters. JS divergence focuses on the difference in information between two probability distributions and is more sensitive to changes in high-probability regions. This means that if the main peaks of two distributions are located at different positions, even if they have a high degree of overlap in low-probability regions, the JS divergence will be significantly higher, reflecting a larger topic difference between the two. The smaller the JS divergence, the more similar the topic distributions of other topic clusters are to the target topic cluster S.
[0148] After the JS divergence is calculated, these divergence values are sorted in ascending order. Then, based on the set threshold (for example, 0.3), all documents in the topic clusters whose JS divergence is less than this threshold are merged. This process essentially expands the range of candidate documents to include not only the documents in the initially determined target topic cluster, but also those in adjacent topic clusters that have a high degree of topic similarity with the target topic cluster. Doing so helps ensure that as many documents as possible are captured that are similar to the target document. t Related documentation.
[0149] This implementation systematically addresses the issue of insufficient kinship documents by averaging the probability distribution between topic clusters, calculating information entropy and JS divergence, and sorting and thresholding divergence values. By properly setting the JS divergence threshold (0.3) and dynamically adjusting the strategy, the algorithm expands the search scope while retaining the core documents of the target topic cluster, significantly improving the recall rate of kinship mining while maintaining a high precision rate through secondary screening.
[0150] The present invention also provides a storage medium,
[0151] The storage medium stores a computer program, which, when executed, implements the steps of any one of the document kinship analysis methods based on the BERT model.
[0152] Therefore, any effect of the document lineage relationship analysis method based on the BERT model can be achieved, and I will not go into details here.
[0153] The present invention again provides a processing device, comprising:
[0154] memory for storing computer programs;
[0155] A processor is used to implement the steps of any one of the document lineage relationship analysis methods based on the BERT model when executing the computer program.
[0156] Therefore, any effect of the document lineage relationship analysis method based on the BERT model can be achieved, and I will not go into details here.
[0157] Anything not described in the present invention can be achieved by adopting or drawing on existing technologies.
[0158] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0159] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A document kinship analysis method based on the BERT model, characterized in that: According to a document set including multiple documents, an embedding vector set including multiple embedding vectors is obtained by using a BERT model, and a dimensionality reduction process is performed to reduce the dimensions of the embedding vectors; According to the embedding vector corresponding to the target document and the embedding vector set, a target topic cluster containing multiple candidate documents is obtained through a hierarchical density clustering algorithm; Obtaining documents having the same blood relationship with the target document by using a statistical distance method based on the target document and the candidate documents; When the number of documents with the same blood relationship is less than a preset threshold, adjacent topic clusters are obtained through the similarity of probability distribution between topic clusters, and the adjacent topic clusters are merged into the target topic cluster. The statistical distance method is then used again to obtain documents with the same blood relationship as the target document.
2. The document kinship analysis method based on the BERT model according to claim 1, characterized in that: According to the document set containing multiple documents, the BERT model is used to obtain an embedding vector set containing multiple embedding vectors, specifically: According to the document, the query vector, key vector, and value vector are obtained one by one through the BERT model; The embedding vector is obtained according to the query vector, the key vector, and the value vector.
3. The document kinship analysis method based on the BERT model according to claim 2, characterized in that: Perform dimensionality reduction to reduce the embedding vector dimension, specifically: The dimensionality reduction process reduces the dimension of the embedding vector by controlling the neighborhood size according to the distance between points.
4. The document kinship analysis method based on the BERT model according to claim 2, characterized in that: The embedding vector is obtained according to the query vector, the key vector, and the value vector, specifically: Obtaining a similarity score based on the query vector and the key vector, and obtaining a probability value based on the similarity score through a softmax function; The value vector is weighted according to the probability value to obtain the embedding vector.
5. The document kinship analysis method based on the BERT model according to claim 1, characterized in that: The hierarchical density clustering algorithm is specifically: The hierarchical density clustering algorithm sets a cluster internal point density requirement, and when the point density obtained by the hierarchical density clustering is greater than or equal to the cluster internal point density requirement, a topic cluster is obtained; The hierarchical density clustering algorithm sets a minimum number of documents for a valid cluster. When the number of documents in a subject cluster is greater than or equal to the minimum number of documents in a valid cluster, the subject cluster is determined to be valid.
6. The document kinship analysis method based on the BERT model according to claim 1, characterized in that: According to the target document and the candidate documents, a document having the same blood relationship with the target document is obtained by a statistical distance method, specifically: Obtaining multiple statistical distances based on the embedding vector of the target document and the embedding vectors of the multiple candidate documents by using a statistical distance method; The plurality of candidate documents are sorted according to the statistical distance, and when the statistical distance is greater than a preset distance threshold, the corresponding candidate document is deleted to obtain a document having the same blood relationship with the target document.
7. The document kinship analysis method based on the BERT model according to claim 1, characterized in that: Through the similarity of probability distribution between topic clusters, adjacent topic clusters are obtained, specifically: According to each topic cluster, the probability distribution of the topic cluster is obtained through averaging processing; Obtaining divergence values between the probability distributions of the plurality of topic clusters and the probability distribution of the target topic cluster by using an information entropy method; According to the divergence value sorting, when the divergence value is less than a preset divergence value threshold, the corresponding topic cluster is determined to be an adjacent topic cluster.
8. The document kinship analysis method based on the BERT model according to claim 7, characterized in that: The information entropy method is specifically as follows: The information entropy method is the JS divergence method. Where p is the probability distribution of the target topic cluster, q is the probability distribution of other topic clusters, and KL represents the KL divergence.
9. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed, implements the steps of the document kinship analysis method based on the BERT model as described in any one of claims 1 to 8.
10. A processing device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the document kinship analysis method based on the BERT model as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Document blood relationship mining method and device based on topic model
CN113032575A
Keyword vectorization method based on topic semantic information and application thereof
CN114298020A
Multi-document scientific abstract generation method based on topic knowledge graph joint enhancement
CN116821371A
Social work field modeling optimization method driven by multi-modal data
CN119939482A