A multi-dimensional analysis method for technical literature based on semantic understanding

By using pre-trained models and clustering algorithms to conduct in-depth semantic understanding and multi-dimensional analysis of technical literature, the shortcomings of deep semantic understanding and multi-dimensional analysis of text in the existing technology are solved, and more accurate technical literature analysis is achieved.

CN119128143BActive Publication Date: 2025-06-20中铁科学研究院集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411033100.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-06-20
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The prior art literature analysis methods have shortcomings in the deep semantic understanding and multi-dimensional analysis of texts, especially when dealing with polysynonyms and relying on large-scale annotation data sets, it is difficult to comprehensively evaluate the value and potential of technical literature.

Method used

Advanced pre-trained models such as BERT or Sentence-BERT are used for text vector processing, combined with UMAP for vector dimensionality reduction, HDBSCAN is used for unsupervised clustering analysis, subject words are extracted through TF-ICF method, and multi-dimensional analysis is performed.

Benefits of technology

It realizes in-depth semantic understanding and comprehensive multi-dimensional analysis of technical literature, reduces dependence on large-scale data sets, and can express text meaning more accurately and reveal technical relationships between patents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119128143B_ABST
    Figure CN119128143B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-dimensional analysis method for technical documents based on semantic understanding. The Sentence-BERT model is used to perform text vectorization processing on the text data to generate a dense vector representation of the text; UMAP is used for vector dimensionality reduction to remove redundant features; HDBSCAN is used for unsupervised clustering analysis to generate clustering results; the TF-ICF method is used to extract topic words from the clustering results; and multi-dimensional analysis is performed on the clustering results. Compared with the prior art, the present invention provides a comprehensive and efficient patent analysis tool through an advanced semantic understanding-based patent clustering method, thereby achieving a deeper interpretation and analysis of technical documents while ensuring cost-effectiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for analyzing technical literature, and particularly to a multi-dimensional analysis method for technical literature based on semantic understanding. Background Art

[0002] Current technical literature analysis techniques widely rely on traditional text processing techniques and algorithms to parse and evaluate patent documents. Although these methods are effective to a certain extent, they have significant limitations. First, when the common word2vec model converts patent text into word vectors, although it can partially capture the associations between words, it performs poorly in dealing with polysemous words, which affects the understanding and mining of the deep semantics of the text. Second, the current analysis methods usually only focus on a single dimension and fail to construct a comprehensive analysis framework, making it difficult to comprehensively evaluate the value and potential of technical literature. This single analysis perspective not only ignores the inherent complexity and richness of the technical literature text but also fails to fully demonstrate the interrelationships between patents and their roles in technological progress.

[0003] More seriously, many advanced analysis techniques based on natural language processing (NLP) require large-scale labeled datasets for effective training to achieve good analysis results. However, constructing such datasets is both time-consuming and laborious, which not only increases the difficulty of analysis but also limits the expansion of analysis dimensions. Therefore, developing a technical literature analysis method that can deeply understand text semantics and perform comprehensive multi-dimensional analysis without relying on large-scale datasets has become a key challenge for improving the quality and efficiency of technical literature analysis. The current solutions on the market have not been able to fully overcome these problems, resulting in limitations in the depth and breadth of technical literature analysis and affecting the accuracy and practicality of analysis results. Summary of the Invention

[0004] The object of the present invention is to provide a multi-dimensional analysis method for technical literature based on semantic understanding. It aims to provide a comprehensive and efficient patent analysis tool through an advanced semantic understanding-based patent clustering method, so as to achieve a deeper interpretation and analysis of technical literature while ensuring cost-effectiveness.

[0005] To achieve the above object, the present invention is implemented according to the following technical solution:

[0006] The present invention includes the following steps:

[0007] S1: Obtain the text data of technical literature, and use the Sentence-BERT model to perform text vectorization processing on the text data to generate a dense vector representation of the text;

[0008] S2: Use UMAP for vector dimensionality reduction to remove redundant features;

[0009] S3: Use HDBSCAN for unsupervised clustering analysis to generate clustering results;

[0010] S4: Adopt the TF-ICF method to extract topic words from the clustering results;

[0011] S5: Conduct multi-dimensional analysis on the clustering results.

[0012] The beneficial effects of the present invention are:

[0013] The present invention is a multi-dimensional analysis method for technical literature based on semantic understanding. Compared with the prior art, the present invention realizes the deep semantic understanding and comprehensive multi-dimensional analysis of patent texts through the following key technical points:

[0014] (1) Application of pre-trained models: The present invention adopts advanced pre-trained models such as BERT or Sentence-Transformers, etc. These models have been pre-trained on large-scale corpora and can better capture the deep semantic information in the text. By mapping the patent abstract text to a vector representation in a high-dimensional space, this method can more accurately express the text meaning, especially for the processing of polysemous words, significantly improving the accuracy of semantic understanding of patent documents.

[0015] (2) Unsupervised text clustering: After converting the patent abstract text into vectors, the present invention uses an efficient clustering algorithm to analyze these vectors and automatically identify the professional technical topics in the patent text. This method not only reduces the dependence on manually constructed data sets but also can discover the implicit patterns and groups in patent literature, thereby revealing the technical associations between patents and providing a more accurate technical field division for users.

[0016] (3) Multi-dimensional keyword recognition and analysis: In addition to traditional text content analysis, the present invention also introduces additional analysis dimensions such as "efficacy", "information technology", etc. By identifying and clustering the keywords in these dimensions, it is possible to further reveal the applications and impacts of patents in different fields. This method broadens the analysis perspective, makes the analysis results more comprehensive, provides users with more diverse insights, and supports more complex decision-making processes. Description of the Drawings

[0017] Figure 1 is the flow chart of the clustering method of the present invention;

[0018] Figure 2 is the SBERT classification objective function;

[0019] Figure 3 is the flow chart of the multi-dimensional analysis method. Detailed Embodiments

[0020] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions of this invention are used to explain the present invention, but do not limit the present invention.

[0021] As Figure 1 shown: The present invention includes the following steps:

[0022] S1: Obtain the text data of technical literature, and use the Sentence-BERT model to perform text vectorization processing on the text data to generate a dense vector representation of the text;

[0023] S2: Use UMAP for vector dimensionality reduction to remove redundant features;

[0024] S3: Use HDBSCAN for unsupervised clustering analysis to generate clustering results;

[0025] S4: Use the TF-ICF method to extract topic words from the clustering results;

[0026] S5: Perform multi-dimensional analysis on the clustering results.

[0027] The present invention uses the Sentence-BERT pre-trained model to process the abstract text of technical literature to generate a dense vector representation of the text. Sentence-BERT is a BERT variant specifically designed for sentences. It can better capture the semantic similarity within sentences through pre-training on a large-scale corpus. In this method, Sentence-BERT is used to extract the key semantic information in the abstract of technical literature and generate vector representations for each sentence. These vectors can accurately reflect the deep semantic structure of the sentences and provide a solid foundation for subsequent analysis.

[0028] SBERT uses a Siamese and Triplet Network to update the weight parameters to achieve that the generated sentence vectors have semantic information. By using the pre-trained model, the SBERT model can obtain discourse vectors that are semantically meaningful enough. SBERT adds a Pooling operation to the output of the last layer of BERT / RoBERTa, mainly using the Mean-pooling strategy, calculating the average value of each Token output vector to represent the sentence vector, thereby generating a fixed-size sentence Embedding vector.

[0029] As Figure 2As shown, SBERT follows the structure of the siamese network. The Encoder for the input sentences uses the same BERT to process, enabling BERT to better capture the relationships between sentences and obtain sentence vectors with semantic features. When SBERT processes text classification, given input sentences a and b, after passing through BERT and the Pooling operation, sentence vectors S a and S b are obtained. The sentence vectors S a and S b along with their difference vector S a -S b are concatenated together to form a new feature vector, which is then multiplied by the trainable weight matrix Wt, i.e.:

[0030] v = softmax(W t (s a , s b |s a -s b |))

[0031] where W t ∈ R 3d*t , d is the dimension of the sentence vector, t is the number of classification labels, and v is a probability distribution vector.

[0032] When two sentences S1 and S2 in the technical literature text are input into SEBRT, the vector representation is:[[]]

[0033]

[0034] where n represents the size of the number of hidden layer units inside BERT.

[0035] UMAP Dimensionality Reduction:

[0036] The Sentence - BERT language model first performs text pre - training using BERT. Since the Chinese model of BERT usually adopts a length limit of 512 characters, the pre - trained technical literature will become a vector matrix of N * 512 (N is the number of technical literature). As N increases, a high - dimensional data set will be formed. To remove redundant features and improve the clustering effect of the text, data dimensionality reduction operation needs to be performed on the vector matrix. For this reason, the present invention proposes a method of using UMAP (Uniform Manifold Approximation and Projection) for dimensionality reduction.

[0037] UMAP is an effective dimensionality reduction algorithm that can project high-dimensional Sentence-BERT vectors into a low-dimensional space while preserving the topological structure of the data. Its theoretical basis is Riemannian geometry and algebraic topology. It mainly uses local manifold approximation and local fuzzy simplicial set representation to construct the topological representation of high-dimensional data. That is, for high-dimensional data, given the low-dimensional representation of some data, a similar process can be used to construct an equivalent low-dimensional topological representation. Currently, UMAP is one of the best methods for text vector dimensionality reduction. In the process of data dimensionality reduction, using the UMAP method can not only reduce the computational complexity and memory usage, but also retain the features of the original data to the greatest extent.

[0038] UMAP dimensionality reduction simplifies the data structure, facilitating subsequent visualization and analysis, and also prepares the input data for the clustering algorithm.

[0039] The vectors after UMAP dimensionality reduction are then immediately input into the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm for clustering. HDBSCAN is a clustering algorithm that does not require pre-setting parameters. It can automatically identify natural clusters in the data and handle noise and outliers. The biggest difference from the traditional DBSCAN is that HDBSCAN can handle clustering problems of clusters with different densities and shows more robust advantages in parameter selection. The HDBSCAN algorithm introduces the idea of hierarchical clustering and places restrictions on the minimum subtrees obtained by pruning the minimum spanning tree, controlling the generated clusters from being too small. In addition, the algorithm is less sensitive to parameters and does not require setting the threshold by itself, only the minimum number of clusters needs to be defined.

[0040] HDBSCAN can evaluate the membership degree for each sample, and the membership degree ranges from [0, 1]. If the membership degree value is 0, it indicates that the sample point is a noise point and does not belong to any cluster; if the membership degree value is 1, it indicates that the sample point is the center point, and the cluster center point attribute can represent the typical features of the cluster. The calculation steps of the algorithm are as follows:

[0041] (1) Spatial transformation of data points: The traditional single-link clustering method clusters according to the minimum distance between data points, and it is easily affected by noise. HDBSCAN uses the mutual reachability distance instead of the minimum distance between data points, which can strengthen the close connection between data points in high-density regions and weaken the association between data points in low-density regions. The calculation formula for the mutual reachability distance between two data points a and b is:

[0042]

[0043] Among them, The neighborhood with a range of K for point a is represented, and d(a, b) refers to the distance between points a and b. It represents the core distance of data point a, that is, the distance between a and its k-th nearest data point.

[0044] When calculating the distance, the Euclidean distance is calculated as:

[0045]

[0046] In the formula, a i , b i are the coordinates of points a and b in the n-dimensional space.

[0047] To improve the efficiency of calculating the core distance, locality-sensitive hashing is used to optimize the search efficiency. The locality-sensitive hashing formula is:

[0048]

[0049] Among them, a is the vector representation of the data point, x is a random vector, y is a random offset, and w is the width of the hash bucket. represents rounding down.

[0050] (2) Construction of the minimum spanning tree. The HDBSCAN algorithm regards the data as a weighted graph, where the data points are vertices, and the weight of the edge between any two points is equal to the mutual reachability distance between these points. The Prim (Prim's algorithm) is used to construct the minimum spanning tree in the weighted graph.

[0051] (3) Establishment of the cluster hierarchy. The algorithm traverses and reorders the edges of the minimum spanning tree with the mutual reachability distance as the weight, and classifies each edge into a new cluster.

[0052] (4) Compression of the clustering tree. When the algorithm compresses and partitions the coarse hierarchy, it compares the number of samples in the newly split cluster with the number of samples in the smallest cluster, and eliminates the smaller one. The size of the smallest cluster is a parameter of HDBSCAN and will be adjusted according to the number of samples in the technical literature.

[0053] (5) Extraction of clustering clusters. HDBSCAN defines a clustering selection method based on stability to identify the clusters with high stability in the dataset. The stability σ c is used to evaluate the stability of the clustering clusters. The calculation formula of σ c is

[0054]

[0055] where x is the data point, c is the clustering cluster, the λ value is the reciprocal of the distance, and λ birth is used to represent the λ value when the clustering cluster is formed, λx represents the λ value separated from the parent cluster.

[0056] In the present invention, the HDBSCAN algorithm is used to identify professional technical topics in technical literature texts, thereby forming meaningful categories that represent potential technical fields and trends in technical literature.

[0057] Improved TF-IDF for extracting topic words:

[0058] After clustering, topic words are extracted from the obtained clusters. The present invention adopts the TF-ICF algorithm of improved TF-IDF (Term Frequency-Inverse Document Frequency). TF-IDF is a statistical method that can judge the importance degree of a word according to its frequency of occurrence in a corpus. The TF-IDF algorithm believes that if a word appears frequently in an article and rarely appears in other articles, then this word has good category discrimination ability. The formula of TF-IDF is TF * IDF, where TF represents the frequency of a word in a document, and the higher the frequency, the more times the word appears; IDF represents the inverse document frequency of a word, and the fewer documents containing this word, the more the word can reflect the theme of the document. The specific calculation formulas of TF and IDF are as follows:

[0059]

[0060] where f k,Di is the frequency of keyword k in document Di, and |Di| is the total number of all words in the i-th document.

[0061]

[0062] where N D is the total number of documents, and I indicates whether keyword k appears in document Di. The more common a word is, the lower its IDF value.

[0063] Based on the TF-IDF algorithm, the present invention regards all documents in a topic cluster as a single document C. Therefore, the inverse document frequency is replaced by the inverse class frequency, and the importance score algorithm TF-ICF for words in a topic cluster can be obtained:

[0064]

[0065] where c is a single cluster document merged from all documents belonging to the same topic cluster, f k,c is the frequency of keyword k in cluster c, |c| is the total number of all words in cluster c, A C is the average number of words in all clusters c containing this keyword, ∑ c∈C f k,c is the frequency of keyword k in all clusters c. Finally, to avoid negative output, one is added to the result of the logarithmic operation.

[0066] In this way, the TF-ICF process simulates the importance of words in clusters rather than individual documents, thus generating a topic word distribution for each cluster. Therefore, this method can be used to mine the topic words in each cluster group, so as to describe and characterize different topics.

[0067] Multi-dimensional analysis method for technical literature:

[0068] In addition to the traditional clustering analysis of the text content of technical literature, the present invention also introduces additional analysis dimensions, such as "efficacy", "information technology", etc. By identifying and clustering the keywords of these dimensions, the application and influence of technical literature in different fields can be further revealed.

[0069] Its core technical process is as Figure 3 shown;

[0070] 1. Data preparation: High-quality technical literature data is the premise to ensure the orderly and correct development of research work. Therefore, the present invention first constructs a retrieval formula to screen technical literature data, and improves the data quality through data preprocessing methods to obtain technical literature data available for analysis.

[0071] 2. Professional technical topic clustering: Adopt a technical literature clustering method based on semantic understanding to perform clustering analysis on the abstract text of technical literature to obtain the clustering results of professional technical topics.

[0072] 3. Other dimension topic clustering:

[0073] (1) Use the Chinese word segmentation library jieba to segment the abstract text of technical literature, count the word frequencies, and sort them in descending order according to the word frequencies.

[0074] (2) Empirically pre-select representative keywords in other dimensions. For example, from the analysis of the "information technology" dimension, select information technology vocabulary such as "computer, machine learning, neural network, algorithm", etc.; then use Sentence-BERT to vectorize the segmented words of technical literature and the keywords. Then use the cosine similarity algorithm to calculate the similarity of the text vectors, and find the words with semantic relationships with the information technology words from the segmented words of technical literature. The cosine similarity algorithm is:

[0075]

[0076] where a i , b i are two text vectors in the n-dimensional space, and represent the magnitudes of vectors a and b respectively.

[0077] This step needs to be repeated several times to evaluate the appropriate initial representative words and cosine similarity threshold.

[0078] (3) After identifying the keywords of the technical literature text in this dimension, use a technical literature clustering method based on semantic understanding to cluster the keywords to obtain the theme in this dimension and the keywords included in this theme.

[0079] (4) Use a rule algorithm to find the keywords of this dimension from the technical literature text. If the technical literature contains the keywords of a theme, the support degree of the technical literature belonging to this theme increases. Finally, select the theme with the highest support degree as the theme of the technical literature in this dimension.

[0080] (5) Repeat steps (2)-(4) to obtain more theme information in other dimensions to support multi-dimensional analysis.

[0081] The technical solution of the present invention is not limited to the limitations of the above specific embodiments. Any technical deformation made according to the technical solution of the present invention falls within the protection scope of the present invention.

Claims

1. A multi-dimensional analysis method of technical documents based on semantic understanding, characterized in that: The following steps are involved: S1: Obtain text data of technical documents, use the Sentence-BERT model to perform text vectorization on the text data, and generate a dense vector representation of the text; The Sentence-BERT model adopts a twin network structure. The encoder of the input sentence is processed by the same BERT. When SBERT processes text classification, it inputs sentence a and sentence b. After BERT and Pooling operations, the sentence vector S can be obtained. a and S b , the sentence vector S a and S b and the difference vector S between them a -S b Spliced ​​together to form a new feature vector, and then multiplied by the trainable weight matrix W t ,Right now: v=softmax(W t (s a ,s b |s a -s b |)) Among them, W t ∈R 3d*t , d is the dimension of the sentence vector, t is the number of classification labels, and v is a probability distribution vector; When two sentences S1 and S2 in the text data are input into SEBRT, the vector is represented as: Among them, n represents the number of hidden layer units inside BERT; S2: Use UMAP to reduce the dimension of vectors and remove redundant features; S3: Use HDBSCAN to perform unsupervised cluster analysis and generate clustering results; S31: Spatial transformation of data points: a and b represent the mutual distance between two data points. The calculation formula is: in, Indicates that the range of point a is the domain of K, and d(a,b) refers to the distance between points a and b. Represents the core distance of data point a, that is, the distance between a and its kth nearest data point; When calculating the distance, the Euclidean distance is calculated as: Where a i ,b i are the coordinates of points a and b in n-dimensional space; Local sensitive hashing is used to optimize search efficiency. The local sensitive hashing formula is: Where a is the vector representation of the data point, x is a random vector, y is a random offset, and w is the width of the hash bucket. Indicates rounding down; S32: Construction of minimum spanning tree: HDBSCAN algorithm regards data as a weighted graph, where data points are vertices and the weight of the edge between any two points is equal to the mutual reachable distance between these points; Prim's algorithm is used to construct the minimum spanning tree in the weighted graph; S33: Establishment of cluster hierarchy: Using mutual reach distance as weight, traverse and reorder the edges of the minimum spanning tree, and classify each edge into a new cluster; S34: Compression of clustering tree: When compressing and segmenting the coarse hierarchical structure, the number of samples of the newly segmented cluster is compared with the number of samples of the smallest cluster, and the smaller one is eliminated. The size of the smallest cluster is a parameter of HDBSCAN and is adjusted according to the number of technical literature samples; S35: Cluster extraction: Identify clusters with higher stability in the data set and use stability σ c To evaluate the stability of clustering; σ c The calculation formula is: Where x is the data point, c is the cluster, and λ is the inverse of the distance. birth To represent the λ value when the cluster is formed, λ x represents the λ value separated from the parent cluster; S4: TF-ICF method was used to extract keywords from clustering results; S5: Perform multi-dimensional analysis on the clustering results.

2. The multi-dimensional analysis method of technical documents based on semantic understanding according to claim 1 is characterized by: In step S2, UMAP performs vector dimensionality reduction by using local manifold approximation and local fuzzy simplex set representation to construct a topological representation of high-dimensional data, that is, a low-dimensional representation of some data is given to the high-dimensional data, and is used to construct an equivalent low-dimensional topological representation.

3. The multi-dimensional analysis method of technical documents based on semantic understanding according to claim 1 is characterized by: The TF-ICF method formula in step S4 is TF*ICF, where TF represents the frequency of a word appearing in a document, and ICF represents the inverse cluster frequency of a word; the specific calculation formula of TF-ICF is as follows: Where c is a single cluster document formed by merging all documents belonging to the same topic cluster, and f k,c is the frequency of keyword k in cluster c, |c| is the total number of all words in cluster c, A C is the average number of words in all clusters c containing the keyword, ∑ c∈C f k,c is the frequency of keyword k in all clusters C; the logarithmic operation result is added by one.

4. The multi-dimensional analysis method of technical documents based on semantic understanding according to claim 3 is characterized by: The multi-dimensional analysis in step S5 includes professional and technical theme clustering and other dimensional theme clustering; firstly, a search-based screening of technical literature data is constructed, and the data quality is improved by a data preprocessing method to obtain technical literature data available for analysis; the professional and technical theme clustering adopts a clustering method based on semantic understanding to perform cluster analysis on the abstract text of the technical literature to obtain professional and technical theme clustering results; The other dimension topic clustering uses the Chinese word segmentation library jieba to segment the abstract text of the technical literature, count the word frequency, and sort it in descending order by word frequency. Then, Sentence-BERT is used to vectorize the technical literature word segmentation and keywords, and the cosine similarity algorithm is used to calculate the similarity of the text vectors. Words with semantic relationship with information technology words are found from the technical literature word segmentation; after identifying the keywords of the technical literature text in this dimension, a clustering method based on semantic understanding is used to cluster the keywords to obtain the theme under this dimension and the keywords contained in this theme; a rule algorithm is used to find the keywords of this dimension from the technical literature text. If the technical literature contains the keywords of a theme, the support of the technical literature belonging to this theme increases. Finally, the theme with the highest support is selected as the theme of the technical literature in this dimension.

Citation Information

Patent Citations

  • Literature generality analysis method and device

    CN116304016A

  • Document duplicate checking method based on hierarchical feature vector search

    CN117951256A