A patent tree and a method of constructing the same
By analyzing patent text data and generating a patent tree structure, the problem of existing technologies being unable to reflect technological evolution and identify key nodes is solved, enabling precise positioning of patent technology elements and optimized resource allocation.
Patent Information
- Application Number
- CN202411756444.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Existing patent management technologies are unable to effectively reflect the technological evolution process and identify key technological nodes, resulting in low resource allocation efficiency and difficulty in optimizing patent layout and strategic planning.
By analyzing patent text data, extracting technical features and classifying and integrating them, and combining time nodes and life cycle stages, a patent tree structure is generated, and the patent tree layout is optimized to reflect the state of technological development.
It enables precise positioning of patent technology elements and cross-domain feature combination analysis, optimizes the hierarchical structure of the patent tree, helps enterprises clearly identify their technology layout, and improves resource allocation efficiency.
Smart Images

Figure CN119808713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of patent data management, and in particular to a patent tree and a construction method thereof. Background Art
[0002] Patent data management technology utilizes systematic data management tools and methods to efficiently organize, store, retrieve, and analyze patent data to support strategic intellectual property management. These technologies typically encompass data mining, patent classification, trend analysis, and patent portfolio management, helping businesses and research institutions unlock the value of patent data, gain in-depth insights into technological intelligence, and optimize decision-making. Patent data management technology plays a vital role in intellectual property protection, R&D planning, technology early warning, and competitive analysis.
[0003] The patent tree construction method is a tool focused on systematically classifying and displaying a company's patent assets. It uses a tree-like structure to hierarchically categorize patents by technical themes, helping companies quickly identify key technology R&D priorities and patent protection gaps. This method is used for intellectual property management, supporting companies in effectively managing patent assets, strategically planning technological routes, and optimizing the layout for high-quality patent development. By constructing a patent tree, companies can gain a clear perspective on technological development, rationally allocate R&D resources, and improve their patent portfolio.
[0004] Existing technologies struggle to effectively reflect the evolution of technology in patent management, and exhibit a significant lag in information changes regarding patent citation growth and technology expansion rates. Existing technologies also struggle to quickly identify key technology nodes during changes in patent lifecycle stages, making it difficult for companies to prioritize patent resources with high development potential in technology planning and resulting in inefficient resource allocation. Furthermore, existing patent model presentations struggle to systematically grasp the inherent hierarchical relationships within the patent structure, increasing the difficulty of technical route planning and strategic planning, hindering companies' refined operations in patent management and R&D deployment. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a patent tree and a construction method thereof.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a patent tree and a method for constructing the same, comprising the following steps:
[0007] S1: Based on patent text data, analyze the technical features in the patent text, classify and integrate the technical features of each patent, and obtain the patent technical feature set;
[0008] S2: Based on the patent technology feature set, the patent data is sorted and integrated by time node, and the data fluctuations at the time node are combined to capture the historical dependency characteristics to obtain the patent technology status prediction results;
[0009] S3: Based on the patent technology status prediction results, the patents are classified according to the life cycle stage. According to the classification results, the path level of each type of patent is divided to generate hierarchical patent path information;
[0010] S4: Based on the patent technical feature data and hierarchical patent path information in the patent technical feature set, repeatedly appearing technical terms and functional descriptions are extracted to determine potential feature item combinations and obtain cross-domain feature combination data;
[0011] S5: Based on the cross-domain feature combination data, the technical features of the patents are converted into vector form, patents with similarity matching are grouped together, the importance and co-occurrence frequency of the technical features are analyzed, and hierarchical nodes are divided to obtain a patent structure layout list;
[0012] S6: Based on the patent structure layout list, assign weights to the patent nodes in the path, optimize and adjust the patent tree layout, and generate a patent tree structure with optimized sorting.
[0013] As a further solution of the present invention, the patent technical feature set includes a set of technical terms of the patent, a functional description structure and an application scenario classification; the patent technical status prediction result includes a change curve of patent citations, an expansion speed index and a citation count change rate; the hierarchical patent path information includes the position of the patent in each stage in the path, the stage division of patent growth and application; the cross-domain feature combination data includes the repeated technical terms in each field and their co-occurrence frequency and correlation score; the patent structure layout list includes the hierarchical arrangement of key technical features in the patent group and the upper and lower node relationships of patents in the same group; the sorted optimized patent tree structure includes the main path sorting, node priority and hierarchical distribution information.
[0014] As a further solution of the present invention, based on patent text data, the technical features in the patent text are analyzed, and the technical features of each patent are classified and integrated to obtain the patent technical feature set. The specific steps are as follows:
[0015] S101: Based on the patent text data, extract the patent technical terms, functional descriptions, and key information of application scenarios, retrieve and match the corresponding technical terms, identify the functional description sentences in the text, screen the key technical terms, annotate the matching functional information, and generate the initial patent information set;
[0016] S102: Based on the initial patent information set, semantic analysis of verb and noun phrases is performed. Verbs are related to the technical objectives and nouns are related to the implementation methods. The technical objective markers and functional implementation contents are integrated to obtain the core content set of the patent.
[0017] S103: Based on the core content set of the patent, the integrated technical objectives and implementation method data are divided into independent functional parts, and classified according to technical functional fields and application scenarios to obtain the patent technical feature set.
[0018] As a further solution of the present invention, based on the patent technology feature set, the patent data is sorted and integrated by time nodes, combined with the data fluctuations at the time nodes, and the historical dependency characteristics are captured to obtain the patent technology status prediction results. The specific steps are as follows:
[0019] S201: Based on the patent technical feature set, extract the number of patent citations, patent family size, and citation count data, segment the data by time dimension, and aggregate the data by year and quarter to obtain a segmented patent data set;
[0020] S202: Based on the segmented patent dataset, sort and integrate each data item by time node, extract time change features, mark the time node of each data change by constructing a sequence, and generate a patent time series set;
[0021] S203: Based on the patent time series set, data fluctuation feature analysis is performed to extract the patent citation growth pattern, patent family expansion rate, and time-varying features of the number of citations. By identifying fluctuations in time-varying trends, historical dependency features are captured to obtain patent technology status prediction results.
[0022] As a further solution of the present invention, based on the patent technology status prediction results, patents are classified according to life cycle stages. According to the classification results, each type of patent is divided into path hierarchies. The specific steps of generating hierarchical patent path information are as follows:
[0023] S301: Based on the patent technology status prediction results, the patent life cycle classification standard is generated by mapping the patent growth rate, citation frequency change and patent family expansion trend data to the life cycle stage;
[0024] S302: Based on the patent life cycle classification standard, the patents are classified into inconsistent categories according to the life cycle stages corresponding to their data, and the patents that meet the life cycle characteristics are grouped to obtain patent life cycle classification results;
[0025] S303: Based on the patent life cycle classification results, according to the grouping results of each category, a path hierarchy is constructed, the life cycle stages are mapped to hierarchical path nodes, and hierarchical patent path information is generated.
[0026] As a further solution of the present invention, the specific steps of extracting recurring technical terms and functional descriptions based on the patent technical feature data and hierarchical patent path information in the patent technical feature set, determining potential feature item combinations, and obtaining cross-domain feature combination data are as follows:
[0027] S401: Based on the patent technical feature set and the hierarchical patent path information, compare patent feature items in multiple technical fields, filter out repeated technical terms and functional descriptions, and obtain patent feature comparison data;
[0028] S402: Based on the patent feature comparison data, collect statistics on the distribution of the selected feature item combinations in the technical field, record the occurrence frequency and application scenarios of the feature combinations, and generate feature distribution data;
[0029] S403: Based on the feature distribution data, calculate the association strength of the feature item combination, analyze the co-occurrence frequency and the association score, screen the matching cross-domain feature combination, and obtain cross-domain feature combination data.
[0030] As a further solution of the present invention, the technical features of the patents are converted into vector form based on the cross-domain feature combination data, patents with similarity matching are grouped together, the importance and co-occurrence frequency of the technical features are analyzed, and hierarchical nodes are divided to obtain a patent structure layout list. The specific steps are as follows:
[0031] S501: Based on the cross-domain feature combination data, separate the technical feature content of the patent, convert each technical feature into a standard vector form using numerical mapping, and integrate all feature vectors for data standardization to generate a vectorized feature set;
[0032] S502: Based on the vectorized feature set, perform feature matching analysis between patents, compare vector values by calculating the similarity between features, combine patents with high similarity, and generate patent similarity analysis results;
[0033] S503: Based on the patent similarity analysis results, feature similarity clustering is performed, hierarchical nodes are established for the matched patent features according to importance and co-occurrence frequency, upper nodes and lower nodes are divided according to the degree of technical impact, and a patent structure layout list is generated.
[0034] As a further solution of the present invention, the similarity between features is calculated by vector value, using the formula:
[0035] ;
[0036] Get the feature similarity between two patents ;
[0037] in, The first patent is The numerical value of the feature, The second patent is in The numerical value of a feature.
[0038] As a further solution of the present invention, based on the patent structure layout list, weights are assigned to the patent nodes in the path, and the patent tree layout is optimized and adjusted to generate a sorted and optimized patent tree structure. The specific steps are as follows:
[0039] S601: Based on the patent structure layout list and according to preset indicators, weights are assigned to patent nodes in each path according to their life cycle stages to obtain a path node weight set;
[0040] S602: Based on the path node weight set, accumulate the node weights in the path, obtain the total weight value and average weight value of each path, sort the paths by weight value, and obtain a path weight sorting set;
[0041] S603: Based on the path weight ranking set, adjust the patent tree layout, set the high-weight path as the main path, and assign the lower-weight paths to secondary branches in order to generate a ranking-optimized patent tree structure.
[0042] A patent tree includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned patent tree and the construction method thereof when executing the computer program.
[0043] Compared with the prior art, the advantages and positive effects of the present invention are:
[0044] In the present invention, the precise positioning of patent technical elements is achieved through intelligent information extraction and semantic analysis. Combined with the deep semantic analysis of verbs and noun phrases, the technical goals and implementation methods of the patent are carefully extracted and their functions are decomposed and integrated, so that the core content of the patent is presented in a clear and structured manner. The technical features are further segmented and integrated according to time nodes and application fields, and combined with the fluctuation feature analysis of the time series, the historical dependency features are identified, which can effectively predict the development status of the patent technology. Through feature matching and similarity calculation, the accurate grouping of patents and the analysis of cross-domain feature combinations can be achieved. Through the path weight sorting set, the high-weight path in the patent tree is set as the main path, which optimizes the hierarchical structure of the patent tree and enables enterprises to clearly identify the relationship between the primary and secondary paths in the patent technology layout, thereby allocating resources more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0046] Figure 2 A detailed flow chart of step S1 of the present invention;
[0047] Figure 3 A detailed flow chart of step S2 of the present invention;
[0048] Figure 4 A detailed flow chart of step S3 of the present invention;
[0049] Figure 5 A detailed flow chart of step S4 of the present invention;
[0050] Figure 6 The flowchart of step S5 of the present invention is refined;
[0051] Figure 7 This is a detailed flow chart of step S6 of the present invention. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0053] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0054] See also Figure 1 The present invention provides a technical solution: a patent tree and a method for constructing the same, comprising the following steps:
[0055] S1: Based on patent text data, analyze the technical features in the patent text, classify and integrate the technical features of each patent, and obtain the patent technical feature set;
[0056] S2: Based on the patent technology feature set, the patent data is sorted and integrated by time node, combined with the data fluctuation of the time node, to capture the historical dependency characteristics and obtain the patent technology status prediction results;
[0057] S3: Based on the patent technology status prediction results, the patents are classified according to the life cycle stage. According to the classification results, the path level of each type of patent is divided to generate hierarchical patent path information;
[0058] S4: Based on the patent technical feature data and hierarchical patent path information in the patent technical feature set, repeatedly appearing technical terms and functional descriptions are extracted to determine potential feature item combinations and obtain cross-domain feature combination data;
[0059] S5: Based on cross-domain feature combination data, the technical features of the patents are converted into vector form, patents with similarity matching are grouped together, the importance and co-occurrence frequency of the technical features are analyzed, and hierarchical nodes are divided to obtain a list of patent structure layouts;
[0060] S6: Based on the patent structure layout list, assign weights to the patent nodes in the path, optimize and adjust the patent tree layout, and generate a sorted and optimized patent tree structure.
[0061] The patent technical feature set includes the patent's technical terminology set, functional description structure and application scenario classification. The patent technology status prediction results include the change curve of patent citations, expansion speed index and citation number change rate. The hierarchical patent path information includes the position of patents in each stage in the path, the stage division of patent growth and application, and the cross-domain feature combination data includes the repeated technical terms in each field and their co-occurrence frequency and correlation score. The patent structure layout list includes the hierarchical arrangement of key technical features in the patent group and the superior and subordinate node relationship of patents in the same group. The sorted and optimized patent tree structure includes the main path sorting, node priority and hierarchical distribution information.
[0062] See also Figure 2 Based on the patent text data, we analyze the technical features in the patent text, classify and integrate the technical features of each patent, and obtain the patent technical feature set in the following specific steps:
[0063] S101: Based on the patent text data, extract the patent technical terms, functional descriptions, and key information of application scenarios, retrieve and match the corresponding technical terms, identify the functional description sentences in the text, screen the key technical terms, annotate the matching functional information, and generate the initial patent information set;
[0064] First, natural language processing is used to segment the text content, and the segmentation results are semantically analyzed in combination with the context. Patent terms that meet the characteristics of the field are identified, frequently appearing terms are marked, and their technical relevance is determined by counting word frequencies. Then, the semantic role labeling method is used to further identify and extract functional description phrases, extract words and phrases related to technical applications and functional implementations, and eliminate technology-irrelevant or repeated words and phrases through syntactic analysis. All extracted technical terms, functional descriptions, application scenarios, etc. are classified and recorded, and an index table is established according to the characteristics to generate an initial patent information set.
[0065] S102: Based on the initial patent information set, perform semantic analysis of verb and noun phrases. Verbs are about the content of technical objectives and nouns are about the content of implementation methods. Integrate the technical objective markers and functional implementation content to obtain the core content set of the patent;
[0066] First, verb phrases in the patent are extracted and marked, and functional verbs are associated with technical goals. The applicability of verbs to application fields is analyzed. Then, the semantic role labeling method is used to extract the content representing the implementation method in the noun phrases. The keywords in these phrases are matched with specific implementation steps. A combination relationship between goals and implementation methods is constructed, and each set of goals and implementation content is compared and analyzed item by item. Content with similar application characteristics is integrated into the same category. Finally, all combined content is classified and grouped according to the technical functional field to obtain the core content set of the patent.
[0067] S103: Based on the core content set of the patent, the integrated technical objectives and implementation method data are divided into independent functional parts, and classified according to technical functional fields and application scenarios to obtain the patent technical feature set;
[0068] First, a hierarchical clustering method is used to divide different technical functions and application scenarios, and the technical goals and corresponding application scenarios are classified separately. The implementation methods contained in each application scenario are verified for correlation, and the combination of goals and implementation methods is mapped to the functional field table. Then, within each category, the classification is refined according to the applicability of the application field, and the technical features of each application field are hierarchically allocated and organized to form independent functional units, thereby obtaining a set of patent technical features.
[0069] See also Figure 3 Based on the patent technology feature set, the patent data is sorted and integrated by time nodes, combined with the data fluctuations at time nodes, and the historical dependency characteristics are captured to obtain the patent technology status prediction results. The specific steps are as follows:
[0070] S201: Based on the patent technical feature set, extract the number of patent citations, patent family size, and citation count data, segment the data by time dimension, and summarize the data by year and quarter to obtain a segmented patent dataset;
[0071] First, each data item is sorted according to the time node of the patent, and each data item is grouped by time. The data is summarized year by year to obtain the growth of the annual data. At the same time, the quarterly data is segmented and statistically analyzed to observe the quarterly changes. Then, a segmented statistical table is established using the time accumulation method. The time dimension of the data is subdivided according to the set time interval. The values of citation counts, patent family size, citation counts, etc. are superimposed and compared according to the specified time period. The differences in data between years and quarters are recorded, and a complete record of segmented values is constructed to obtain a segmented patent data set.
[0072] S202: Based on the segmented patent dataset, sort and integrate each data item by time node, extract time change features, and mark the time node of each data change by constructing a sequence to generate a patent time series set;
[0073] First, the citation counts, patent family size, and citation counts at each time node are arranged in chronological order, and the changing trends of each time node are recorded. By constructing a time series analysis framework, the temporal changes of each data item are marked, and the growth and decline at different time points are recorded in turn. The change values of each time node are accumulated, and the amplitude of the numerical changes is recorded by year and quarter, gradually forming a continuous time series. By classifying the trend of the change rate and amplitude of each time node, significant growth or decline trends are marked to generate a patent time series set.
[0074] S203: Based on the patent time series set, data fluctuation characteristics are analyzed to extract the patent citation growth pattern, patent family expansion rate, and time-varying characteristics of the number of citations. By identifying fluctuations in time-varying trends and capturing historical dependency characteristics, the patent technology status prediction results are obtained.
[0075] Using differential analysis, we capture the fluctuation characteristics between time nodes. We obtain data on citation counts, patent family size, and citation counts from the time series, perform adjacent node differences, and record these differences in a fluctuation table to identify the changing trends between time nodes. We then use a sliding average method to average the difference results, recording the fluctuation mean in the time series. Time nodes with significant differences are labeled as fluctuation nodes, and periods of significant differential fluctuation are classified as high-volatility intervals. These high-volatility intervals are then compared and analyzed with stable intervals. The characteristics of time nodes during periods of continuous high-volatility are recorded. Assuming the citation count of the previous time node as a benchmark, we set a 10% difference threshold. For example, if the number of citations in the previous quarter is 500 and the number of citations in the next quarter is 560, the increase is (560-500) / 500=0.12, or 12%. This increase is greater than the 10% threshold, and therefore the time node is considered to have a significant difference. Based on these consecutive periods of high volatility, we can identify trends in technological change. For example, high volatility at several consecutive time points (rapid increases in citations and patent family size) indicates that the technology is in an expansion phase, foreshadowing the possibility of continued growth. If a period of high volatility is followed by several periods of plateauing or declining activity, it may be entering a period of stability or decline. Therefore, by analyzing the characteristics of these time points, we can infer future activity and potential lifecycle stages, thereby generating a prediction of the patent technology status.
[0076] See also Figure 4 Based on the prediction results of patent technology status, patents are classified according to the life cycle stage. According to the classification results, the path level of each type of patent is divided. The specific steps for generating hierarchical patent path information are as follows:
[0077] S301: Based on the prediction results of patent technology status, the patent life cycle stage is divided into categories by mapping the growth rate, citation frequency change and patent family expansion trend data to the life cycle stage to generate a patent life cycle classification standard;
[0078] Using a lifecycle mapping method, we match data on patent growth rates, changes in citation frequency, and patent family expansion trends. We first analyze the annual growth rate of patents, setting several indicator ranges: high growth rate, stable growth, and negative growth. For example, a patent is classified as in the high growth stage if its annual citation growth rate exceeds 20%, its citation frequency continues to rise over the past three years, and its patent family rapidly expands from an initial 2 patents to over 10 patents. In the stable growth stage, a patent is classified as in the stable growth stage if its annual citation growth rate fluctuates between 5% and 10%, its citation frequency remains stable or increases slightly over the past three years, and its patent family remains within a stable range of 5-8 patents. In the decline stage, a patent is classified as in the decline stage if its annual citation growth rate drops to a negative value (e.g., -5%), its citation frequency decreases significantly in recent years, and its patent family expansion stagnates, remaining at fewer than 3 patents. Patents within each indicator range are classified, and then the changes in the patent's citation frequency are correlated with the life cycle stage. High-frequency citations, continuous growth, and reduction are mapped to different life cycle intervals. Then, the expansion trend is judged based on the growth curve of patent family expansion. If the patent family expands rapidly in the short term, it is marked as a high expansion stage, while if the growth slows down, it is marked as a stable expansion stage. The patent family expansion is comprehensively mapped with the data on citation frequency and growth rate to generate a patent life cycle classification standard.
[0079] S302: Based on the patent life cycle classification standard, the patents are classified into inconsistent categories according to the life cycle stages corresponding to their data, and the patents that meet the life cycle characteristics are grouped to obtain the patent life cycle classification results;
[0080] Patents are divided into their respective categories according to the corresponding life cycle stages through the classification screening method. First, the characteristic values of each stage in the classification standard are compared, and the growth rate, citation frequency change and patent family expansion characteristics of the patent are confirmed one by one to see whether they meet a certain category. According to the set value of the standard and the degree of conformity of the actual patent data, patents that meet specific life cycle characteristics are classified into corresponding groups. If the growth rate and citation frequency change of a patent meet different categories relatively, it will be preferentially classified into the most suitable stage based on the main characteristics. The patents in each life cycle are arranged in order, numbered and marked with characteristics in the table to obtain the patent life cycle classification results.
[0081] S303: Based on the patent lifecycle classification results, a path hierarchy is constructed according to the grouping results of each category, and the lifecycle stages are mapped to hierarchical path nodes to generate hierarchical patent path information;
[0082] The classification results are processed into hierarchical paths through the hierarchical path construction method, and each life cycle category is mapped to a hierarchical path structure. First, the high-growth category in the life cycle classification is set as the main node of the path, and the patents in this category are set as the upper node of the path. Then, the categories with steady growth and slow expansion are set as the secondary nodes of the path, and the association hierarchy is formed one by one. The technical correlation within each category is used to construct a patent path table, and a hierarchical path structure is formed according to the life cycle stage. The path hierarchy is adjusted and arranged to obtain hierarchical patent path information.
[0083] See also Figure 5 ,Based on the patent technical feature data and hierarchical patent path information in the patent technical feature set,,repeated technical terms and functional descriptions are extracted, potential feature item combinations are determined, and the specific steps for,obtaining cross-domain feature combination data are as follows:
[0084] S401: Based on the patent technical feature set and hierarchical patent path information, compare patent feature items in multiple technical fields, filter out repeated technical terms and functional descriptions, and obtain patent feature comparison data;
[0085] Patent feature items in various technical fields are compared through comparative analysis methods. First, the technical terms, functional descriptions, application scenarios and other features of each patent are extracted. Then, keyword matching screening tools are used for repeated screening. The technical terms and functional description items that appear repeatedly in various fields are identified and summarized, and common terms that appear multiple times are marked. For terms or descriptions that appear multiple times but have similar meanings, semantic analysis tools are used to compare word meanings to determine whether they are classified as similar terms. Then, the frequency of occurrence of all terms and functional description items is classified, common items and field-unique items are marked, and patent feature comparison data is output.
[0086] S402: Based on the patent feature comparison data, collect statistics on the distribution of the selected feature item combinations in the technical field, record the occurrence frequency and application scenarios of the feature combinations, and generate feature distribution data;
[0087] The distribution of the selected feature item combinations in the technical fields is statistically analyzed through the frequency statistics method. First, the frequency of occurrence of the feature combination in each field is obtained, and the high-frequency occurrence items of the feature items are marked as general items or cross-field combinations using the frequency classification method. Then, the high-frequency usage scenarios, specific applications and applicable technical fields of each feature combination are registered. For combinations with higher frequencies, their corresponding application scenarios and technical branches are recorded one by one, and a feature distribution table is constructed. Finally, the distribution of feature items in all fields is sorted out to generate feature distribution data including application scenarios and occurrence frequencies.
[0088] S403: Based on the feature distribution data, calculate the correlation strength of the feature item combination, analyze the co-occurrence frequency and correlation score, screen the matching cross-domain feature combination, and obtain cross-domain feature combination data;
[0089] Based on the characteristic distribution data, according to the formula Calculate the association strength of cross-domain feature combinations. It represents the correlation strength of feature combinations, and is used to measure the strong correlation of feature combinations across fields. n is the total number of fields covered by each feature combination. It represents the frequency of occurrence of the feature combination in the i-th field, which is calculated by dividing the number of occurrences of the feature combination by the total number of patents in the field. It represents the correlation score of the feature combination in the i-th field, which is quantified based on the applicable scope of the features and the cross-field correlation score, and is scored by the breadth of application scenario coverage. Assuming that the feature combination is "image recognition and edge detection", its correlation score is scored in the following fields: Field A: Medical image analysis, score: 0.9. In medical image analysis, image recognition and edge detection are crucial. For example, they are widely used in tumor boundary recognition, so the score is high. Field B: Autonomous driving technology, score: 0.7. In autonomous driving, image recognition and edge detection are used to identify roads and obstacles, but their applications are less concentrated in specific scenarios, and the correlation score is relatively low. Field C: Agricultural detection, score: 0.5. In agriculture, this feature combination is used for plant health monitoring and soil analysis, but it mainly relies on other feature recognitions, so the correlation score is low.
[0090] If the feature combination appears 50 times in a certain field record and the total number of field records is 500, the calculated frequency is If the feature combination is widely used in multiple fields and has significant technical effects, the score is set to 0.8. If the frequency and score of this feature combination in n different fields are respectively , , , , the calculation formula is:
[0091] ;
[0092] If the average correlation strength of similar feature combinations in the past was 0.15, the resulting value of 0.22 is higher than the historical average, indicating that the correlation strength of this feature combination in the current application has been significantly improved. The results show that this feature combination has gradually strengthened in cross-domain applications, its potential value has increased significantly, and it is suitable for further promotion to other fields to expand the coverage of the technology.
[0093] See also Figure 6,Based on the cross-domain feature combination data, the technical features of the patents are converted into vector form, the patents with similarity matching are grouped together, the importance and co-occurrence frequency of the technical features are analyzed and the hierarchical nodes are divided. The specific steps to obtain the patent structure layout list are as follows:
[0094] S501: Based on cross-domain feature combination data, separate the technical features of the patent, convert each technical feature into a standard vector form using numerical mapping, and integrate all feature vectors for data standardization to generate a vectorized feature set;
[0095] This is achieved through a feature extraction and numerical mapping process. First, technical keywords are extracted from each patent. Natural language processing (NLP) tools (such as Python's NLTK or spaCy) are used for text segmentation and part-of-speech tagging. Technical vocabulary and terminology in the feature description are identified, stop words are removed, and high-frequency technical terms are retained. Next, word embedding models (such as Word2Vec or GloVe) are used to convert these keywords into numerical vector representations, obtaining the initial vector value for each keyword. The vectors generated by word embeddings are further encoded and adjusted to ensure that each vector value reflects the similarity between technical features. Subsequently, data normalization methods are used to adjust the vector data to a uniform interval (such as [0, 1]) to maintain a consistent scale across the data for subsequent similarity analysis. Finally, the vector values of all technical keywords are integrated to form a complete technical feature vector for the patent, ensuring that each patent has an independent and standardized representation in the vector space. The technical features of all patents are vectorized and aggregated into a standard vectorized feature set.
[0096] S502: Based on the vectorized feature set, feature matching analysis is performed between patents. The vector values are numerically compared by calculating the similarity between the features, and patents with high similarity are combined to generate patent similarity analysis results.
[0097] To calculate the similarity between features, we use the following formula:
[0098] ;
[0099] Get the feature similarity between two patents ;
[0100] in, The closer the cosine value of is to 1, the higher the similarity between patent features is. The first patent is The numerical value of the item feature is obtained through feature extraction and numerical mapping. After keyword extraction with the help of feature extraction tools (such as Python's NLTK or spaCy), it can be converted into a numerical vector through a vocabulary embedding model (such as Word2Vec, GloVe). The second patent is in The value of the feature is obtained in the same way as Same, ensure that the corresponding dimensions between numerical vectors are consistent, and : Represent the modulus lengths of eigenvectors A and B respectively.
[0101] Use text extraction tools (such as NLTK) to obtain core technical vocabulary from patent documents, and then use the word embedding model (Word2Vec) to convert these vocabulary into vector form. Assume that the feature vectors of patent A and patent B are and , respectively represent the relative weights between technical features.
[0102] Dot product part: ;
[0103] Vector modulus: For patent A, ;
[0104] For Patent B, ;
[0105] Cosine similarity calculation: ;
[0106] In order to determine which patent features have significant similarity, a similarity threshold is set. Assume that the threshold is set to 0.75, that is, if The two patents are considered to have similar features and share common technical features or application scenarios. The threshold can be adjusted based on the specific research or technical field.
[0107] Suppose the similarity of the following patent features has been calculated, resulting in the following results: the similarity between Patents A and B is 0.82, the similarity between Patents A and C is 0.6, the similarity between Patents B and C is 0.78, and the similarity between Patents A and D is 0.85. Based on the set similarity threshold of 0.75, these similarity results are grouped as follows: High Similarity Group 1: Patents A and B, and Patents A and D. The similarity values of these two pairs are 0.82 and 0.85, respectively, both exceeding the threshold, indicating that these three patents are highly similar in terms of technical features. They can be placed in the same group, such as under the same node in the tree structure, to highlight the close connection between their technical content. High Similarity Group 2: Patents B and C, with a similarity of 0.78, meets the threshold, indicating that Patents B and C also have high similarity in terms of features and are suitable for inclusion in another similarity group. By comparing and grouping patent similarities, technical branches can be constructed within the patent tree. For example, a group containing patents A, B, and D can be marked on a branch of the patent tree to show their close technical connection, and their common features can be noted in the branch name to more clearly view their development trends in subsequent technology evolution and patent landscape analysis. The grouping of patents B and C can be marked on another branch, which can be further expanded to include other patents with similar technical features.
[0108] S503: Based on the results of the patent similarity analysis, feature similarity clustering is performed. Hierarchical nodes are established for the matched patent features according to importance and co-occurrence frequency. Upper and lower nodes are divided according to the degree of technical impact, and a patent structure layout list is generated.
[0109] First, using the similarity values between patents in the similarity matrix as initial data, a hierarchical clustering algorithm was used to cluster the patent technical features. Following the principle of prioritizing groups with high similarity, patents with higher similarity were gradually grouped together to form clusters. During the clustering process, the co-occurrence frequency of feature vectors was used as a key metric. Frequency statistics were used to record the co-occurrence frequency of technical features within each cluster, and frequently occurring technical features were labeled to reflect their core position within each cluster. After clustering, a cluster partitioning method was used to establish a hierarchical structure, classifying feature nodes according to their technological influence and importance. Technical feature nodes with high co-occurrence frequencies and high values in the similarity matrix were designated as superior nodes, representing the core technology. Nodes with low co-occurrence frequencies and relatively low similarity values were designated as subordinate nodes, representing secondary or auxiliary technologies within the field. Once the hierarchical structure was completed, a patent structure layout list containing the stratified results of each patent feature was generated, presenting a list of patent structure layouts based on technical feature similarity.
[0110] See also Figure 7Based on the patent structure layout list, the patent nodes in the path are weighted and the patent tree layout is optimized and adjusted. The specific steps to generate the sorted optimized patent tree structure are as follows:
[0111] S601: Based on the patent structure layout list and preset indicators, weight the patent nodes in each path according to their life cycle stages to obtain a path node weight set;
[0112] Using pre-set lifecycle stage indicators, each patent node in the path is assigned a weight value based on its lifecycle stage. First, based on the lifecycle stage, nodes are divided into high-growth stage, mature stage, and decline stage, and a corresponding weight is assigned to each stage (for example, high-growth stage is 3, mature stage is 2, and decline stage is 1). Next, each node in the path is traversed in turn, and the corresponding weight value is added to the node attribute set based on its lifecycle stage to form an initial path node weight set. Using a document processing tool, the weight value of each node is recorded one by one, and the weight values of all nodes are recorded in the path weight set, so that the weights of each path can be subsequently integrated into the path weight set.
[0113] S602: Based on the path node weight set, accumulate the node weights in the path, obtain the total weight value and average weight value of each path, sort the paths by weight value, and obtain a path weight sorting set;
[0114] By counting the node weights in the path, we first traverse the nodes in each path, call the weight values stored in the path node weight set, and add the weight values of all nodes to form the total weight value of each path. Then, we divide the total weight value by the number of path nodes to obtain the average weight value of the path. For different paths, we record the total weight values and average weight values of all paths in sequence, pass these values to the patent management tool for sorting, and use the sorted path weight set to construct the final path weight result to ensure that the weight sorting is reasonable and the data is clear.
[0115] S603: Based on the path weight ranking set, adjust the patent tree layout, set the high-weight path as the main path, and assign the lower-weight paths to secondary branches in order, to generate a ranking-optimized patent tree structure;
[0116] The patent tree layout is adjusted and optimized based on the weight sorting results. By comparing the weight values in the path weight sorting set, the paths with higher weights in the sorting are assigned as main paths. The main paths are marked at the core branch positions of the patent tree and are displayed preferentially on the main branches of the tree structure. At the same time, the paths with the second highest weights are arranged to the secondary branches according to their weights. By appropriately allocating the weight sorting sets of the secondary paths, it is ensured that the high-weight paths are displayed preferentially in the main paths of the patent tree, ultimately forming a patent tree structure layout that has been sorted and optimized.
[0117] A patent tree includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the above-mentioned patent tree and the construction method thereof are implemented.
[0118] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.
Claims
1. A method for constructing a patent tree, characterized in that: The following steps are involved: Based on patent text data, we analyze the technical features in the patent text, classify and integrate the technical features of each patent, and obtain a patent technical feature set; Based on the patent technology feature set, the patent data is sorted and integrated by time nodes, and the data fluctuations at the time nodes are combined to capture the historical dependency characteristics to obtain the patent technology status prediction results; Based on the patent technology status prediction results, the patents are classified according to the life cycle stage. According to the classification results, the path level of each type of patent is divided to generate hierarchical patent path information; Based on the patent technical feature data and hierarchical patent path information in the patent technical feature set, recurring technical terms and functional descriptions are extracted to determine potential feature item combinations and obtain cross-domain feature combination data; Based on the cross-domain feature combination data, the technical features of the patents are converted into vector form, patents with similarity matching are grouped together, the importance and co-occurrence frequency of the technical features are analyzed, and hierarchical nodes are divided to obtain a patent structure layout list; Based on the patent structure layout list, weights are assigned to the patent nodes in the path, and the patent tree layout is optimized and adjusted to generate a patent tree structure with optimized sorting; The patent technical feature set includes a set of patent technical terms, a functional description structure, and a classification of application scenarios; the patent technical status prediction results include a curve of patent citation changes, an expansion speed index, and a rate of change in the number of citations; the hierarchical patent path information includes the position of patents in each stage in the path, and the stage division of patent growth and application; the cross-domain feature combination data includes recurring technical terms in each field, their co-occurrence frequency, and correlation scores; the patent structure layout list includes a hierarchical arrangement of key technical features in a patent group and the superior-subordinate node relationship of patents in the same group; the sorted and optimized patent tree structure includes main path sorting, node priority, and hierarchical structure distribution information.
2. The method for constructing a patent tree according to claim 1, characterized in that: Based on patent text data, we analyze the technical features in the patent text, classify and integrate the technical features of each patent, and obtain the patent technical feature set in the following specific steps: Based on patent text data, we extract patent technical terms, functional descriptions, and key information of application scenarios, retrieve and match corresponding technical terms, identify functional description sentences in the text, screen key technical terms, annotate matching functional information, and generate an initial patent information set. Based on the initial patent information set, a semantic analysis of verb and noun phrases is performed. Verbs are about the content of the technical objectives and nouns are about the content of the implementation methods. The technical objective markers and functional implementation content are integrated to obtain the core content set of the patent; Based on the core content set of the patent, the integrated technical objectives and implementation method data are divided into independent functional parts, and classified according to technical functional fields and application scenarios to obtain the patent technical feature set.
3. The method for constructing a patent tree according to claim 1, characterized in that: Based on the patent technology feature set, the patent data is sorted and integrated by time nodes, combined with the data fluctuations at time nodes, to capture historical dependency features and obtain the patent technology status prediction results. The specific steps are as follows: Based on the patent technical feature set, extract the number of patent citations, patent family size, and citation count data, segment the data by time dimension, and summarize the data by year and quarter to obtain a segmented patent data set; Based on the segmented patent dataset, each data item is sorted and integrated by time node, time change characteristics are extracted, and the time node of each data change is marked by constructing a sequence to generate a patent time series set; Based on the patent time series set, data fluctuation characteristics are analyzed to extract the patent citation growth pattern, patent family expansion rate, and time-varying characteristics of the number of citations. By identifying fluctuations in time-varying trends and capturing historical dependency characteristics, the patent technology status prediction results are obtained.
4. The method for constructing a patent tree according to claim 1, characterized in that: Based on the patent technology status prediction results, the patents are classified according to the life cycle stage. According to the classification results, the path level of each type of patent is divided. The specific steps for generating hierarchical patent path information are as follows: Based on the patent technology status prediction results, the patent life cycle classification standard is generated by mapping the patent growth rate, citation frequency changes and patent family expansion trend data to the life cycle stages; Based on the patent life cycle classification standard, patents are classified into inconsistent categories according to the life cycle stages corresponding to their data, and patents that meet the life cycle characteristics are grouped to obtain patent life cycle classification results; Based on the patent life cycle classification results, a path hierarchy is constructed according to the grouping results of each category, and the life cycle stages are mapped to hierarchical path nodes to generate hierarchical patent path information.
5. The method for constructing a patent tree according to claim 1, characterized in that: Based on the patent technical feature data and hierarchical patent path information in the patent technical feature set, the specific steps for extracting recurring technical terms and functional descriptions, determining potential feature item combinations, and obtaining cross-domain feature combination data are as follows: Based on the patent technical feature set and hierarchical patent path information, patent feature items in multiple technical fields are compared, and repeated technical terms and functional descriptions are screened to obtain patent feature comparison data; Based on the patent feature comparison data, statistics are collected on the distribution of the selected feature item combinations in the technical field, the occurrence frequency and application scenarios of the feature combinations are recorded, and feature distribution data is generated; Based on the feature distribution data, the association strength of the feature item combination is calculated, the co-occurrence frequency and the association score are analyzed, and matching cross-domain feature combinations are screened to obtain cross-domain feature combination data.
6. The method for constructing a patent tree according to claim 1, characterized in that: Based on the cross-domain feature combination data, the technical features of the patents are converted into vector form, patents with similarity matching are grouped together, the importance and co-occurrence frequency of the technical features are analyzed, and hierarchical nodes are divided to obtain the patent structure layout list. The specific steps are as follows: Based on the cross-domain feature combination data, the technical feature content of the patent is separated, each technical feature is converted into a standard vector form using numerical mapping, and all feature vectors are integrated for data standardization to generate a vectorized feature set; Based on the vectorized feature set, feature matching analysis is performed between patents. The vector values are numerically compared by calculating the similarity between the features, and patents with high similarity are combined to generate patent similarity analysis results. Based on the results of the patent similarity analysis, feature similarity clustering is performed, hierarchical nodes are established for the matched patent features according to importance and co-occurrence frequency, upper and lower nodes are divided according to the degree of technical impact, and a patent structure layout list is generated.
7. The method for constructing a patent tree according to claim 6, characterized in that: To calculate the similarity between features, we use the following formula: ; Get the feature similarity between two patents ; in, The first patent is The numerical value of the feature, The second patent is in The numerical value of a feature.
8. The method for constructing a patent tree according to claim 1, characterized in that: Based on the patent structure layout list, weights are assigned to the patent nodes in the path, and the patent tree layout is optimized and adjusted to generate a sorted and optimized patent tree structure. The specific steps are as follows: Based on the patent structure layout list, according to preset indicators, the patent nodes in each path are weighted according to the life cycle stage to which they belong, and a path node weight set is obtained; Based on the path node weight set, accumulating the node weights in the path, obtaining the total weight value and average weight value of each path, and sorting the paths by weight value to obtain a path weight sorting set; Based on the path weight ranking set, the patent tree layout is adjusted, high-weight paths are set as main paths, and lower-weight paths are assigned to secondary branches in order to generate a ranking-optimized patent tree structure.
9. A patent tree comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, the steps of the method for constructing a patent tree according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Generation method and device of patent technology route and computer equipment
CN118427356A