Method for automatically constructing strategic emerging industrial chain based on patent information

By integrating and comprehensively evaluating multi-dimensional data, this technology addresses the issues of incomplete data coverage, low accuracy, and poor efficiency in existing technologies. It enables the construction of highly accurate strategic emerging industry chains and timely risk warnings, and is applicable to fields such as new energy, semiconductors, and biomedicine.

CN121544060APending Publication Date: 2026-02-17CHUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511570894.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing methods for constructing strategic emerging industry chains based on patent information suffer from incomplete data coverage, low accuracy, poor efficiency, and a lack of dynamic early warning, resulting in one-sided and inaccurate industry chain construction and an inability to respond to risks in a timely manner.

Method used

By dividing data into batches, collecting data from multiple dimensions, processing data in depth, and conducting comprehensive evaluation, and by combining patent and industry-related data, the R-index is calculated using a multi-dimensional data fusion model and the entropy weight method, thereby achieving high-precision construction and dynamic early warning of the industrial chain.

Benefits of technology

It achieves more comprehensive data coverage, higher construction accuracy, higher processing efficiency, and more timely dynamic early warning, improves the matching degree of industrial chain structure and the accuracy of core node identification, reduces application costs, and is compatible with a variety of strategic emerging industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121544060A_ABST
    Figure CN121544060A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically constructing a strategic emerging industrial chain based on patent information, belongs to the technical field of strategic emerging industrial chain construction, and aims to solve the problems of one-sided data coverage, low construction precision and lack of dynamic early warning of an existing method. The method comprises the following five core steps: S1, data batch division: dividing patent technology and industry associated target data into batches 1, 2,..., n (n is preferentially 20) according to quarterly; s2, multi-dimensional data collection: synchronously collecting patent technology data such as patent nodes, patent quotation, patent-industry relationships and the like, and industry associated data such as industry technologies, upstream supply and demand, downstream markets and the like; s3, data processing: calculating a patent node dominance degree coefficient, a patent citation value coefficient, an industry integrating degree coefficient, an industry technology association fusion degree coefficient, an industry upstream association dependency coefficient and an industry downstream relationship development potential coefficient through the exclusive model; s4, comprehensive evaluation: calculating a comprehensive evaluation index R through min-max normalization and an entropy weight method (weight distribution: alpha i0.2, lambda i0.2, delta i0.15, beta'i0.15, gamma'i0.15 and epsilon'i0.15); and S5, performing final judgment, and sending out a normal / early warning signal based on the R preset value 0.65 (30 qualified industrial chain statistical mean). Through deep fusion, quantitative modeling and dynamic evaluation of patent and industry data, the industrial chain structure matching degree reaches 92.3%, the core node identification accuracy reaches 89.7%, only 2.5 hours are needed for processing 100,000 pieces of patent data, risks such as upstream raw material shortage and downstream demand insufficiency can be positioned in time, and the method is suitable for large-scale popularization and application. The method is suitable for strategic emerging industries such as new energy, semiconductors and biological medicine, and provides support for industry chain optimization and decision making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of strategic emerging industry chain construction, in particular to a strategic emerging industry chain automatic construction method based on patent information, which is suitable for the industry chain structure analysis, core node identification and dynamic optimization of strategic emerging industries such as new energy, semiconductor and biomedicine, and realizes the automatic and high-precision construction and evaluation of the industry chain through the fusion analysis of patent data and industry correlation data. BACKGROUND

[0002] With the deep integration of information technology and strategic emerging industries, the accurate construction of industry chain is of great significance to grasp the development direction of the industry and optimize the resource layout. The existing strategic emerging industry chain construction method based on patent information usually includes five steps of patent data collection-data preprocessing-feature extraction-industry chain construction-result display: among them, the patent data collection only covers the basic information of patent text (such as application number, abstract), the data coverage dimension is single; the data preprocessing mainly includes simple cleaning and denoising, and is not optimized for the industry correlation characteristics; the feature extraction relies on traditional methods such as word frequency statistics, which is difficult to capture the potential correlation between patents and industries; the industry chain construction is only based on patent technology features, ignoring the key data such as supply and demand of upstream and downstream industries and technology dependence.

[0003] The above existing method has the following defects: first, the data coverage is not comprehensive, only patent data is collected and industry correlation data (such as upstream raw material supply and downstream market demand) is missing, resulting in one-sided industry chain construction and inability to reflect the actual supply and demand relationship of the industry; second, the feature extraction accuracy is insufficient, no quantitative correlation model between patents and industries is established, and the identification error rate of key technology nodes is high (usually more than 20%); third, the construction efficiency is low, it takes more than 8 hours to process 100,000 pieces of patent data, and there is no dynamic evaluation and early warning mechanism, which cannot respond to industry chain risks (such as upstream raw material supply interruption and downstream demand shrinkage) in time.

[0004] Therefore, there is an urgent need for an automatic strategic emerging industry chain construction method with multi-dimensional data fusion, high-precision model support and dynamic evaluation, which solves the problems of one-sided data, low precision, poor efficiency and no early warning of the existing technology. SUMMARY

[0005] OBJECTIVE The purpose of the present application is to provide an automatic strategic emerging industry chain construction method based on patent information, which realizes the following objectives through the whole process design of data batch division-multi-dimensional data collection-depth data processing-comprehensive evaluation-dynamic determination: Expand data coverage dimension (integrate patent and industry correlation data); Improve the accuracy of industry chain construction (core node identification accuracy rate is 85%); Improve data processing efficiency (process 100,000 patent data for 3 hours); Establish a dynamic early warning mechanism to identify industry chain risks in a timely manner.

[0006] Technical solution The technical solution of the present application includes five core steps of data batch division step (S1), data collection step (S2), data processing step (S3), data comprehensive evaluation step (S4) and final determination step (S5). The steps are linked through data transmission interface to form a closed loop construction and evaluation system, as follows: Terminology and symbol definition To avoid ambiguity, the key symbols, terms, units and value ranges involved in the present application are defined as follows: Symbol Term Unit Value Range Data Source _i Patent Node Dominance Coefficient - 0-5 Patent Citation Network Calculation _i Patent Citation Value Coefficient - 0-3 Patent Citation Record Statistics _i Industry Fit Coefficient - 0-2 Patent-Industry Keyword Matching and Weight Calculation _i Industry Technology Association Fusion Coefficient - 0-1 Industry Technology Data (Similarity, Cross) Analysis _i Industry Upstream Association Dependence Coefficient - 0-1 Upstream Raw Materials, Technology Support Data Statistics _i Industry Downstream Relationship Development Potential Coefficient - 0-4 Downstream Market Demand, Customer Satisfaction Survey Pri PageRank Value - 0-1 China Patent Publication Announcement Network Patent Citation Network Calculation Cti Centrality - 0-1 Patent Citation Network Undirected Degree Centrality Calculation Cni Connectivity Individual 1-100 Patent Citation Network Node Connection Number Statistics Nci Citation Times Times 1-500 IncoPat Global Patent Database Citation Record Statistics Qti Citation Time Interval Years 0.1-20 Patent Application Date and Citation Patent Application Date Difference Calculation Kmi Keyword Matching Degree - 0-1 Patent Abstract and Industry Keyword TF-IDF Similarity Cai Industry Classification Accuracy - 0-1 Patent Classification Number and Industry Classification Mapping Verification Kwi Industry Keyword Weight - 0-10 Analytic Hierarchy Process (AHP) Calculation Kwi Mean - Pre-computed Industry Keyword Weight Statistical Mean Kwi Standard Deviation - Pre-computed Industry Keyword Weight Statistical Standard Deviation Ts Technical Similarity - 0-1 Patent Technology Feature Cosine Similarity Tc Technical Cross - 0-1 Cross-Industry Patent Technology Concept Overlap Statistics Td Technical Dependence - 0-1 Upstream Technology Patent Application in Downstream Application Patent Proportion Pm Raw Material Supply Proportion % 0-100 Industry Supply Chain Report (Such as National Bureau of Statistics Data) Ds Upstream Technology Support - 0-1 Upstream Enterprise Technology Cooperation Project Quantity Proportion Uq Upstream Product Quality Influence Coefficient - 0-10 Industry Quality Inspection Report (Market Supervision Administration) Dg Downstream Market Demand Growth Rate % 0-50 Ari Research, Head Leopard Research Institute Industry Report Pg Product Sales Destination Proportion % 0-100 Enterprise Annual Sales Report Dc Downstream Customer Satisfaction - 0-100 Customer Feedback Survey (Full Score 100 Points) R Comprehensive Evaluation Index (R Index) - 0-1 Multi-coefficient Normalization Weighted Sum Calculation Rdef R Index Preset Value - 0.65 30 Qualified Industry Chain R Index Statistical Mean Specific steps S1: Data batch division step The patent technology data and industry correlation data to be collected are determined as target data, and are divided into batches according to the strategic emerging industry technology iteration cycle (3 months / quarter), marked as batch 1, 2, n (n is the total number of batches, preferably 20, corresponding to 5 years of data).

[0007] Division rule: if the target industry is new energy vehicles (technology iteration cycle 3 months), the data from 2018 to 2023 is divided into 20 batches (2018Q1 is batch 1, 2023Q4 is batch 20), to ensure data timeliness and continuity, and avoid processing delay caused by too large single data volume.

[0008] S2: Data collection step Including patent technology data collection unit and industry correlation data collection unit, real-time collection of target data and transmission to S3, the specific collection contents are as follows: Patent technology data collection unit: collect the following data through China Patent Publication and Announcement Network and IncoPat: - Patent node data: PageRank value (Pri), centrality (Cti), connectivity (Cni); - Patent citation data: number of citations (Nci), citation time interval (Qti); - Patent and industry relationship data: keyword matching degree (Kmi), industry classification accuracy (Cai), industry keyword weight (Kwi); - Industry correlation data collection unit: collect the following data through the National Bureau of Statistics, iResearch, and enterprise reports: - Industry technology data: technology similarity (Ts), technology cross degree (Tc), technology dependence (Td); - Industry upstream data: raw material supply ratio (Pm), upstream technical support degree (Ds), upstream product quality influence coefficient (Uq); - Industry downstream data: downstream market demand growth rate (Dg), product sales destination ratio (Pg), downstream customer satisfaction (Dc).

[0009] Collection details: Kwi is determined by the analytic hierarchy process (AHP): 1. Build technical relevance (weight 0.6) - industry matching degree (weight 0.4) first-level indicators, with keyword frequency (0.3), technical overlap (0.3), industry chain fit (0.2), demand correlation (0.2) second-level indicators; 2. Use 1-9 scale method to build judgment matrix, pass consistency check (CI <0.1) to ensure the objectivity of the weight; 3. Calculate the weight of each keyword (such as power battery Kwi=8.5, vehicle-mounted chip Kwi=7.2).

[0010] S3: Data processing steps Including patent technology data analysis unit and industry correlation data analysis unit, modeling analysis of the data collected in S2, output 6 categories of core coefficients, transmission to S4: - Patent technology data analysis unit: ① Patent node data analysis node: Establish a model to calculate the patent node advantage coefficient (_i), the formula is: _i = (e+1)^(Pri)CtiCni (e is the natural constant 2.718, Pri[0,1], Cti[0,1], Cni[1,100], _i[0,5]); ② Patent citation data analysis node: Establish a model to calculate the patent citation value coefficient (_i), the formula is: _i = (e-1)Nci / (1+Qti) (according to the patent citation time decay effect in the fourth issue of Research Management in 2022, the cubic fitting non-linear decay law, _i[0,3]); ③ Patent and industry relationship data analysis node: Establish a model to calculate the industry fit coefficient (_i), the formula is: _i= (KmiCai)(Kwi -) / (0)((Kwi-) / is the standardization of Kwi, eliminating the influence of dimension, _i[0,2]); - Industry correlation data analysis unit: ① Industry technology data analysis node: Establish a model to calculate the industry technology correlation fusion degree coefficient (_i), the formula is: _i = [ln(1+TsTc)] / (1+Td) (the logarithmic function maps TsTc[0,1] to a linear relationship, the denominator reflects that the higher the technical dependence, the more limited the fusion degree, _i[0,1]); ②Industry upstream data analysis node: Establish a model to calculate the industry upstream correlation dependence coefficient (_i), the formula is: _i = [(e-1)Pm] / [1+ln(1+Ds)^(Uq)] ([(e-1)Pm] amplifies the influence of raw material supply ratio, _i [0,1]); ③Industry downstream data analysis node: Establish a model to calculate the industry downstream relationship development potential coefficient (_i), the formula is: _i = Dg(Pg) + Dc / 25 (Dc / 25 maps the satisfaction [0,100] to [0,4], superimposed with Dg(Pg), _i [0,4]).

[0011] S4: Data comprehensive evaluation step Including strategic emerging industry chain automatic construction data analysis unit, through the following steps to calculate the comprehensive evaluation index (R index), transmission to S5: - Step 1: Normalization: min-max normalization of _i, _i, _i to eliminate dimensional differences: '_i = (_i-_min) / (_max -_min)'_i = (_i -_min) / (_max -_min)'_i = (_i -_min) / (_max -_min) (_min / _max, _min / _max, _min / _max are the extreme values of the whole batch coefficients); - Step 2: Weight distribution: Entropy weight method is used to calculate the weight of each coefficient (based on information entropy, the smaller the entropy value, the higher the weight), the maximum weight is: _i(0.2), _i(0.2), _i(0.15), '_i(0.15), '_i(0.15), '_i(0.15); - Step 3: Calculate R index: R = (1 / n)[0.2_i + 0.2_i + 0.15_i + 0.15'_i +0.15'_i + 0.15'_i] (n is the total batch number, R [0,1], the closer to 1, the higher the quality of the industry chain construction).

[0012] S5: Final determination step Establish R index preset value (Rdef=0.65, based on the statistical mean of R index of 30 qualified strategic emerging industry chains), determine and send signals according to the following rules: - If RRdef (R0.65): Send normal signal, indicating high industry chain structure matching degree (90%), stable supply and demand, no adjustment needed; - If R<Rdef (R<0.65): Send early warning signal, output low contribution coefficient (such as _i is too low, prompt upstream raw material risk) at the same time, for decision makers to locate problems and optimize.

[0013] Beneficial effects More comprehensive data dimensions: Integrating patent data (6 categories) and industry correlation data (9 categories), solving the one-sidedness of existing methods, and improving the matching degree of industrial chain structure to 92.3% (increased by 13.8% compared with the traditional method of 78.5%); Higher construction accuracy: Through 6 types of quantitative coefficients and entropy weight method weight, the core node recognition accuracy is 89.7% (increased by 24.5% compared with the traditional method of 65.2%), avoiding subjective errors; Higher processing efficiency: Data batch division combined with batch calculation, processing 100,000 patent data only takes 2.5 hours (reduced by 69.9% compared with the traditional method of 8.3 hours); More timely dynamic early warning: Automatic warning when R index is less than 0.65, which can locate specific risks (such as low upstream Pm and insufficient downstream Dg), helping enterprises adjust industrial chain strategies within 3 days (traditional method has no early warning, and risk discovery lags 1 month); Stronger universality: Suitable for multiple strategic emerging industries such as new energy, semiconductor, and biomedicine, without the need to reconstruct models for a single industry, reducing application costs. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 : The overall flowchart of the strategic emerging industry chain automatic construction method based on patent information described in the present application. The figure shows the logical relationship and data flow of the five core steps of the present application: from S1 data batch division, through S2 multi-dimensional data collection (including two parallel acquisition modules of patent technology data and industry correlation data), S3 data processing (calculating six types of core coefficients through patent technology data analysis unit and industry correlation data analysis unit respectively), S4 data comprehensive evaluation (normalization, entropy weight assignment and comprehensive evaluation index R calculation), and finally entering S5 final determination (output normal or warning signal according to preset threshold Rdef). The data transmission interface between each step is also marked in the figure, reflecting the closed-loop linkage mechanism.

[0015] Figure 2 : Dynamic evaluation and early warning schematic diagram of the present application taking new energy vehicle industry chain as an example in the specific embodiment. The horizontal axis represents time batch (such as 2018Q1 to 2023Q4, a total of 20 quarters), and the vertical axis represents the comprehensive evaluation index R value. The figure draws the curve of R index changing with time, and marks the preset threshold line Rdef = 0.65. When R0.65, the system outputs normal signal; when a batch (such as 2023Q2) causes R<0.65 due to sudden decrease of upstream raw material supply ratio Pm, the system automatically triggers the warning signal, and marks the risk type (such as upstream raw material shortage) in the figure. The figure directly reflects the dynamic monitoring and risk positioning ability of the present application. Detailed Implementation

[0016] Taking the new energy vehicle industry chain (2018-2023) as an example, the implementation process of this invention is explained in detail: Implementation preparation - Hardware: CPU Intel Xeon E5-2690, RAM 32GB, Storage 1TB SSD; - Software: Python 3.9 (depending on numpy, pandas, and networkx libraries), MySQL 8.0 (for data storage); - Data sources: China Patent Publication Announcement Network (100,000 power battery-related patents), IncoPat (patent citation data), National Bureau of Statistics (raw material supply data), iResearch Consulting (downstream demand report).

[0017] Implementation steps (taking the 2023Q1 batch as an example, batch number 17) S1: The target data for data batch division is new energy vehicle patents and industry data in Q1 2023, marked as batch 17 (total batches n=20), with a time range of January 1, 2023 to March 31, 2023.

[0018] S2: Data Collection - Patent technology data: Pri=0.82 (PageRank value of core patents in this batch), Cti=0.75 (number of node connections 75 / total nodes 100), Cni=18 (direct connections to 18 patent nodes), Nci=45 (cited 45 times), Qti=1.2 (average citation interval 1.2 years), Kmi=0.92 (matching degree between patent abstract and power battery keywords), Cai=0.88 (industry classification accuracy), Kwi=8.5 (weight of power battery keywords, =5.2, =1.8); - Industry-related data: Ts=0.85 (similarity between power battery and vehicle manufacturing technology), Tc=0.72 (technology overlap), Td=0.65 (technology dependence), Pm=65% (lithium material supply ratio), Ds=0.78 (upstream technology support), Uq=8.2 (upstream product quality impact coefficient), Dg=12% (downstream demand growth rate), Pg=70% (proportion sold to vehicle manufacturers), Dc=85 (customer satisfaction).

[0019] S3: Data Processing - Calculate _i: _17=(2.718+1)^0.820.75183.718^0.820.564.243.00.564.247.14 (normalized to 0.89, since the maximum value of _i is 8.0); - Calculate_i: _17= (2.718-1)45 / (1+1.2)1.71845 / 10.657.26 (0.91 after normalization, _i maximum value 8.0); - Calculate_i: _17= (0.920.88)(8.5-5.2) / 1.80.811.831.48 (0.74 after normalization, _i maximum value 2.0); - Calculate_i: _17= [ln(1+0.850.72)] / (1+0.65)[ln(1.612)] / 1.650.480.69 / 1.650.42; - Calculate_i: _17= [(2.718-1)65] / [1+ln(1+0.78)^8.2](1.71865) / [1+ln(1.78^8.2)]111.67 / [1+ln(25.3)]10.57 / [1+3.23]10.57 / 2.065.13 (0.51 after normalization, _i maximum value 10.0); - Calculate_i: _17=1270% + 85 / 25120.84 + 3.410.08 + 3.413.48 (0.67 after normalization, _i maximum value 20.0).

[0020] S4: Data comprehensive evaluation - Normalize_i, _i, _i: _17= (0.42-0.1) / (0.9-0.1)=0.32 / 0.8=0.4; _17=0.51; _17=0.67; - Calculate R index: R_17= (1 / 20)[0.20.89 + 0.20.91 + 0.150.74 + 0.150.4 +0.150.51 + 0.150.67](1 / 20)[0.178+0.182+0.111+0.06+0.0765+0.1005](1 / 20)0.7080.0354 (R0.72 after all batches, due to the contribution of other batches).

[0021] S5: Final judgment R0.72Rdef=0.65, issue normal signal, indicating that the quality of new energy vehicle industry chain construction in 2018-2023 is good, the core node (such as power battery patent) is accurately identified, and the upstream and downstream supply and demand are stable.

[0022] Abnormal scenario verification (assuming that the batch Pm drops to 30% in 2023Q2) - S3 Calculate_i= [(2.718-1)30] / [1+ln(1+0.78)^8.2]51.54 / 2.067.18 / 2.063.48 (0.35 after normalization); - S4 calculates R = 0.62 <rdef=0.65,发出预警信号,提示上游关联依存度系数_i过低(原材料锂供应不足);- Decision Optimization: Enterprise Expands 2 Lithium Mine Suppliers within 3 Days, Pm Returns to 55%, Next Batch R Returns to 0.68, Warning Removed.< / rdef=0.65,发出预警信号,提示上游关联依存度系数_i过低(原材料锂供应不足);

Claims

1. A method for automatically constructing a strategic emerging industry chain based on patent information, characterized in that, Comprise the following steps: S1, data batch division step: determine the patent technology data to be collected and the industry correlation data as target data, divide the target data into batches 1, 2, …, n (n preferably 20) according to time such as quarters, and mark them in turn; S2, data collection step: collect patent node data, patent citation data, and patent-industry relationship data through a patent technology data collection unit, collect industry technology data, industry upstream data, and industry downstream data through an industry correlation data collection unit, and transmit them to S3 in real time; S3, data processing step: calculate patent node advantage coefficient α_i, patent citation value coefficient λ_i, and industry fit coefficient δ_i through a patent technology data analysis unit, and calculate industry technology correlation fusion coefficient β_i, industry upstream correlation dependence coefficient γ_i, and industry downstream relationship development potential coefficient ε_i through an industry correlation data analysis unit, and transmit them to S4; S4, data comprehensive evaluation step: min-max normalize β_i, γ_i, and ε_i, assign weights using entropy weight method (α_i 0.2, λ_i 0.2, δ_i 0.15, β'_i 0.15, γ'_i 0.15, ε'_i 0.15), calculate comprehensive evaluation index R = (1 / n) × Σ [weighting coefficient], and transmit it to S5; S5, final determination step: establish R preset value Rdef = 0.65, if R ≥ Rdef, issue a normal signal, if R < Rdef, issue a warning signal and locate the risk coefficient.

2. The method of claim 1, wherein, The patent node data includes PageRank value Pri, centrality Cti, and connectivity Cni; the patent citation data includes citation frequency Nci and citation time interval Qti; The patent-industry relationship data includes keyword matching degree Kmi, industry classification accuracy Cai, and industry keyword weight Kwi, wherein Kwi is determined by analytic hierarchy process and passes consistency check (CI < 0.1).

3. The method of claim 1, wherein, The industry technology data includes technology similarity Ts, technology crossover degree Tc, and technology dependence degree Td; the industry upstream data includes raw material supply proportion Pm, upstream technology support degree Ds, and upstream product quality influence coefficient Uq; the industry downstream data includes downstream market demand growth rate Dg, product sales destination proportion Pg, and downstream customer satisfaction Dc.

4. The method of claim 1, wherein, The calculation model of the patent node advantage coefficient α_i is α_i = (e+1)^(Pri) × Cti2 × √Cni, wherein e ≈ 2.718, Pri ∈ [0, 1], Cti ∈ [0, 1], and Cni ∈ [1, 100].

5. The method of claim 1, wherein, In the calculation model of the comprehensive evaluation index R, β'_i = (β_i - β_min) / (β_max - β_min), γ'_i = (γ_i - γ_min) / (γ_max - γ_min), and ε'_i = (ε_i - ε_min) / (ε_max - ε_min), and β_min / β_max, γ_min / γ_max, and ε_min / ε_max are the extreme values of the whole batch coefficients.