Priority development technology identification method based on multi-source information

Through semantic analysis and modeling of multi-source data, combined with the use of knowledge graphs and objective functions, the limitations of traditional technical evaluation methods are solved, accurate prediction of technological development trends and scientific identification of priority development technologies are achieved, and long-term stability of technological development and market adaptability are ensured.

CN120124636APending Publication Date: 2025-06-10INST POLICY & MANAGEMENT CHINESE ACADEMY SCI +1

Patent Information

Application Number
CN202510194816.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology evaluation methods rely on a single information source or simple indicators, lack comprehensiveness and forward-lookingness, resulting in the neglect of potential technologies, and over-investigation of popular but insufficient competitiveness technologies. At the same time, the data formats and semantic expressions between different information sources are different, making it difficult to form comprehensive and accurate knowledge.

Method used

The priority development technology identification method based on multi-source information is adopted, and multi-source data composed of academic papers, patents, government reports and market application cases are obtained, and the relationship between technology development and time characteristics is used to determine the relationship, and the technology development is modeled and the technology development is predicted. Then, select keywords whose technology life cycle is in the growth stage as the target node, build a knowledge graph for technological development, obtain a set of technological development paths based on the target node traversal, and build an objective function based on technological improvement and market returns, and select the optimal path from the technological development path.

Benefits of technology

Accurate prediction of technological development trends has been achieved, ensuring that technological development focuses on the direction of change, and the chosen technological development path can balance market returns and technological improvement, adapt to different development stages and strategic goals, and improve the comprehensiveness and forward-looking nature of technical evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124636A_ABST
    Figure CN120124636A_ABST
Patent Text Reader

Abstract

The invention discloses a preferential development technology identification method based on multi-source information, and the method comprises the steps: obtaining multi-source data containing academic papers, patents, government reports and market application cases, mining the multi-source data through semantic analysis, determining the correlation between technology development and time, carrying out the modeling, and predicting the technology development trend, the method comprises the steps of selecting keywords according to a technology development trend, comparing historical hotspot technologies to calculate semantic similarity, determining a technology life cycle, selecting keywords in a growth period as target nodes, constructing a knowledge graph of the keywords, traversing the graph based on the target nodes, obtaining a technology development path set, and constructing a target function according to technology improvement and market income. And an optimal technology development path is screened out. According to the method, by integrating multi-source information and applying semantic analysis and knowledge graph technologies, the technology development trend can be comprehensively and accurately predicted, the efficiency and accuracy of technology development decision making are improved, the decision making risk is reduced, and meanwhile good interpretability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information strategic planning and processing, and particularly to a method for identifying priority development technologies based on multi-source information. Background Art

[0002] With the rapid development of technology, the technologies in various fields have shown an explosive growth. How to screen out the technologies with potential for priority development from numerous technologies has become a key issue.

[0003] In traditional technology evaluation methods, they often rely on a single information source or simple indicators. For example, only based on short-term market demand or a single academic research direction to judge the priority of technologies, lacking comprehensiveness and foresight. This limitation has led to many potential technologies being ignored, while some popular but possibly lacking long-term competitiveness technologies have received excessive resource investment. At the same time, the data formats and semantic expressions among different information sources are diverse, lacking effective integration means, and it is difficult to form a comprehensive and accurate understanding of technology development. For example, the professional terms in academic papers, the legal descriptions in patents, the macro-guidance in government reports and the actual feedback in market application cases are difficult to analyze synergistically, resulting in biases in technology evaluation results.

[0004] A method for identifying priority development technologies based on multi-source information, which can integrate multi-source information, comprehensively consider various factors and accurately predict the technology development trend, so as to scientifically determine the priority development technologies and promote the efficient progress of science and technology and industry. Summary of the Invention

[0005] The object of the present invention is to provide a method for identifying priority development technologies based on multi-source information.

[0006] To achieve the above object, the present invention is implemented according to the following technical solution:

[0007] The first aspect of the present invention provides a method for identifying priority development technologies based on multi-source information, including:

[0008] Step S1, obtaining multi-source data composed of academic papers, patents, government reports and market application cases;

[0009] Step S2, applying semantic analysis to the multi-source data to determine the relationship between technology development and time characteristics, and modeling the technology development in combination with the relationship to predict the technology development trend;

[0010] Step S3, selecting keywords based on the technology development trend, determining the technology life cycle of the keywords according to the semantic similarity between the keywords and historical hot technologies, and selecting the keywords with the technology life cycle in the growth stage as the target nodes;

[0011] Step S4: Conduct correlation analysis among the keywords, construct a knowledge graph of technological development, and traverse based on the target node in the knowledge graph to obtain a set of technological development paths;

[0012] Step S5: Construct an objective function according to technological improvement and market returns, and select the optimal technological development path from the set of technological development paths according to the objective function.

[0013] As a further method, a method for applying semantic analysis to the multi-source data to determine the relationship between technological development and time characteristics and modeling technological development in combination with the relationship includes:

[0014] Perform text cleaning on the multi-source data, including identifying and deleting noise data such as duplicate information, garbled characters, and incomplete records, and convert the cleaned text data into word vectors using a word vector model;

[0015] Input the word vectors into a hybrid neural network architecture combining a convolutional neural network and a recurrent neural network, where the convolutional neural network is used to capture local features of the text, and the recurrent neural network is used to identify long-sequence dependencies in the text to obtain a set of text features;

[0016] Classify the set of features through a time series clustering algorithm, and for each cluster, obtain the technological innovation frequency of technological development;

[0017] Construct a technological development feature matrix based on the technological innovation frequency of technological development and use it as a training set to train a support vector regression model to complete the modeling of technological development.

[0018] As a further method, the method for selecting keywords based on the technological development trend includes:

[0019] Obtain the curve of the technological innovation frequency in the technological development trend, and use variational mode decomposition to decompose the curve to obtain multiple representative modal functions;

[0020] Calculate the influence weight of the vocabulary based on multiple representative modal functions corresponding to the vocabulary. The expression is:

[0021]

[0022] where, W i is the influence weight of the i-th vocabulary, n is the number of representative modal functions, m is the number of vocabularies, α k and β k are respectively the word frequency weight adjustment coefficient and the inverse document frequency weight adjustment coefficient of the k-th representative modal function, TF i,k is the word frequency of the i-th vocabulary in the k-th representative modal function, IDF i,kThe inverse document frequency of the i-th term in the k-th representation modality function, IMF k is the weight of the k-th representation modality function;

[0023] Select terms with influence weights higher than the average as keywords.

[0024] As a further method, the method of determining the technology life cycle of keywords according to the semantic similarity between keywords and historical hot technologies and selecting keywords with technology life cycles in the growth stage as target nodes includes:

[0025] Obtain a document set containing keywords and historical hot technologies from multi-source data;

[0026] Calculate the semantic similarity between keywords and historical hot technologies. The expression is:

[0027]

[0028] where D i and D j are the document sets corresponding to keywords and historical hot technologies respectively, α is the balance coefficient of feature terms, h and g are the numbers of feature terms in document sets D i and D j respectively, is the eigenvalue of the k-th feature term in document set D i , is the eigenvalue of the l-th feature term in document set D j , M kl is and the semantic similarity matrix between them, TF i and TF j are the TF-IDF values of document sets D i and D j respectively, M kk is in document set D i the normalization factor, M ll is in document set D j the normalization factor;

[0029] Normalize the semantic similarity from 0 to 1. If the semantic similarity shows an upward trend during the observation period and the average value of the semantic similarity with historical hot technologies is less than 0.4, then the technology life cycle of this keyword is in the growth stage, and this keyword is used as the target node.

[0030] As a further method, the method of performing association analysis between the keywords and constructing a knowledge graph of technology development includes:

[0031] Statistically analyze the co-occurrence frequencies of different keywords in the same sentence, the same paragraph, and different text types as the association strength between keywords.

[0032] Use text mining techniques to extract the co-occurrence relationships between keywords from multi-source data as the connection relationships between nodes. Take the keywords as nodes and construct the edges between nodes based on the association strength to build a knowledge graph.

[0033] As a further method, the method of traversing in the knowledge graph based on the target node to obtain a set of technology development paths includes:

[0034] Analyze the association strength between nodes in historical technology development paths and calculate the average value as the association strength threshold.

[0035] Starting from the target node, traverse adjacent nodes according to the association strength between nodes. Specifically: if the association strength value is greater than the association strength threshold and the adjacent node is not the target node, then continue to traverse the adjacent node.

[0036] For the traversal results of each target node, sort them in descending order according to the technology development potential of the nodes as the technology development priority order. Form technology development paths by traversing the nodes to obtain a set of technology development paths, where the calculation formula for the technology development potential of a node is:

[0037]

[0038] where P i is the technology development potential of the i-th node, Z is the normalization constant, ψ is the number of adjacent nodes of the i-th node, is the semantic association weight coefficient, S ij is the semantic similarity between the i-th node and the j-th node, M i is the maximum semantic similarity between the i-th node and its adjacent nodes, C ij is the association strength between the i-th node and the j-th node, D i is the minimum association degree between the i-th node and its adjacent nodes, F ij is the sum of the frequencies of occurrence of the i-th node and the j-th node in multi-source data, F i is the sum of the frequencies of occurrence of all node pairs in the knowledge graph, and δ is the frequency influence index.

[0039] As a further method, the method of constructing an objective function according to technology improvement and market return and selecting the optimal technology development path from the set of technology development paths includes:

[0040] Construct a first objective function with the maximum technology improvement, and the expression is:

[0041]

[0042] Among them, x is the technical parameter vector, q is the number of performance indicators, ω i is the weight of the i-th performance indicator, P i (x) is the actual performance value of the i-th performance indicator, P i,base is the baseline value of the i-th performance indicator, i.e., the lowest performance level of the historical technology, P i,max (x) is the maximum theoretical value of the i-th performance indicator, a is the non-linear influence coefficient for adjusting performance improvement, p is the number of resource consumption indicators, v j is the weight of the j-th resource consumption, C j (x) is the actual value of the j-th resource consumption, b is the non-linear influence coefficient for adjusting resource consumption;

[0043] Construct the second objective function with the maximum market return, and the expression is:

[0044]

[0045] Among them, D is the number of different stages in the technology development path, ξ i is the market potential of the i-th technology development stage, λ i is the market share of the i-th technology development stage, t i is the unit profit of the i-th technology development stage, R i is the market risk probability of the i-th technology development stage, F is the number of R & D cost indicators in the technology development path, Q j is the value of the j-th R & D cost indicator, W j is the weight of the j-th R & D cost;

[0046] According to the strategic focus of technology development, determine the relative weights of technology improvement and market return in the overall decision-making. Based on the relative weights, construct a comprehensive evaluation function and select the technology development path with the optimal comprehensive evaluation.

[0047] Compared with the existing technology, the embodiments of the present invention have at least the following advantages or beneficial effects:

[0048] (1) Through the comprehensive analysis and modeling of multi-source data, the present invention can accurately and objectively predict the development trend of technology. Compared with the methods relying on expert experience and single data sources, it can better capture the dynamic changes of technology development and improve the accuracy and reliability of prediction;

[0049] (2) The present invention selects keywords based on the technological development trend, combines the semantic similarity between the keywords and historical hot technologies, determines the technological life cycle of the keywords, and analyzes the technologies in the growth stage, so that the technological development focuses on the direction of technological change;

[0050] (3) The present invention constructs an objective function by combining technological improvement and market return, selects the optimal path from the set of technological development paths, ensures that the selected technological development path can balance market return and technological improvement, is conducive to long-term stable development, and can adapt to different development stages and strategic goals. Description of the Drawings

[0051] Figure 1 It is a flowchart of the steps of a method for identifying technologies with priority development based on multi-source information in an embodiment of the present invention. Detailed Embodiment

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] Refer to Figure 1 As shown, the present invention provides a method for identifying technologies with priority development based on multi-source information, including:

[0054] Step S1, obtaining multi-source data composed of academic papers, patents, government reports, and market application cases;

[0055] In actual evaluation, in the field of smart home lighting technology, using a professional academic database, 847 academic papers on intelligent control algorithms for intelligent lighting systems, LED light source control algorithms, etc. in the past five years were collected. On the global patent retrieval platform, more than 610 relevant patents and 24 government reports were found. Among them, it was proposed to increase the penetration rate of smart home lighting systems in newly built residences in the next three years, promote the integration of intelligent lighting with technologies such as the Internet of Things and big data, and provide R & D subsidies and demonstration project construction funds. And starting from collecting market application cases, the sales situation and user evaluation feedback of lighting products of mainstream smart home brands in the market were investigated, and project reports were collected.

[0056] Step S2, performing semantic analysis on the multi-source data, determining the relationship between technological development and time characteristics, and modeling the technological development in combination with the relationship to predict the technological development trend;

[0057] It should be noted that multi-source data includes academic papers, patents, government reports, and market application cases, with complex and professional texts. Semantic analysis uses natural language processing technology to extract key technical terms such as "intelligent dimming" and "human body sensing lighting", identify the expression differences of these terms in different text types and at different time nodes, and determine the internal relationship between technology development and time by analyzing these differences. Based on this relationship modeling, machine learning algorithms are used to construct a prediction model, laying a foundation for subsequent accurate judgment of technology trends and assisting enterprise decisions;

[0058] In the actual evaluation, after preprocessing, the GloVe word vector model is used to convert the cleaned text data into word vectors, generating approximately 30,000 word vectors. The word vectors are input into an architecture combined by a convolutional neural network and a long short-term memory network. Among them, the CNN is configured with 3 convolutional layers, with 64 convolutional kernels in each layer, which is good at capturing local key information in the text and identifying key parameters and material names in standard documents; the LSTM has 2 hidden layers, with 128 neurons in each layer, to mine the internal logic of the long sequence of the text, and completely outline the development context of the smart home lighting technology from basic research, product development to market promotion, and finally obtain a text feature set with more than 2,500 features;

[0059] In the actual evaluation, a density-based time series clustering algorithm is adopted to classify the feature set according to the technology innovation timeline, clustering it into 8 clusters, specifically: intelligent dimming control algorithm, intelligent lighting human body sensing interaction, intelligent lighting and Internet of Things integration, intelligent lighting energy-saving technology, intelligent lighting light environment simulation and health, intelligent lighting voice control technology, intelligent lighting scene customization and automation, and intelligent lighting distributed control system. Among them, for the "intelligent dimming control algorithm" cluster, the annual growth rate of the number of relevant academic papers published in the past 3 years reaches 70%, and the annual increase in patent applications is 65%. A feature matrix of the development of intelligent lighting technology is constructed as a training set to deeply train the support vector regression model.

[0060] Step S3: Select keywords based on the technological development trend, determine the technology life cycle of the keywords according to the semantic similarity between the keywords and historical hot technologies, and select the keywords with the technology life cycle in the growth stage as the target nodes;

[0061] It should be noted that the technical life cycle of a keyword is determined based on the semantic similarity between the keyword and historical hot technologies. As technology advances, if the semantic similarity shows an upward trend, it indicates that the technology is absorbing historical experience, integrating some advantages of traditional technologies, and beginning to expand applications and optimize performance. For example, the emerging "intelligent dimming" technology in smart home lighting gradually approaches some principles of traditional dimming technology, but also expands intelligent interaction features. At this time, it can be judged that it is in the growth stage and has great development potential. When the semantic similarity reaches a high and stable state, it reflects that the technology has fully integrated past experience and is approaching maturity, with limited room for subsequent innovation, and enters the mature stage. Once the similarity turns downward, it indicates that the technology is being replaced by a more novel solution and enters the decline stage. In this way, the stage where the technology is located can be accurately positioned, providing a key basis for subsequent decisions.

[0062] In actual evaluation, extract the technology innovation frequency curve in the technology development trend, and use variational mode decomposition (VMD) to decompose the curve into 5 representative mode functions. For "intelligent dimming", the word frequencies in each representative mode function are 10, 16, 8, 12, 10 in turn, and the inverse document frequencies correspond to 2.8, 3.2, 2.2, 2.6, 2.8. Combining the weights of each mode function, the calculated influence weight is 0.72; screen the words with influence weights exceeding the average threshold of 0.45, and obtain 40 keywords including "intelligent dimming", "human body sensing lighting", "linkage between intelligent lighting and the Internet of Things", etc. Extract the document materials containing the keywords and historical hot quantum technologies (traditional incandescent lamp dimming technology, early energy-saving lamp technology) from multi-source data, calculate the semantic similarity, and conduct life cycle judgment to obtain the target nodes, including: intelligent dimming, human body sensing, Internet of Things linkage, dimming algorithm, and lighting light health simulation.

[0063] Step S4: Conduct correlation analysis between the keywords, construct a knowledge graph of technology development, and traverse based on the target nodes in the knowledge graph to obtain a set of technology development paths;

[0064] It should be noted that the reason for conducting correlation analysis between keywords is that each keyword represents different technical points and they do not exist in isolation. For example, in the field of smart home lighting, "intelligent dimming" and "human body sensing lighting" may often appear together in relevant texts, indicating that there is a connection between them, either in terms of functional cooperation or overlapping application scenarios. Quantify this correlation strength by statistically analyzing their co-occurrence frequencies in different text types, and then construct a knowledge graph with keywords as nodes and correlation strength as edges to visually present the internal connection network between technologies;

[0065] In the actual evaluation, the co-occurrence frequencies of different keywords are counted in the same sentence of academic papers, the same paragraph of patent specifications, and different text types such as various policy reports and market cases, so as to quantify the association strength between keywords. Among them, the co-occurrence frequency of "intelligent dimming" and "human body sensing lighting" is as high as 30 times, and the co-occurrence frequency of "intelligent lighting and Internet of Things linkage" and "intelligent lamp heat dissipation technology" is 6 times. According to the text mining algorithm, the keyword co-occurrence associations in multi-source data are obtained as the connections between the nodes of the knowledge graph. The keywords are used as nodes, and the edges between the nodes are constructed according to the association strength to obtain the knowledge graph. This graph covers 31 nodes and 143 edges; calculate the association strength between the nodes of the historical development path of smart home lighting technology, and obtain the average value as the association strength threshold, which is 0.15. Starting from "intelligent dimming", depth traversal is carried out on adjacent nodes in order according to the association strength between the nodes to obtain the technology development path set, including: (1) Intelligent dimming: intelligent lighting scene preset, intelligent lighting and intelligent speaker connection, intelligent lighting lamp layout optimization; (2) Intelligent lighting light environment simulation and health: intelligent eye protection dimming technology, intelligent lighting growth companion mode, intelligent lighting mobile phone control linkage; (3) Human body sensing lighting: intelligent lighting delay extinguishing function, intelligent lighting and door and window sensor linkage, intelligent lighting lamp self-cleaning assistance.

[0066] Step S5, construct an objective function according to technology improvement and market revenue, and select the optimal technology development path from the technology development path set according to the objective function.

[0067] In the actual evaluation, the enterprise is in the stage of equal emphasis on technology upgrading and market expansion. Set the technology improvement weight to 0.6 and the market revenue weight to 0.4. The optimal technology development path is: (3) Human body sensing lighting: intelligent lighting delay extinguishing function, intelligent lighting and door and window sensor linkage, intelligent lighting lamp self-cleaning assistance.

[0068] In this embodiment, the method of using semantic analysis for the multi-source data to determine the relationship between technology development and time characteristics and modeling the technology development in combination with the relationship includes:

[0069] Perform text cleaning on the multi-source data, including identifying and deleting noise data such as duplicate information, garbled codes, and incomplete records, and converting the cleaned text data into word vectors using a word vector model;

[0070] Input the word vectors into a hybrid neural network architecture combined with a convolutional neural network and a recurrent neural network. Among them, the convolutional neural network is used to capture the local features of the text, and the recurrent neural network is used to identify the long-sequence dependence relationships in the text to obtain the feature set of the text;

[0071] Classify the feature set through a time series clustering algorithm. For each cluster, obtain the technology innovation frequency of technology development;

[0072] Construct a technology development feature matrix based on the technology innovation frequency of technology development, and use it as a training set to train a support vector regression model to complete the modeling of technology development.

[0073] In this embodiment, the method for selecting keywords based on the technology development trend includes:

[0074] Obtain the curve of the technology innovation frequency in the technology development trend, and use variational mode decomposition to decompose the curve to obtain multiple representation mode functions;

[0075] Calculate the influence weight of the vocabulary based on the multiple representation mode functions corresponding to the vocabulary. The expression is:

[0076]

[0077] where W i is the influence weight of the i-th vocabulary, n is the number of representation mode functions, m is the number of vocabularies, α k and β k are the word frequency weight adjustment coefficient and inverse document frequency weight adjustment coefficient of the k-th representation mode function respectively, TF i,k is the word frequency of the i-th vocabulary in the k-th representation mode function, IDF i,k is the inverse document frequency of the i-th vocabulary in the k-th representation mode function, and IMF k is the weight of the k-th representation mode function;

[0078] Select the vocabulary with an influence weight higher than the average as the keyword.

[0079] In this embodiment, the method for determining the technology life cycle of keywords according to the semantic similarity between keywords and historical hot technologies, and selecting the keywords with the technology life cycle in the growth stage as the target nodes includes:

[0080] Obtain a document set containing keywords and historical hot technologies from multi-source data;

[0081] Calculate the semantic similarity between keywords and historical hot technologies. The expression is:

[0082]

[0083] where D i and D j are the document sets corresponding to keywords and historical hot technologies respectively, α is the balance coefficient of feature items, h and g are the numbers of feature items in the document sets D i and D j respectively, is the document set Di The eigenvalue of the k-th feature term in is the document collection D j The eigenvalue of the l-th feature term in M kl is and is the semantic similarity matrix between i and TF j are respectively the TF-IDF values of the document collections D i and D j in M kk is in the document collection D i in the normalization factor M ll is in the document collection D j in the normalization factor;

[0084] Normalize the semantic similarity from 0 to 1. If the semantic similarity shows an upward trend during the observation period and the average value of the semantic similarity with historical hot technologies is less than 0.4, then the technology life cycle of this keyword is in the growth stage, and this keyword is used as the target node.

[0085] In this embodiment, the method for performing association analysis on the keywords and constructing a knowledge graph of technology development includes:

[0086] Statistical co-occurrence frequencies of different keywords in the same sentence, the same paragraph, and different text types as the association strength between keywords;

[0087] Using text mining technology to extract co-occurrence relationships between keywords from multi-source data as the connection relationships between nodes, using keywords as nodes, and constructing edges between nodes based on the association strength to construct a knowledge graph.

[0088] In this embodiment, the method for traversing in the knowledge graph based on the target node to obtain a set of technology development paths includes:

[0089] Analyze the association strength between nodes in historical technology development paths and calculate the average value as the association strength threshold;

[0090] Starting from the target node, traverse adjacent nodes according to the association strength between nodes. Specifically: if the association strength value is greater than the association strength threshold and the adjacent node is not the target node, then continue to traverse the adjacent node;

[0091] For the traversal results of each target node, sort them in descending order according to the technology development potential of the nodes as the technology development priority order, form technology development paths by traversing the nodes, and obtain a set of technology development paths, where the calculation formula for the technology development potential of the nodes is:

[0092]

[0093] Among them, P i is the technological development potential of the i-th node, Z is the normalization constant, ψ is the number of adjacent nodes of the i-th node, is the semantic association weight coefficient, S ij is the semantic similarity between the i-th node and the j-th node, M i is the maximum semantic similarity between the i-th node and its adjacent nodes, C ij is the association strength between the i-th node and the j-th node, D i is the minimum association degree between the i-th node and its adjacent nodes, F ij is the sum of the frequencies of occurrence of the i-th node and the j-th node in multi-source data, F i is the sum of the frequencies of occurrence of all node pairs in the knowledge graph, and δ is the frequency influence index.

[0094] In this embodiment, the method of constructing an objective function according to technological improvement and market revenue and selecting an optimal technological development path from the set of technological development paths includes:

[0095] Construct a first objective function with the maximum technological improvement, and the expression is:

[0096]

[0097] Among them, x is the technological parameter vector, q is the number of performance indicators, ω i is the weight of the i-th performance indicator, P i (x) is the actual performance value of the i-th performance indicator, P i,base is the baseline value of the i-th performance indicator, i.e., the lowest performance level of the historical technology, P i,max (x) is the maximum theoretical value of the i-th performance indicator, a is the non-linear influence coefficient for adjusting performance improvement, p is the number of resource consumption indicators, v j is the weight of the j-th resource consumption, C j (x) is the actual value of the j-th resource consumption, and b is the non-linear influence coefficient for adjusting resource consumption;

[0098] Construct a second objective function with the maximum market revenue, and the expression is:

[0099]

[0100] Among them, D is the number of different stages in the technological development path, ξ i is the market potential of the i-th technological development stage, λ i is the market share of the i-th technological development stage, t iis the unit profit of the i-th stage of technological development, R i is the market risk probability of the i-th stage of technological development, F is the number of R & D cost indicators in the technological development path, Q j is the value of the j-th R & D cost indicator, W j is the weight of the j-th R & D cost;

[0101] According to the strategic focus of technological development, determine the relative weights of technological improvement and market returns in the overall decision-making. Based on the relative weights, construct a comprehensive evaluation function and select the technological development path with the optimal comprehensive evaluation.

[0102] The second aspect of the present invention also provides a system for determining priority development technologies based on multi-source information, including:

[0103] A data acquisition module for obtaining multi-source data composed of academic papers, patents, government reports, and market application cases;

[0104] A semantic modeling module for performing semantic analysis on the multi-source data to determine the relationship between technological development and time characteristics, and modeling technological development in combination with the relationship to predict technological development trends;

[0105] A keyword extraction module for selecting keywords based on the technological development trend, determining the technological life cycle of the keywords according to the semantic similarity between the keywords and historical hot technologies, and selecting the keywords with the technological life cycle in the growth stage as target nodes;

[0106] A knowledge graph module for performing correlation analysis between the keywords, constructing a knowledge graph of technological development, and traversing based on the target nodes in the knowledge graph to obtain a set of technological development paths;

[0107] A path screening module for constructing an objective function according to technological improvement and market returns, and selecting the optimal technological development path from the set of technological development paths according to the objective function.

[0108] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, as long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.

Claims

1. A method for identifying priority technologies based on multi-source information, characterized in that: The following steps are involved: Step S1, obtaining multi-source data consisting of academic papers, patents, government reports and market application cases; Step S2: applying semantic analysis to the multi-source data to determine the relationship between technology development and time characteristics, and modeling technology development based on the relationship to predict technology development trends; Step S3: Select keywords based on the technology development trend, determine the technology life cycle of the keywords according to the semantic similarity between the keywords and historical hot technologies, and select keywords whose technology life cycle is in the growth stage as target nodes; Step S4: performing association analysis between the keywords, constructing a knowledge graph of technology development, traversing the knowledge graph based on the target node, and obtaining a set of technology development paths; Step S5: construct an objective function according to technology improvement and market benefits, and select the optimal technology development path from the technology development path set according to the objective function.

2. The method for identifying priority technologies based on multi-source information according to claim 1, characterized in that: The method of applying semantic analysis to the multi-source data to determine the relationship between technology development and time characteristics, and modeling technology development based on the relationship, includes: Perform text cleaning on multi-source data, including identifying and deleting duplicate information, garbled characters, and incomplete records, and converting the cleaned text data into word vectors using the word vector model; The word vector is input into a hybrid neural network architecture that combines a convolutional neural network and a recurrent neural network. The convolutional neural network is used to capture the local features of the text, and the recurrent neural network is used to identify the long sequence dependencies in the text to obtain the feature set of the text. The feature set is classified by a time series clustering algorithm, and for each cluster, the frequency of technological innovation of technological development is obtained; Based on the technological innovation frequency of technological development, a technological development feature matrix is ​​constructed, and used as a training set to train the support vector regression model to complete the modeling of technological development.

3. The method for identifying priority technologies based on multi-source information according to claim 1, characterized in that: The method for selecting keywords based on the technology development trend includes: Obtain the curve of technological innovation frequency in the technological development trend, decompose the curve using variational mode decomposition, and obtain multiple characterizing mode functions; The influence weight of a word is calculated based on multiple representation modal functions corresponding to the word. The expression is: Among them, W i is the influence weight of the i-th word, n is the number of modal functions, m is the number of words, α k and β k are the term frequency weight adjustment coefficient and inverse document frequency weight adjustment coefficient of the kth modal function, TF i,k is the frequency of the i-th word in the k-th representation modal function, IDF i,k is the inverse document frequency of the i-th word in the k-th representation modal function, IMF k is the weight of the k-th characterization modal function; Filter words with influence weights higher than the average as keywords.

4. The method for identifying priority technologies based on multi-source information according to claim 1, characterized in that: The method of determining the technical life cycle of a keyword based on the semantic similarity between the keyword and the historical hot technology and selecting a keyword whose technical life cycle is in the growth stage as a target node includes: Obtain document collections containing keywords and historical hot technologies from multi-source data; Calculate the semantic similarity between keywords and historical hot technologies. The expression is: Among them, D i and D j are the document sets corresponding to keywords and historical hot technologies, α is the balance coefficient of feature items, h and g are the document sets D i and D j The number of feature items in , For the document set D i The eigenvalue of the kth eigenvalue in , For the document set D j The eigenvalue of the lth eigenvalue in M kl for and The semantic similarity matrix between i and TF j They are document sets D i and D j TF-IDF value, M kk for In document collection D i The normalization factor in M ll for In document collection D j The normalization factor in ; The semantic similarity is normalized from 0 to 1. If the semantic similarity shows an upward trend during the observation period and the average semantic similarity with historical hot technologies is less than 0.4, the technical life cycle of the keyword is in the growth stage, and the keyword is taken as the target node.

5. The method for identifying priority technologies based on multi-source information according to claim 1, characterized in that: The method of performing association analysis between the keywords and constructing a knowledge graph of technology development includes: Count the co-occurrence frequencies of different keywords in the same sentence, paragraph, and different text types as the correlation strength between keywords; Text mining technology is used to extract the co-occurrence relationship between keywords from multi-source data as the connection relationship between nodes. Keywords are used as nodes, and edges between nodes are constructed based on the strength of association to construct a knowledge graph.

6. The method for identifying priority technologies based on multi-source information according to claim 1, characterized in that: The method for obtaining a set of technology development paths by traversing the target node in the knowledge graph includes: Analyze the correlation strength between nodes in the historical technology development path and calculate the average value as the correlation strength threshold; Starting from the target node, the adjacent nodes are traversed according to the association strength between the nodes. Specifically, if the association strength value is greater than the association strength threshold and the adjacent node is not the target node, the adjacent nodes are continued to be traversed; For the traversal results of each target node, the nodes are sorted from large to small according to their technological development potential as the technological development priority order. The traversed nodes are formed into a technological development path to obtain a set of technological development paths, where the calculation formula for the technological development potential of the node is: Among them, P i is the technological development potential of the ith node, Z is the normalization constant, ψ is the number of adjacent nodes of the ith node, is the semantic association weight coefficient, S ij is the semantic similarity between the i-th node and the j-th node, M i is the maximum semantic similarity between the ith node and its adjacent nodes, C ij is the strength of the association between the i-th node and the j-th node, D i is the minimum correlation between the ith node and its adjacent nodes, F ij is the sum of the frequencies of the i-th node and the j-th node in the multi-source data, F i is the sum of the occurrence frequencies of all node pairs in the knowledge graph, and δ is the frequency influence index.

7. The method for identifying priority technologies based on multi-source information according to claim 1, characterized in that: The method of constructing an objective function according to technology improvement and market benefits, and selecting the optimal technology development path from the technology development path set according to the objective function, comprises: The first objective function is constructed with the maximum technical improvement, and the expression is: Where x is the technical parameter vector, q is the number of performance indicators, ω i is the weight of the i-th performance indicator, P i (x) is the actual performance value of the i-th performance indicator, P i,base is the baseline value of the i-th performance indicator, i.e., the lowest performance level of the historical technology, P i,max (x) is the maximum theoretical value of the i-th performance indicator, a is the nonlinear influence coefficient of adjusting performance improvement, p is the number of resource consumption indicators, and v j is the weight of the jth resource consumption, C j (x) is the actual value of the jth resource consumption, and b is the nonlinear influence coefficient for adjusting resource consumption; The second objective function is constructed with the maximum market return, and the expression is: Where D is the number of different stages in the technology development path, ξ i is the market potential of the i-th technology development stage, λ i is the market share of the i-th technology development stage, t i is the unit profit at the i-th technological development stage, R i is the market risk probability of the i-th technology development stage, F is the number of R&D cost indicators in the technology development path, Q j is the value of the j-th R&D cost indicator, W j is the weight of the j-th R&D cost; According to the strategic focus of technological development, determine the relative weights of technological improvement and market benefits in the overall decision-making. Based on the relative weights, construct a comprehensive evaluation function and select the best technological development path based on comprehensive evaluation.

Citation Information

Patent Citations

  • Social media-oriented topic life cycle trend prediction method and system, and medium

    CN114817761A

  • Method for evaluating new energy bearing capacity of power distribution network based on comprehensive similarity

    CN116247675A

  • Graphene technical route prediction method based on large model and knowledge graph

    CN118840145A

  • Graphene industry application discovery method based on large model and knowledge graph analysis

    CN119168423A

  • Data-driven cross-domain intelligent asset knowledge reasoning and value evaluation method and system

    CN119476499A

Cited By

  • Software development application data processing method based on AI large model

    CN121722360A

  • Cost imposition analysis method based on technology competition deduction

    CN122694481A