Industrial chain innovation chain data processing method and device and electronic equipment
By constructing an integrated graph of the industrial chain and innovation chain and a corpus of nodes, the limitations and low efficiency of traditional methods in analyzing results have been solved. This has enabled intelligent analysis of the industrial chain and innovation chain and differentiated development strategies, thereby improving the efficiency and accuracy of the analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional industrial chain and innovation chain analysis methods involve limited industrial information, have significant limitations in analysis results, and rely heavily on manual methods, which are inefficient and prone to errors.
By acquiring industry research report data, industry data, innovation data, policy data, and demand data, we construct an integrated map of the industrial chain and innovation chain. We use natural language processing technology to extract keywords, identify industrial nodes and innovation nodes, generate a node corpus, and generate differentiated development strategies based on the strength level of the nodes.
It has achieved the convergence and intelligent analysis of industrial and scientific and technological innovation data, identified the main entities in the industrial chain and innovation chain, reduced the workload of integrated analysis, provided targeted development and improvement strategies, and improved the efficiency and accuracy of analysis.
Smart Images

Figure CN117009543B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to an industry chain innovation chain data processing method and device and electronic equipment. BACKGROUND
[0002] Based on the Internet big data and artificial intelligence, the analysis and service of the industry chain innovation chain become more and more important. In the traditional industry chain innovation chain analysis method, the analysis result has great limitation due to less involved industry information, and most of them are performed by artificial means, which is low in efficiency and easy to make mistakes. SUMMARY
[0003] To overcome the problems in the related art, the embodiments of the present application provide an industry chain innovation chain data processing method, device and electronic equipment.
[0004] The present application is realized by the following technical solutions:
[0005] In a first aspect, the embodiments of the present application provide an industry chain innovation chain data processing method, comprising:
[0006] Obtaining target data, the target data including industry research report data, industry data, innovation data, policy data and appeal data, the industry research report data being data related to research, industry standards and planning of a target industry, the industry data being information of each enterprise in the target industry, the innovation data being data related to the innovation capability of the target industry, the policy data being a policy document related to the target industry, and the appeal data being demand information of the target industry;
[0007] Extracting an industry node and an innovation node of the target industry from the industry research report data, and generating an industry chain innovation chain fusion graph according to the industry node and the innovation node, the industry node representing an industrialized technology and product, and the innovation node representing an unindustrialized innovation direction, technology and product;
[0008] Constructing a node corpus of each industry node and each innovation node in the industry data and the innovation data according to the industry chain innovation chain fusion graph;
[0009] Determining the strength level of each node in the node corpus, and generating an industry innovation fusion development strategy of the node according to the strength level of each node and the industry data, the innovation data, the policy data and the appeal data.
[0010] The industrial chain innovation chain data processing method realizes the convergence, management and intelligent analysis of industrial and scientific and technological data through technical means, constructs an industrial chain innovation chain fusion graph, identifies industrial chain innovation chain subjects, analyzes the strengths and weaknesses of the industrial chain innovation chain, and provides differentiated and targeted development promotion strategies for nodes of different strengths. The embodiments of the application can realize the convergence, management and intelligent analysis of industrial and scientific and technological data, reduce the workload of industrial chain innovation chain fusion analysis, and empower the development of industrial chain innovation chain fusion through technical means.
[0011] Based on the first aspect, in some embodiments, the industrial data includes enterprise business data and information data, and the innovation data includes patent data. The industrial chain innovation chain data processing method further includes a step of preprocessing target data.
[0012] The preprocessing of the target data includes:
[0013] A patent keyword is extracted from the patent data by using natural language processing technology to obtain a subject technology feature data set. The patent subject technology feature data set includes a patent subject name, a patent subject type and a patent subject feature. The patent subject name is an inventor and an applicant of a patent. The patent subject type is the type of the inventor and the applicant. The patent subject feature is the extracted patent keyword.
[0014] An information keyword is extracted from the information data by using natural language processing technology to obtain a subject information feature data set. The subject information feature data set includes an information subject name, an information subject type and an information subject feature. The information subject name is an enterprise name, a scientific research institution name, a university name or a person name. The information subject type is an enterprise, a scientific research institution, a university and a talent. The information subject feature is the extracted information keyword.
[0015] An operating range keyword is extracted from the enterprise business data by using natural language processing technology to obtain a subject operating feature data set. The subject operating feature data set includes a business subject name, a business subject type and a business subject feature. The business subject name is an enterprise name. The business subject type is an enterprise. The business subject feature is the extracted operating range keyword.
[0016] Based on the first aspect, in some embodiments, the industrial chain innovation chain data processing method includes:
[0017] A first similarity between a first node feature of a feature type being a technology feature and a subject feature in the subject technology feature data set is calculated. If the first similarity is greater than a first threshold value, the subject feature is marked with an industrial label and a node label. The name of the industrial label and the name of the node label are the industrial name and the node name of the first node feature.
[0018] calculating a second similarity between a second node feature of a feature type of information feature and a subject feature in the subject information feature dataset, if the second similarity is greater than a second threshold, the subject feature corresponds to a subject to be labeled with an industry label and a node label, the name of the industry label and the name of the node label are the industry name and the node name of the second node feature;
[0019] calculating a third similarity between a third node feature of a feature type of operation feature and a subject feature in the subject operation feature dataset, if the third similarity is greater than a third threshold, the subject feature corresponds to a subject to be labeled with an industry label and a node label, the name of the industry label and the name of the node label are the industry name and the node name of the third node feature;
[0020] if the node name is consistent with the node label name of the subject, the subject is identified as a subject belonging to the node; when the industry label of the node is consistent with the industry name, the subject is identified as a subject of the industry, and a subject belongs to one or more nodes and one or more industries.
[0021] Based on the first aspect, in some embodiments, the industry node and the innovation node of the target industry are extracted from the industry research report data, and an industry chain-innovation chain fusion graph is generated according to the industry node and the innovation node, comprising:
[0022] According to the industry research report data, the industry nodes of the target industry are divided, and an industry chain graph is constructed, the industry chain graph comprising node ID, node name, node level, node type and upper node ID;
[0023] According to the industry research report data, the non-industrialized innovation direction, technology and product of the target industry are divided to obtain the innovation node, and the upper node and development stage of the innovation node are marked to form an innovation chain graph, the innovation chain graph comprising node ID, node name, node level, node type, development stage and upper node ID;
[0024] The industry chain graph and the innovation chain graph are merged to form the industry chain-innovation chain fusion graph.
[0025] Based on the first aspect, in some embodiments, the node corpus of each industry node and each innovation node in the industry data and the innovation data is constructed according to the industry chain-innovation chain fusion graph, comprising:
[0026] extracting typical subjects and industry information of the nodes from the industry research report data by using natural language processing technology, the industry information including subject name, subject type, industry name and node name, the subject type being divided into enterprise, scientific research institution, university and individual, the subject name being the name of enterprise, scientific research institution, university and talent, the industry name being the industry to which the subject belongs, and the node name being the node to which the subject belongs;
[0027] For each node of the industry chain innovation connection graph, a node corpus is constructed according to the following steps:
[0028] Step A: querying all subject names of the same node in the typical subjects of the node according to the node name;
[0029] Step B: querying the subject technical feature data set according to the subject name obtained in step A, and importing the feature names of the data items meeting the query condition into the node corpus, to generate a data set including industry name, node name, node feature name and feature type, the industry name being the industry name of the current research, the node name being the node name of the current research, the node feature being the subject feature in the subject technical feature data set, and the feature type being the technical feature;
[0030] Step C: querying the subject information feature data set according to the subject name obtained in step A, and importing the feature names of the data items meeting the query condition into the node corpus, to generate a data set including industry name, node name, node feature and feature type, the industry name being the industry name of the current research, the node name being the node name of the current research, the node feature being the subject feature in the subject information feature data set, and the feature type being the information feature;
[0031] Step D: querying the subject operation feature data set according to the subject name obtained in step A, and importing the feature names of the data items meeting the query condition into the node corpus, to generate a data set including industry name, node name, node feature and feature type, the industry name being the industry name of the current research, the node name being the node name of the current research, the node feature being the subject feature in the subject operation feature data set, and the feature type being the operation feature.
[0032] Based on the first aspect, in some embodiments, the determining of the strength levels of the nodes in the node corpus comprises:
[0033] setting evaluation dimensions of the nodes and weights of each evaluation dimension, the evaluation dimensions including development index, innovation index, scale index, stability index and benefit index;
[0034] According to the level of the target region, sample other regions of the same level, and sort the secondary indicators of the sample regions. According to the ranking, assign scores to the secondary indicators of the regions in equal intervals. The score corresponding to the region with the minimum distance to the secondary indicator value of the target region is the secondary indicator score of the target region. The secondary indicators include at least one of the number of enterprises, the number of listed enterprises, the number of patents, the number of registered funds, the average establishment time of enterprises, and the number of new enterprises;
[0035] By The score of the node A is calculated, i is the i-th evaluation dimension, j is the j-th secondary evaluation indicator under each dimension, s ij is the score of the j-th secondary indicator under the i-th dimension, ω ij is the weight of the j-th secondary indicator under the i-th dimension, ω i is the weight of the i-th dimension, and Score(A) is the score of node A.
[0036] If Score(A) = 0, the strength level of node A is marked as empty, indicating that node A is a missing link.
[0037] When Score(A) is greater than 0 and less than or equal to threshold α, the strength level of node A is marked as weak, indicating that node A is in a disadvantaged state.
[0038] When Score(A) is greater than α and less than or equal to threshold β, the strength level of node A is marked as medium, indicating that node A has neither advantage nor disadvantage.
[0039] When Score(A) is greater than or equal to threshold β, the strength level of node A is marked as strong, indicating that node A is in a dominant state.
[0040] Based on the first aspect, in some embodiments, the innovation data further includes scientific research institution data, instrument equipment data, and public technology service platform data.
[0041] The generation of the industrial innovation integration and development strategy of the node according to the strength level of each node and the industrial data, the innovation data, the policy data, and the appeal data includes:
[0042] Matching the relying units of instrument equipment and public technology platforms with the subject names of node subjects, and establishing a first association relationship between instrument equipment, public technology platforms, and nodes and node subjects.
[0043] Matching the appeal proposers with the subject names of node subjects, and establishing a second association relationship between appeals and nodes and node subjects.
[0044] According to the policy data, the industry and node to which the policy belongs are identified and marked by using natural language processing technology;
[0045] Enterprise research and development foundation analysis: for each enterprise belonging to the node, query all instrument equipment and public technology service platform information associated with the node enterprise according to the first association relationship, and query the patent data of the enterprise according to the subject patent data, and highlight and warn all instrument equipment and public technology service platform information that are not associated and have no patent data for the listed enterprises;
[0046] Investment and financing analysis: identify enterprises and scientific research institutions outside the target analysis area of the node, and for enterprises, select one or more enterprises for recommendation according to the establishment time, registered capital, paid-in capital and patent quantity; for scientific research institutions, select one or more scientific research institutions for recommendation according to the number of patents;
[0047] Talent introduction analysis: identify universities and talents outside the target area of the node, and select one or more universities for recommendation according to the number of patents; for talents, select one or more talents for recommendation according to the number of patents;
[0048] Policy support analysis: according to the marked policy node and the industry to which the policy belongs, query the policies of the node and the industry to which the policy belongs, and push the policy name and policy text to the node subject;
[0049] Equipment sharing analysis: according to the first association relationship and the second association relationship, query the instrument equipment associated with the node and the appeal associated with the node and the appeal type for instrument equipment sharing, and push the instrument equipment name and the associated enterprise to the appeal proposer, and push the appeal proposer and the appeal content to the instrument equipment owning institution;
[0050] Joint research analysis: according to the second association relationship, query the joint research appeal related to the node, and push the appeal proposer and the appeal content to the node subject, and push the node scientific research institution, enterprise name, address and contact information to the appeal proposer;
[0051] Public technology service platform analysis: according to the first association relationship and the second association relationship, query the public service platform associated with the node and the appeal type for public service platform, and push the public service platform name and the relying unit to the appeal subject, and push the appeal proposer and the appeal content to the public technology service platform relying unit;
[0052] For the industry node and innovation node with empty strong and weak level marks, provide investment and financing analysis and talent introduction analysis;
[0053] For the industry nodes and innovation nodes marked as weak in the strength level, enterprise foundation analysis, investment attraction analysis, talent attraction analysis, policy support analysis, equipment sharing analysis, joint research analysis and public service platform analysis are provided.
[0054] For the industry nodes and innovation nodes marked as strong or medium in the strength level, enterprise foundation analysis, policy support analysis, equipment sharing analysis, joint research analysis and public service platform analysis are provided.
[0055] In a second aspect, an embodiment of the present application provides an industrial chain and innovation chain data processing apparatus, comprising:
[0056] An acquisition module is configured to acquire target data, wherein the target data comprises industrial research report data, industrial data, innovation data, policy data and appeal data, the industrial research report data is data related to research, industry standards and planning of a target industry, the industrial data is information of each enterprise in the target industry, the innovation data is data related to innovation capability of the target industry, the policy data is a policy document related to the target industry, and the appeal data is demand information of the target industry;
[0057] A graph generation module is configured to extract industry nodes and innovation nodes of the target industry from the industrial research report data, and generate an industrial chain and innovation chain fusion graph according to the industry nodes and the innovation nodes, wherein the industry nodes represent industrialized technologies and products, and the innovation nodes represent non-industrialized innovation directions, technologies and products;
[0058] A corpus construction module is configured to construct a node corpus of each industry node and each innovation node in the industrial data and the innovation data according to the industrial chain and innovation chain fusion graph;
[0059] A development strategy generation module is configured to determine a strength level of each node in the node corpus, and generate an industrial and innovation fusion development strategy of the node according to the strength level of each node and the industrial data, the innovation data, the policy data and the appeal data.
[0060] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the industrial chain and innovation chain data processing method according to any one of the first aspect when executing the computer program.
[0061] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable by a processor to implement the industrial chain and innovation chain data processing method according to any one of the first aspect.
[0062] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on an electronic device, causes the electronic device to perform the industry chain innovation chain data processing method according to any one of the first aspect.
[0063] It can be understood that beneficial effects of the second aspect to the fifth aspect can be referred to the related description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0065] Figure 1 The flowchart of the industry chain innovation chain data processing method provided by the embodiment of the present application is shown in the figure.
[0066] Figure 2 The structure diagram of the industry chain innovation chain data processing device provided by the embodiment of the present application is shown in the figure.
[0067] Figure 3 The structure diagram of the electronic device provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0068] In the following description, specific details are set forth such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it should be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.
[0069] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0070] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0071] As used in the description of the application and the appended claims, the term “if’ can be interpreted as meaning “when” or “upon” or “in response to determining” or “in response to detecting” depending on the context. Similarly, the phrase “if it is determined” or “if [the described condition or event] is detected” can be interpreted as meaning “upon determining” or “in response to determining” or “upon detecting [the described condition or event]” or “in response to detecting [the described condition or event]”, depending on the context.
[0072] In addition, in the description of the application and the appended claims, the terms “first”, “second”, “third”, etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0073] In the description of the application, the reference “one embodiment” or “some embodiments” and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the application. Therefore, the statements “in one embodiment”, “in some embodiments”, “in other some embodiments”, “in further some embodiments” and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean “one or more but not all embodiments”, unless otherwise specifically emphasized. The terms “include”, “contain”, “have” and their variants mean “include but not limited to”, unless otherwise specifically emphasized.
[0074] The embodiments of the present application mainly use big data and artificial intelligence technology to analyze the fusion of industry chain and innovation chain, realize the convergence, governance and intelligent analysis of industry and scientific and technological data, build an industry chain and innovation chain fusion map, identify industry chain and innovation chain subjects, analyze the strengths and weaknesses of industry chain and innovation chain, and provide differentiated and targeted development promotion strategies for nodes with different strengths. On the basis of realizing the convergence, governance and intelligent analysis of industry and scientific and technological data, the workload of industry chain and innovation chain fusion analysis is reduced, and technical means are used to empower the development of industry chain and innovation chain fusion.
[0075] Figure 1 The schematic flowchart of the industry chain and innovation chain data processing method provided by an embodiment of the present application is shown in FIG. 1. Referring to FIG. 1, the method comprises the following steps. Figure 1 The detailed description of the industry chain and innovation chain data processing method is as follows:
[0076] In step 101, target data is acquired, which includes industry research report data, industry data, innovation data, policy data and appeal data.
[0077] The industry research report data is related to research, industry standards and planning of the target industry, the industry data is information of each enterprise in the target industry, the innovation data is related to the innovation capability of the target industry, the policy data is a policy document related to the target industry, and the appeal data is demand information of the target industry.
[0078] For example, the industry research report data can include industry research report, industry standard report and industry planning document, the industry data can include enterprise business data, top enterprise data and enterprise information data, the innovation data can include scientific research institution data, patent data, equipment and instrument data and public technology service platform data.
[0079] Specifically, the top enterprise data can include top list name, enterprise name and evaluation time, the scientific research institution data can include scientific research institution name, scientific research institution type, establishment time and relying unit, the instrument and equipment data can include instrument and equipment name and owning institution, the public technology service platform data can include public service platform name and relying unit, the policy data can include policy name, issuing unit, issuing time and policy text, the appeal data can include appeal type, appeal person and appeal content, and the appeal type can include instrument and equipment sharing appeal, joint public relations appeal and public service platform appeal. The patent data can include patent title, applicant, inventor, patent abstract and right claim, the industry information data can include information title and information text, the industry research report data can include research report title and research report text, and the enterprise business data can include enterprise name, unified social credit code and business scope.
[0080] In some scenarios, the industry research report data can be collected through industry research websites by using a crawler or manually, the enterprise business data can be collected through enterprise websites or third-party enterprise websites, the top enterprise data, scientific research institution data, equipment and instrument data, public technology service platform data, policy data and appeal data can be collected through government official websites, the patent data can be collected through patent websites, and the industry information data can be collected through information websites. After obtaining the data, the data can be imported into a corresponding database for storage.
[0081] In some embodiments, after step 101, the above-mentioned industry chain and innovation chain data processing method can further include a step of preprocessing the target data.
[0082] Specifically, the preprocessing of the target data can include:
[0083] extracting patent keywords from the patent data using natural language processing technology to obtain a subject technology feature dataset, the patent subject technology feature dataset containing patent subject names, patent subject types, and patent subject features, the patent subject names being inventors and applicants of the patent, the patent subject types being types of the inventors and the applicants (for example, enterprises, scientific research institutions, colleges and universities, and individuals), and the patent subject features being the extracted patent keywords;
[0084] extracting information keywords from the information data using natural language processing technology to obtain a subject information feature dataset, the subject information feature dataset containing information subject names, information subject types, and information subject features, the information subject names being enterprise names, scientific research institution names, college and university names, or personal names, the information subject types being enterprises, scientific research institutions, colleges and universities, and talents, and the information subject features being the extracted information keywords;
[0085] extracting operating range keywords from the enterprise business data using natural language processing technology to obtain a subject operation feature dataset, the subject operation feature dataset containing business subject names, business subject types, and business subject features, the business subject names being enterprise names, the business subject types being enterprises, and the business subject features being the extracted operating range keywords.
[0086] Step 102, extracting industry nodes and innovation nodes of the target industry from the industry research report data, and generating an industry chain-innovation chain fusion graph according to the industry nodes and the innovation nodes.
[0087] The industry nodes represent technologies and products that have been industrialized, and the innovation nodes represent innovation directions, technologies, and products that have not been industrialized. For example, the industry nodes and the innovation nodes are divided and defined by manual analysis according to the industry research report data. The innovation nodes can be cutting-edge innovation directions, technologies, and products, and the industry nodes can be mature technologies and products that have been industrialized
[0088] For example, step 102 can include:
[0089] According to the industry research report data, the industry nodes of the target industry are divided, an industry chain graph is constructed, and the industry chain graph contains node ID, node name, node level, node type, and superior node ID;
[0090] According to the industry research report data, the non-industrialized innovation directions, technologies, and products of the target industry are divided to obtain innovation nodes, the superior nodes and development stages of the innovation nodes are marked, and an innovation chain graph is formed, the innovation chain graph containing node ID, node name, node level, node type, development stage, and superior node ID;
[0091] The industry chain graph and the innovation chain graph are merged to form the industry chain innovation chain fusion graph.
[0092] The superior node of the innovation node can be an industry node or an innovation node, and the development stage of the innovation node can include experimental research, demonstration application, commercialization promotion, and industrialization diffusion. An example of the innovation chain graph is as follows (041, wafer-level integrated AI chip, 4 levels, innovation node, demonstration application, 031), and an example of the industry chain graph is as follows (031, AI chip, 3 levels, industry node, 022).
[0093] In step 103, a node corpus of each industry node and each innovation node in the industry data and the innovation data is constructed according to the industry chain innovation chain fusion graph.
[0094] For example, step 103 can include:
[0095] Typical subjects and industry information of the nodes are extracted from the industry research report data by using a natural language processing technology, including subject names, subject types, industry names, and node names. The subject types are divided into enterprises, scientific research institutions, colleges and universities, and individuals. The subject names are the names of enterprises, scientific research institutions, colleges and universities, and talents. The industry names are the industries to which the subjects belong. The node names are the nodes to which the subjects belong. For example, (Tsinghua University, college, chip, AI chip), and the processed data is imported into a node typical subject data set.
[0096] For each node of the industry chain innovation chain graph, a node corpus is constructed according to the following steps:
[0097] In step A, all subject names of the same node are queried in the node typical subject according to the node name.
[0098] In step B, the subject technical feature data set is queried according to the subject names obtained in step A, and the feature names of the data items that meet the query condition are imported into the node corpus. The generated data set includes the industry name, the node name, the node feature name, and the feature type. The industry name is the industry name of the current research. The node name is the node name of the current research. The node feature is the subject feature in the subject technical feature data set. The feature type is the technical feature.
[0099] In step C, the subject information feature data set is queried according to the subject names obtained in step A, and the feature names of the data items that meet the query condition are imported into the node corpus. The generated data set includes the industry name, the node name, the node feature, and the feature type. The industry name is the industry name of the current research. The node name is the node name of the current research. The node feature is the subject feature in the subject information feature data set. The feature type is the information feature.
[0100] Step D, querying the subject operating feature dataset according to the subject name obtained in step A, and importing the feature name of the data item meeting the query condition into the node corpus to generate a dataset as follows: industry name, node name, node feature, feature type, the industry name is the industry name of the current research, the node name is the node name of the current research, the node feature is the subject feature in the subject operating feature dataset, and the feature type is the operating feature.
[0101] Step E, de-duplicating the obtained node corpus training set, and screening the data in the node corpus through expert analysis.
[0102] For example, after step 103, the above-mentioned industry chain innovation chain data processing method can further include:
[0103] calculating a first similarity between the first node feature with the feature type of technical feature and the subject feature in the subject technical feature dataset, if the first similarity is greater than a first threshold, the subject feature corresponds to the subject labeled industry label and node label, and the name of the industry label and the name of the node label are the industry name and the node name of the first node feature;
[0104] calculating a second similarity between the second node feature with the feature type of information feature and the subject feature in the subject information feature dataset, if the second similarity is greater than a second threshold, the subject feature corresponds to the subject labeled industry label and node label, and the name of the industry label and the name of the node label are the industry name and the node name of the second node feature;
[0105] calculating a third similarity between the third node feature with the feature type of operating feature and the subject feature in the subject operating feature dataset, if the third similarity is greater than a third threshold, the subject feature corresponds to the subject labeled industry label and node label, and the name of the industry label and the name of the node label are the industry name and the node name of the third node feature;
[0106] If the node name is consistent with the node label name of the subject, the subject is identified as belonging to the node; when the industry label of the node is consistent with the industry name, the subject is identified as the subject of the industry, and one subject belongs to one or more nodes and one or more industries.
[0107] Then, according to the newly identified node subject, the node corpus is continuously optimized according to the foregoing content until no new node subject can be identified.
[0108] Step 104, determining the strength level of each node in the node corpus, and generating the industry innovation integration development strategy of the node according to the strength level of each node and the industry data, the innovation data, the policy data and the appeal data.
[0109] For example, the "determining the strength level of each node in the node corpus" in step 104 can include:
[0110] Setting the evaluation dimensions of the nodes and the weight of each evaluation dimension, the evaluation dimensions include development index, innovation index, scale index, stability index and benefit index; the average weighted operator can be used to aggregate the scores of each dimension, and the preferences of industrial researchers can be flexibly configured;
[0111] According to the level of the target region, sampling other regions of the same level, and sorting the secondary indicators of the sample regions, assigning scores according to the ranking, and the score corresponding to the region with the smallest distance to the secondary indicator value of the target region is the secondary indicator score of the target region, the secondary indicators include at least one of the number of enterprises, the number of listed enterprises, the number of patents, the number of registered funds, the average establishment time of enterprises and the number of new enterprises, and the secondary indicators of each dimension can be flexibly selected according to the preferences of industrial researchers;
[0112] Setting the weight of each dimension and the secondary indicators and weights, selecting secondary evaluation indicators for the set evaluation dimensions, and setting the weight of each evaluation indicator, the average weighted operator can be used to aggregate the scores of the secondary indicators of each dimension;
[0113] By Calculating the score of the node, A is the evaluated node, i is the ith evaluation dimension, j is the jth secondary evaluation indicator under each dimension, s ij is the score of the jth secondary indicator under the ith dimension, ω ij is the weight of the jth secondary indicator under the ith dimension, ω i is the weight of the ith dimension, and Score(A) is the score of node A;
[0114] If Score(A) = 0, the strength level of node A is marked as empty, indicating that node A is a missing link;
[0115] When Score(A) is greater than 0 and less than or equal to the threshold value a, the strength level of node A is marked as weak, indicating that node A is in a disadvantaged state;
[0116] When Score(A) is greater than a and less than or equal to the threshold value b, the strength level of node A is marked as medium, indicating that node A has neither advantage nor disadvantage;
[0117] When Score(A) is greater than or equal to the threshold value b, the strength level of node A is marked as strong, indicating that node A is in a dominant state.
[0118] Exemplarily, the "generating an industrial innovation integration development strategy of the node according to the strength level of each node and the industrial data, the innovation data, the policy data and the appeal data" in step 104 can include:
[0119] Matching the name of the relying unit of the instrument equipment and the public technology platform with the name of the node subject, establishing a first association relationship between the instrument equipment, the public technology platform and the node and the node subject;
[0120] Matching the name of the appeal presenter with the name of the node subject, establishing a second association relationship between the appeal and the node and the node subject;
[0121] According to the policy data, using natural language processing technology to identify and mark the industry and node to which the policy belongs;
[0122] Enterprise research foundation analysis: for each enterprise belonging to the node, querying all instrument equipment and public technology service platform information associated with the node enterprise according to the first association relationship, and querying the patent data of the enterprise according to the subject patent data, highlighting and warning all instrument equipment and public technology service platform information that are not associated and all enterprises without patent data;
[0123] Investment attraction analysis: identifying enterprises and scientific research institutions outside the target analysis area of the node, selecting one or more enterprises for recommendation according to the enterprise establishment time, registered capital, paid-in capital and patent quantity for enterprises; for scientific research institutions, selecting one or more scientific research institutions for recommendation according to the number of patents;
[0124] Talent introduction analysis: identifying colleges and universities and talents outside the target area of the node, selecting one or more colleges and universities for recommendation according to the number of patents; for talents, selecting one or more talents for recommendation according to the number of patents;
[0125] Policy support analysis: querying the policies of the node and the industry to which the node belongs according to the marked node and industry to which the policy belongs, and pushing the policy name and policy text to the node subject;
[0126] Equipment sharing analysis: querying the instrument equipment associated with the node and the appeal associated with the node and the appeal type of instrument equipment sharing according to the first association relationship and the second association relationship, and pushing the instrument equipment name and the enterprise to which the instrument equipment belongs to the appeal presenter, and pushing the appeal presenter and the appeal content to the instrument equipment owning institution;
[0127] Joint research analysis: querying the joint research appeal associated with the node according to the second association relationship, and pushing the appeal presenter and the appeal content to the node subject, and pushing the names, addresses and contact information of the scientific research institutions and enterprises to the appeal presenter;
[0128] Public technical service platform analysis: according to the first association relationship and the second association relationship, the public service platform associated with the node and the appeal type for the public service platform are queried, and the public service platform name and the relying unit are pushed to the appeal subject, and the appeal proposer and the appeal content are pushed to the public technical service platform relying unit;
[0129] For the industry nodes and innovation nodes with empty strong and weak level marks, investment attraction analysis and talent attraction analysis are provided;
[0130] For the industry nodes and innovation nodes with weak strong and weak level marks, enterprise foundation analysis, investment attraction analysis, talent attraction analysis, policy support analysis, equipment sharing analysis, joint research analysis and public service platform analysis are provided;
[0131] For the industry nodes and innovation nodes with strong or medium strong and weak level marks, enterprise foundation analysis, policy support analysis, equipment sharing analysis, joint research analysis and public service platform analysis are provided.
[0132] The above-mentioned industry chain and innovation chain data processing method realizes the convergence, governance and intelligent analysis of industry and scientific and technological innovation data through technical means, constructs an industry chain and innovation chain fusion graph, identifies industry chain and innovation chain subjects, analyzes the strengths and weaknesses of the industry chain and innovation chain, and provides differentiated and targeted development promotion strategies for nodes of different strong and weak levels. The embodiment of the present application can realize the convergence, governance and intelligent analysis of industry and scientific and technological innovation data, reduce the workload of industry chain and innovation chain fusion analysis, and empower the development of industry chain and innovation chain fusion through technical means.
[0133] It should be understood that the size of the serial number of each step in the above-mentioned embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0134] Corresponding to the industry chain and innovation chain data processing method described in the above embodiment, Figure 2 The structural block diagram of the industry chain and innovation chain data processing device provided by the embodiment of the present application is shown, and only the part related to the embodiment of the present application is shown for easy explanation.
[0135] Referring to Figure 2 The industry chain and innovation chain data processing device in the embodiment of the present application can include
[0136] The acquisition module 201 is configured to acquire target data, the target data including industry research report data, industry data, innovation data, policy data and appeal data, the industry research report data being data related to research, industry standards and planning of a target industry, the industry data being information of each enterprise in the target industry, the innovation data being data related to innovation capability of the target industry, the policy data being a policy document related to the target industry, and the appeal data being demand information of the target industry.
[0137] The graph generation module 202 is configured to extract an industry node and an innovation node of the target industry from the industry research report data, and generate an industry chain-innovation chain fusion graph according to the industry node and the innovation node, the industry node representing an industrialized technology and product, and the innovation node representing an unindustrialized innovation direction, technology and product.
[0138] The corpus construction module 203 is configured to construct a node corpus of each industry node and each innovation node in the industry data and the innovation data according to the industry chain-innovation chain fusion graph.
[0139] The development strategy generation module 204 is configured to determine a strength level of each node in the node corpus, and generate an industry-innovation fusion development strategy of the node according to the strength level of each node and the industry data, the innovation data, the policy data and the appeal data.
[0140] It should be noted that the information interaction and execution process between the above apparatus / units are based on the same concept as the method embodiments, and the specific functions and technical effects can be referred to the method embodiments, which will not be described here.
[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the above described functions. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or software. In addition, the specific names of the functional units and modules are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can be referred to the corresponding process in the method embodiments, which will not be described here.
[0142] The embodiment of the present application also provides an electronic device, which is shown inFigure 3 The electronic device 300 can include at least one processor 310, a memory 320, and a computer program stored in the memory 320 and executable on the at least one processor 310, the processor 310 implementing the steps in any of the method embodiments described above when executing the computer program, for example Figure 1 the steps 101 to 104 in the illustrated embodiments. Alternatively, the processor 310 implements the functions of the modules / units in the apparatus embodiments described above when executing the computer program, for example Figure 2 the functions of the modules 201 to 204 illustrated.
[0143] By way of example, the computer program can be segmented into one or more modules / units, one or more of which are stored in the memory 320 and executed by the processor 310 to accomplish the present application. The one or more modules / units can be a series of computer program segments that can accomplish a specific function, which are used to describe the execution process of the computer program in the electronic device 300.
[0144] Those skilled in the art can understand that Figure 3 is merely an example of an electronic device and does not constitute a limitation on the electronic device, which can include more or fewer components than those shown, or combine some components, or different components, for example, input / output devices, network access devices, buses, etc.
[0145] The processor 310 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0146] The memory 320 can be an internal storage unit of the electronic device, or an external storage device of the electronic device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 320 is used to store the computer program and other programs and data required by the electronic device. The memory 320 can also be used to temporarily store data that has been output or will be output.
[0147] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.
[0148] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps in each embodiment of the industrial chain innovation chain data processing method.
[0149] The embodiment of the present application provides a computer program product, when the computer program product is run on a mobile terminal, so that the mobile terminal executes the steps in each embodiment of the industrial chain innovation chain data processing method.
[0150] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application realizes all or part of the processes in the above-mentioned embodiments, which can be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium, and the computer program can realize the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer readable medium at least includes any entity or device capable of carrying the computer program code to an electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0151] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0152] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A data processing method for the industrial chain and innovation chain, characterized in that, include: Acquire target data, which includes industry research report data, industry data, innovation data, policy data, and demand data. The industry research report data is data related to research, industry standards, and planning of the target industry. The industry data is information on various enterprises in the target industry. The innovation data is data related to the innovation capabilities of the target industry. The policy data is policy documents related to the target industry. The demand data is demand information of the target industry. The industry nodes and innovation nodes of the target industry are extracted from the industry research report data. An industrial chain and innovation chain integration map is generated based on the industry nodes and innovation nodes. The industry nodes represent industrialized technologies and products, and the innovation nodes represent unindustrialized innovation directions, technologies and products. Based on the industrial chain and innovation chain integration map, a node corpus is constructed for each industrial node and each innovation node in the industrial data and the innovation data. The node corpus includes multiple nodes composed of each industrial node and each innovation node. Determine the strength level of each node in the node corpus, and generate a node-based industrial innovation integration and development strategy based on the strength level of each node, the industry data, the innovation data, the policy data, and the demand data. The construction of a node corpus for each industry node and each innovation node in the industry data and innovation data based on the industry chain and innovation chain fusion map includes: Natural language processing technology is used to extract typical subject and industry information of nodes from the industry research report data, including subject name, subject type, industry name and node name. The subject type is divided into enterprises, research institutions, universities and individuals. The subject name is the name of the enterprise, research institution, university and talent. The industry name is the industry to which the subject belongs. The node name is the node to which the subject belongs. For each node in the integration map of the industrial chain and innovation chain, construct a node corpus according to the following steps: Step A: Search for all subject names of the same node in the typical subjects of the node by node name; Step B: Based on the subject name obtained in Step A, query the subject technical feature dataset, import the feature names of data items that meet the query conditions into the node corpus, and generate the dataset as: industry name, node name, node feature name and feature type. The industry name is the name of the industry currently being studied, the node name is the name of the node currently being studied, the node feature is the subject feature in the subject technical feature dataset, and the feature type is the technical feature. Step C: Based on the subject name obtained in Step A, query the topic information feature dataset, import the feature names of data items that meet the query conditions into the node corpus, and generate the following dataset: industry name, node name, node feature and feature type. The industry name is the industry name currently being studied, the node name is the node name currently being studied, the node feature is the subject feature in the subject information feature dataset, and the feature type is information feature. Step D: Based on the subject name obtained in Step A, query the subject business feature dataset, import the feature names of data items that meet the query conditions into the node corpus, and generate the following dataset: industry name, node name, node feature, feature type. The industry name is the name of the industry currently being studied, the node name is the name of the node currently being studied, the node feature is the subject feature in the subject business feature dataset, and the feature type is the business feature. The innovation data also includes data from research institutions, instrument and equipment, and public technology service platforms. The process of generating industrial innovation and integration development strategies for nodes based on their strength levels, industry data, innovation data, policy data, and demand data includes: Match the supporting units of instruments and equipment and public technology platforms with the main entity names of node entities to establish the primary association between instruments and equipment, public technology platforms and nodes and node entities; Match the name of the person making the request with the name of the node entity to establish a second association between the request and the node, and the node entity. Based on policy data, natural language processing technology is used to identify and label the industries and nodes to which the policies belong; Enterprise R&D Foundation Analysis: For each enterprise belonging to a node, query all instruments, equipment and public technology service platform information related to the node enterprise based on the first association relationship, and query the enterprise's patent data based on the main patent data. Highlight and issue warnings for all instruments, equipment and public technology service platform information that are not related and for the list of enterprises without patent data. Investment attraction analysis: Identify enterprises and research institutions located outside the target analysis area. For enterprises, select one or more for recommendation based on their establishment time, registered capital, paid-in capital, and number of patents; for research institutions, select one or more for recommendation based on the number of patents. Talent recruitment and attraction analysis: Identify universities and talents whose nodes are outside the target region, and recommend one or more universities based on the number of patents held by the universities; for talents, recommend one or more talents based on the number of patents held. Policy support analysis: Based on the marked policy's node and industry, query the policies of that node and industry, and push the policy name and policy text to the node entity; Equipment sharing analysis: Based on the first and second association relationships, query the instruments and equipment associated with the node and the requests associated with the node with the request type of instrument and equipment sharing, and push the instrument and equipment name and the company to which it belongs to to the requester, and push the requester and the request content to the institution that owns the instrument and equipment. Joint research analysis: Based on the second relationship, query the joint research requests related to the node, and push the requester and the request content to the node entity, and push the name, address and contact information of the research institution and enterprise to the requester; Public technology service platform analysis: Based on the first and second association relationships, query the public service platforms associated with the node and the requests with the request type of public service platform, and push the name of the public service platform and the supporting unit to the request subject, and push the requester and the request content to the supporting unit of the public technology service platform; Evaluation scores for industry nodes and innovation nodes are calculated based on evaluation dimensions and secondary indicators. The evaluation dimensions include development index, innovation index, scale index, stability index, and efficiency index. The secondary indicators include at least one of the following: number of enterprises, number of listed enterprises, number of patents, amount of registered capital, average establishment time of enterprises, and number of newly added enterprises. The strength level of a node is determined based on the evaluation scores. A score of 0 indicates an empty strength level; a score greater than 0 and less than or equal to a first threshold indicates a weak strength level; a score greater than the first threshold and less than a second threshold indicates a medium strength level; and a score greater than or equal to the second threshold indicates a strong strength level. For industry nodes and innovation nodes with empty strength / weakness ratings, we provide analysis on investment attraction and talent attraction. For industry nodes and innovation nodes marked as weak in the strength and weakness level, we provide enterprise production and research foundation analysis, investment attraction analysis, talent attraction analysis, policy support analysis, equipment sharing analysis, joint research analysis, and public technology service platform analysis. For industry nodes and innovation nodes marked as strong or medium in strength level, we provide enterprise production and research foundation analysis, policy support analysis, equipment sharing analysis, joint research analysis, and public technology service platform analysis.
2. The data processing method for the industrial chain and innovation chain as described in claim 1, characterized in that, The industry data includes enterprise business registration data and information data, the innovation data includes patent data, and the industrial chain and innovation chain data processing method also includes a step of preprocessing the target data; The preprocessing of the target data includes: Natural language processing technology is used to extract patent keywords from the patent data to obtain a patent subject technical feature dataset. The patent subject technical feature dataset includes the patent subject name, patent subject type and patent subject features. The patent subject name is the inventor and applicant of the patent, the patent subject type is the type of inventor and applicant, and the patent subject features are the extracted patent keywords. Natural language processing technology is used to extract information keywords from the information data to obtain a main information feature dataset. The main information feature dataset includes the information subject name, information subject type, and information subject features. The information subject name is the name of an enterprise, research institution, university, or person. The information subject type is an enterprise, research institution, university, or talent. The information subject features are the extracted information keywords. Natural language processing technology is used to extract business scope keywords from the enterprise's business registration data to obtain a main business feature dataset. The main business feature dataset includes the business entity name, business entity type, and business entity features. The business entity name is the enterprise name, the business entity type is an enterprise, and the business entity features are the extracted business scope keywords.
3. The data processing method for the industrial chain and innovation chain as described in claim 2, characterized in that, The data processing method for the industrial chain and innovation chain also includes: Calculate the first similarity between the first node feature (characteristic type: technical feature) and the main feature in the main technical feature dataset. If the first similarity is greater than the first threshold, then label the subject corresponding to the main feature with an industry label and a node label. The name of the industry label and the name of the node label are the industry name and node name of the first node feature. Calculate the second similarity between the second node feature (feature type: information feature) and the main feature in the main information feature dataset. If the second similarity is greater than the second threshold, then label the subject corresponding to the main feature with an industry label and a node label. The name of the industry label and the name of the node label are the industry name and node name of the second node feature. Calculate the third similarity between the third node feature of the feature type of business feature and the main feature in the main business feature dataset. If the third similarity is greater than the third threshold, then label the main body corresponding to the main feature with industry label and node label. The name of the industry label and the name of the node label are the industry name and node name of the third node feature. If the node name matches the node label name of the subject, the subject is identified as belonging to that node; when the industry label of the node matches the industry name, the subject is identified as the subject of that industry. A subject belongs to one or more nodes or one or more industries.
4. The data processing method for the industrial chain and innovation chain as described in claim 1, characterized in that, The process of extracting industry nodes and innovation nodes of the target industry from the industry research report data, and generating an integrated industrial chain and innovation chain map based on the industry nodes and innovation nodes, includes: Based on the industry research report data, the industry nodes of the target industry are divided to construct an industry chain map. The industry chain map includes node ID, node name, node level, node type and parent node ID. Based on the industry research report data, the un-industrialized innovation directions, technologies and products of the target industry are divided to obtain innovation nodes, and the superior nodes and development stages of the innovation nodes are marked to form an innovation chain map. The innovation chain map includes node ID, node name, node level, node type, development stage and superior node ID. The industrial chain map and the innovation chain map are merged to form the integrated industrial chain and innovation chain map.
5. The data processing method for the industrial chain and innovation chain as described in claim 1, characterized in that, Determining the strength level of each node in the node corpus includes: Set the evaluation dimensions for each node and the weight of each evaluation dimension; Based on the target region's level, other regions at the same level are sampled, and the secondary indicators of the sample regions are ranked. The scores of the secondary indicators of the regions are assigned equally spaced according to the ranking. The score of the region with the smallest distance from the secondary indicator value of the target region is the secondary indicator score of the target region. pass Calculate the score for each node, where A is the node being evaluated. i For the first i One evaluation dimension, j The first in each dimension j Two secondary evaluation indicators, For the first i The first dimension j The scores of each secondary indicator, For the first i The first dimension j The weights of each secondary indicator, For the first i Weights of each dimension The score for node A; like =0, the strength level of node A is empty, indicating that node A is a missing link; when When the value is greater than 0 and less than or equal to the threshold α, the strength level of node A is marked as weak, indicating that node A is in a disadvantageous state. when When the strength level of node A is greater than α and less than the threshold β, the strength level of node A is marked as medium, indicating that node A has neither advantage nor disadvantage. when When the strength level of node A is greater than or equal to the threshold β, it is marked as strong, indicating that node A is in a dominant state.
6. A data processing device for industrial chain and innovation chain, used to implement the method as described in any one of claims 1-5, characterized in that, include: The acquisition module is used to acquire target data, which includes industry research report data, industry data, innovation data, policy data, and demand data. The industry research report data is data related to research, industry standards, and planning of the target industry. The industry data is information on various enterprises in the target industry. The innovation data is data related to the innovation capabilities of the target industry. The policy data is policy documents related to the target industry. The demand data is demand information of the target industry. The graph generation module is used to extract industry nodes and innovation nodes of the target industry from the industry research report data, and generate an industrial chain and innovation chain integration graph based on the industry nodes and innovation nodes. The industry nodes represent industrialized technologies and products, and the innovation nodes represent unindustrialized innovation directions, technologies and products. The corpus construction module is used to construct a node corpus for each industry node and each innovation node in the industry data and the innovation data based on the industry chain and innovation chain fusion graph. The development strategy generation module is used to determine the strength level of each node in the node corpus, and generate the industrial innovation integration and development strategy of the node based on the strength level of each node, as well as the industry data, innovation data, policy data and demand data.
7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Industrial chain analysis system and method based on big data
CN111080132A
Industrial chain strength evaluation method and device, equipment and readable medium
CN113869768A