Future industry innovation system international comparison method based on hypergraph motif
Through the method based on supermap model, the construction and analysis of institutions cooperated with supermap has solved the problems of insufficient adaptability and lack of interactive modeling in the research on innovation systems in existing technologies, and achieved more accurate portrayal of future industrial innovation systems and promotion of technological breakthroughs.
Patent Information
- Application Number
- CN202510573415.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When studying the national innovation system, the existing technology has problems such as insufficient context adaptability, lack of advanced interactive modeling and single model analysis dimensions, which are difficult to adapt to the high integration and nonlinear characteristics of future industries.
Using a method based on the supergraph model, the cooperative relationship data between different institutions is collected and preprocessed, and the institutional cooperative hypergraph is constructed. The enumeration method is used to determine the total number of k-order non-isomorphic connected sub-supermaps, and a random zero model is generated to calculate its average value to calculate the abundance of the k-order non-isomorphic connected sub-supermaps in the cooperative supermaps of each institution and generate a significant profile.
It has achieved a more accurate representation of the future industrial innovation system, reduced estimation errors, better understood the interactive modes between multiple innovation entities, and promoted technological breakthroughs.
Smart Images

Figure CN120086697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network data processing, and particularly to an international comparison method for the future industrial innovation system based on hypergraph motifs. Background Art
[0002] Currently, the research methods for national innovation systems have the following significant limitations:
[0003] I. Insufficient context adaptability: Existing technologies mostly focus on traditional industries such as automobile manufacturing or strategic emerging industries such as nanotechnology. Their analysis frameworks are based on linear innovation assumptions and are difficult to adapt to the high integration and non-linear characteristics of future industries. For example, although technologies such as Chinese patents CN116842182A and CN118656713A attempt to predict the future industrial technology directions, they only focus on the role of single enterprise entities and ignore the synergy effects of multiple entities such as national laboratories and universities. Such methods cannot reveal the impact of multi-entity interactions on the evolution path of the national innovation system, resulting in biases in technology route assessments.
[0004] II. Lack of high-order interaction modeling: Current research mostly uses classical network models, assuming only pairwise connection relationships between innovation entities. However, major technological breakthroughs in future industries often occur in the form of "large projects" and require multiple institutions to form a collaborative giant system to participate together. Although the network motif analysis method proposed in Chinese patent CN115766476A focuses on local structures, it is limited by the binary connection assumption and cannot represent the high-order interaction characteristics of multi-node collaboration, resulting in estimation errors in the collaborative efficiency of innovation networks.
[0005] III. Single motif analysis dimension: Although motifs, as topological primitives of complex networks, have been proven to be decisive for network functions (Science, 2002, 298(5594): 824-827), existing technologies such as CN113486217A only use motifs to identify node importance and do not extend it to network structure comparison analysis. While traditional global metrics (such as node centrality) can characterize the macroscopic features of networks, they cannot analyze the micro-driving mechanisms of specific collaborative patterns (such as the government-industry-university triangular closed loop) on the evolution of the national innovation system, limiting the ability to compare the differences between cross-national and cross-industry innovation systems. Summary of the Invention
[0006] The purpose of the present invention is to provide an international comparison method for the future industrial innovation system based on hypergraph motifs to solve the problems of insufficient context adaptability, lack of high-order interaction modeling, and single motif analysis dimension existing in the research methods of national innovation systems in the prior art.
[0007] To achieve the above object, the present application adopts the following technical solutions: An international comparison method for the future industrial innovation system based on hypergraph motifs of the present application includes the following steps: Collect the original dataset representing the cooperation relationships between different institutions within the target industry, and preprocess the original dataset to obtain several institution cooperation datasets divided by country and the attributes of each institution; Based on the institution cooperation datasets, construct an institution cooperation hypergraph for each country. The nodes in the institution cooperation hypergraph represent institutions and are attached with institution attributes, and the hyperedges represent cross-institution cooperation events; For each institution cooperation hypergraph, use the enumeration method to determine the total number of k-order non-isomorphic connected sub-hypergraphs, where k is an integer greater than 1, and generate multiple random null models for it, and calculate the average value of the k-order non-isomorphic connected sub-hypergraphs in the multiple random null models; Calculate the abundance of the k-order non-isomorphic connected sub-hypergraphs in each institution cooperation hypergraph according to each total number and its corresponding average value, and generate a significance profile to compare the target industrial innovation systems of different countries.
[0008] Preferably, the preprocessing of the original dataset to obtain several institution cooperation datasets divided by country and the attributes of each institution includes: Extract all institution names included in the original dataset, and classify all institutions by country; Disambiguate all institution names, and label the attributes of each institution according to the research institution registry; Divide the original dataset into several institution cooperation datasets distinguished by country according to the classified institutions and the disambiguated institution names. The institution cooperation datasets include the attributes of each institution.
[0009] Preferably, the use of the enumeration method to determine the total number of k-order non-isomorphic connected sub-hypergraphs includes: Enumerate all hyperedges containing k nodes in the institution cooperation hypergraph to obtain the number of basic non-isomorphic connected sub-hypergraphs of the k-order motif; For hyperedges with a node number less than k, expand its node number to k, and generate extended k-order non-isomorphic connected sub-hypergraphs according to the expanded node set; Add the number of basic induced sub-hypergraphs of the k-order motif to the number of the extended k-order non-isomorphic connected sub-hypergraphs to obtain the total number of k-order non-isomorphic connected sub-hypergraphs.
[0010] Preferably, the expansion of its node number to k includes: If adding a node to the node set included in a hyperedge with a node number less than k expands its node number to k, then the added node exists in the union of the exclusive neighborhood of the nodes included in the hyperedge and its adjacent hyperedges; If at least two nodes are added to the node set included in a hyperedge with a node count less than k, and its node count is only then expanded to k, the added nodes belong to its adjacent hyperedges, and the later-added nodes are the exclusive neighborhoods of the earlier-added nodes.
[0011] Preferably, generating the expanded k-order non-isomorphic connected sub-hypergraph according to the expanded node set includes: Generating the power set of the k nodes included in the expanded node set, determining the combinations existing in the corresponding institutional cooperation hypergraph in the power set, and generating the expanded k-order non-isomorphic connected sub-hypergraph accordingly.
[0012] Preferably, before calculating the average value of the k-order non-isomorphic connected sub-hypergraphs in the multiple random null models, it further includes: Determining the quantity of order k in each random null model according to the above-mentioned enumeration method.
[0013] Preferably, calculating the abundance of the k-order non-isomorphic connected sub-hypergraphs in each institutional cooperation hypergraph according to each total number and its corresponding average value and generating a significance profile includes: Calculating the difference and sum between each total number and its corresponding average value respectively; Using the sum obtained plus a preset adjustment factor as the denominator and the difference as the numerator to calculate the abundance of the k-order non-isomorphic connected sub-hypergraph in the corresponding institutional cooperation hypergraph; Normalizing the abundance to obtain its significance profile.
[0014] An electronic device includes a memory and a processor, the memory is used to store one or more computer instructions, wherein, the one or more computer instructions are executed by the processor to implement a method for international comparison of future industrial innovation systems based on hypergraph motifs as described in any one of the above.
[0015] A computer-readable storage medium storing a computer program, the computer program causes a computer to implement a method for international comparison of future industrial innovation systems based on hypergraph motifs as described in any one of the above when executed.
[0016] A computer program product includes a computer program or instructions, the computer program or instructions implement a method for international comparison of future industrial innovation systems based on hypergraph motifs as described in any one of the above when executed by a processor.
[0017] The present invention has the following beneficial effects:
[0018] 1. For the complex system of the national innovation system in future industries, by introducing network analysis tools to collect and analyze cooperation events among institutions in different countries, the non-linear interaction relationships among institutions such as national laboratories, high-level research universities, and leading technology enterprises can be characterized more accurately.
[0019] 2. By introducing the hypergraph structure, the assumption of pairwise interaction of nodes in the classical network structure is extended to higher-order interactions among multiple nodes. Technically, this can reduce the estimation error, more accurately characterize the impact of the local topological structure, i.e., the motif, on the network evolution process, and is more conducive to understanding the changes in the interaction patterns among future industrial innovation systems over time.
[0020] 3. By introducing network motifs, the interaction patterns among multiple innovation entities in the entire industry can be intuitively characterized, and then the main innovation systems driving technological breakthroughs in future industries can be determined. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 is a flowchart of an international comparison method for future industrial innovation systems based on hypergraph motifs provided by an embodiment of the present application; Figure 2 is a schematic diagram of six non-isomorphic connected sub-hypergraphs composed of 3 nodes in an embodiment of the present application; Figure 3 is a schematic diagram of an electronic device for implementing an international comparison method for future industrial innovation systems based on hypergraph motifs provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the technical solutions of the present application clearer, the following further elaborates on the present invention in detail with reference to the accompanying drawings and specific embodiments. The terms "first", "second", etc. in the claims and the description of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances. This is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product, or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products, or devices.
[0024] As Figure 1 shown in the figure, this embodiment provides an international comparison method for the future industrial innovation system based on hypergraph motifs, including the following steps: S110. Collect the original data set representing the cooperation relationships between different institutions in the target industry, and preprocess the original data set to obtain several institution cooperation data sets divided by country and the attributes of each institution; S120. Construct an institution cooperation hypergraph for each country based on the institution cooperation data set. The nodes in the institution cooperation hypergraph represent institutions and are attached with institution attributes, and the hyperedges represent cross-institution cooperation events; S130. For each institution cooperation hypergraph, use the enumeration method to determine the total number of k-order non-isomorphic connected sub-hypergraphs therein, where k is an integer greater than 1, and generate multiple random null models for it, and calculate the average value of the k-order non-isomorphic connected sub-hypergraphs in the multiple random null models; S140. Calculate the abundance of the k-order non-isomorphic connected sub-hypergraphs in each institution cooperation hypergraph according to each total number and its corresponding average value, and generate a significance profile to compare the target industrial innovation systems of different countries.
[0025] The first step is data collection and preprocessing.
[0026] First, determine the future industry to be investigated, that is, the target industry, and obtain the original data set that can represent the cooperation relationships between different institutions in the target industry for subsequent construction of the institution cooperation hypergraph.
[0027] Taking the co - author relationship between institutions in publicly available literature data such as papers and patents as an example, after determining the target industry, rely on domain expert knowledge to determine the keywords of the target industry. Using the keywords of the target industry as the retrieval conditions, retrieve relevant papers in publicly available data platforms such as OpenAlex, Dimensions, Web of Science, Scopus, PubMed, etc., and screen out the target papers according to the set document types. The set document types include journal papers (article), conference papers (conference), books (book), and book chapters (book chapter). At the same time, screen out the key information of the papers, including title, abstract, journal name (conference name), author, affiliated institution, publication year, DOI number, etc. Among them, OpenAlex is a comprehensive and freely accessible academic literature search engine and service, aiming to provide wide access to global academic achievements, including detailed information on entities such as papers, authors, institutions, journals, etc.; Dimensions is a modern research information platform developed by Digital Science, providing one - stop services from funding to publications to patents, covering a wide range of scientific research information types, including but not limited to journal articles, books, conference papers, patents, and clinical trial records; Web of Science is a large - scale citation index database maintained by Clarivate Analytics, one of the earliest scientific citation indexes, mainly collecting high - quality journal articles in the field of natural sciences and expanding to fields such as social sciences, arts, and humanities; Scopus is one of the world - leading abstract and citation databases created by Elsevier; PubMed is a free literature retrieval system in the fields of life sciences and biomedicine provided by the National Library of Medicine (NLM) of the United States, mainly focusing on journal articles in medicine and related disciplines, especially in the areas of life sciences and biomedicine; DOI (Digital Object Identifier) is a standardized system for identifying digital resources, such as academic papers, book chapters, datasets, etc. Each DOI number is unique, and it provides a persistent link to a specific digital resource. Usually, a DOI number consists of a prefix and a suffix, separated by a slash. The prefix starts with "10.", followed by one or more digits, which are usually related to the publishing institution, and the suffix is defined by the publisher itself to uniquely identify the resource. For example, a typical DOI number can be 10.1000 / 123456.
[0028] Furthermore, extract all the institution names contained in the original dataset and classify all institutions by country. Disambiguate all institutional names and label the attributes of each institution according to the research institution registry; Divide the original dataset into several institution cooperation datasets differentiated by country according to the classified institutions and the disambiguated institutional names, where each institution cooperation dataset contains the attributes of each institution.
[0029] Next, classify the co-authored institutions of the selected papers according to the country. Specifically, the co-authored institutions located in the same country are grouped into one category.
[0030] Before classification, if the co-authored institution of a certain paper is missing, first retrieve it again in other databases according to the key information of the paper. If the missing institutional information can be obtained, use it to fill in the missing value. If it cannot be obtained, retrieve the author with the same name in the collected data and use the co-authored institution of this author that is not missing in other data to fill in. If it still cannot be obtained, delete the data corresponding to the paper with the missing co-authored institution. This operation is mainly for the cooperation mode of co-authorship relationships and can be selected whether to use according to actual needs.
[0031] Then, disambiguate the institutional names again. Ambiguity of institutional names may occur in two cases. One is that the same institution corresponds to multiple institutional names, such as inconsistent abbreviations, former names, or incorrect signatures; the other is that different authors may sign different hierarchical sub-institutions of the same institution. In the case of disambiguation, for the first case, select one name as the standard name of the institution and uniformly use this standard name for the same institution; for the second case, it is agreed that different hierarchical sub-institutions of the same institution are uniformly signed as the standard name of the highest-level institution.
[0032] Finally, the organization types are labeled according to the public database ROR (Research Organization Registry), where the organization type is the organization attribute. ROR is a global open registry that aims to provide unique identifiers for research organizations and contains rich metadata information (such as organization name, address, type, etc.). Labeling the organization types based on ROR specifically includes: accessing the official website of ROR or downloading its open data. Each organization in ROR contains a unique ROR ID and related metadata such as name, address, and type. Then, traverse the ROR dataset, extract the types field of each organization. The types field is the field in the ROR database that clearly labels the organization type. According to the extracted organization type information, assign corresponding type labels to each organization. For example, educational institutions are labeled as Education, healthcare institutions are labeled as Healthcare, companies / enterprises are labeled as Company, and government institutions are labeled as Government, etc. If an organization has multiple types, for example, it is both an educational institution and a healthcare institution, multiple types can be labeled simultaneously. By labeling the organization types, a finer-grained perspective can be provided for innovation network analysis, thereby revealing the roles, cooperation patterns of different types of organizations in the network, and their impacts on the overall system. This not only helps to understand the current innovation system but also provides a scientific basis for subsequent policy-making and network optimization.
[0033] Extract the cooperation events between organizations within each country from the original dataset to obtain several organization cooperation datasets distinguished by country, and each dataset records the organization attributes.
[0034] Data preprocessing can help identify and correct potentially incorrect or inconsistent data, ensuring the quality of the subsequent generated hypergraph.
[0035] For the complex system of the national innovation system in future industries, in this embodiment, by introducing network analysis tools to collect and analyze the cooperation events between organizations in different countries, the non-linear interaction relationships between organizations such as national laboratories, high-level research universities, and leading technology enterprises can be characterized more accurately.
[0036] The second step is to construct the cooperation hypergraph.
[0037] Construct an organization cooperation hypergraph corresponding to each country according to the cooperation relationships between organizations within each country , and each cooperation hypergraph represents the cooperation events between organizations within a country. Among them, represents the node set, and each node represents an organization, represents the number of nodes, represents the set of hyperedges, Each element in represents a hyperedge, which is mathematically represented as a non-empty subset of the node set, that is , represents the number of hyperedges, and and are both greater than 1. It can also be represented by an incidence matrix: , when , , conversely, . Each node 's incident hyperedges are defined as the set of all hyperedges containing the node , that is . If the intersection of two hyperedges is not empty, that is, the two hyperedges contain common nodes, then the two hyperedges are said to be adjacent.
[0038] In the co-authorship relationship, a paper can be regarded as a hyperedge composed of its publishing institutions, and the number of nodes contained in this hyperedge is the number of publishing institutions of this paper. Suppose in a certain country, paper A is jointly completed by institution X and institution Y, paper B is jointly completed by institution Y, institution Z, and institution W, and paper C is only completed by institution X. Then the hyperedge corresponding to paper A contains the nodes , the hyperedge corresponding to paper B contains the nodes , and the hyperedge corresponding to paper C only contains the node . Then in the constructed institutional cooperation hypergraph , the node set , and the hyperedge set .
[0039] By constructing the institutional cooperation hypergraph of each country, it is possible to analyze the cooperation patterns among institutions within each country in more detail, which is beneficial to subsequent cross-country comparative studies and facilitates understanding the differences in cooperation patterns, cooperation intensities, and cooperation types between different countries.
[0040] In this embodiment, by introducing the hypergraph structure, the assumption of pairwise interaction of nodes in the classical network structure is generalized to higher-order interactions among multiple nodes. Technically, this can reduce the estimation error, more accurately characterize the impact of the local topological structure, i.e., the motif, on the network evolution process, and is more conducive to understanding the changes in the interaction patterns among future industrial innovation systems over time.
[0041] The third step is motif detection.
[0042] After obtaining the institutional cooperation hypergraphs corresponding to each country respectively After that, the number of k-order motifs in each institutional collaboration hypergraph is detected, where k is an integer greater than 1, representing the number of nodes contained in the sub-hypergraph. Considering the practical significance and computational complexity, k is usually taken as 3, 4, or 5. In a hypergraph, a k-order motif is defined as a connected subgraph whose quantity is statistically significantly non-random. These subgraphs contain k nodes and are interconnected by higher-order interactions of any order.
[0043] In any institutional collaboration hypergraph in for any subset of nodes define the open neighborhood as the set of all nodes that come from and are adjacent to at least one of the nodes in , where represents the set composed of the remaining elements after removing the subset of nodes from the node set . Here, the symbol " " is the symbol for the set difference operation. Suppose in , the node set is , and the hyperedge set is . If the subset of nodes , then , and its open neighborhood ; for any node , the exclusive neighborhood of node relative to is the set of all nodes that are adjacent to but do not belong to the union of and . Taking the above hypergraph as an example, . At the same time, for any subset of nodes , all the hyperedges in the induced sub-hypergraph formed by these nodes must completely retain all the hyperedges existing between these vertices in the original hypergraph. That is, in the induced sub-hypergraph, only the hyperedges that already exist in the original hypergraph and are completely contained within will be retained. Taking the above hypergraph as an example again, because contains the three nodes , and there is a hyperedge between these three nodes, so the node set in the finally formed induced sub-hypergraph is , and six non-isomorphic connected sub-hypergraphs are formed, whose edge sets are respectively or or , such asFigure 2 As shown in Pattern ①, since these three cases are isomorphic, that is, there is no essential difference in structure, they are grouped into one category. For example, Figure 2 as shown in Pattern ① in [reference], when enumerating, these isomorphic sub-hypergraphs will all be enumerated and counted in the number of sub-hypergraphs of the same category. For example, here Pattern ① is counted as 3, and the same principle applies to the following examples; , such as Figure 2 as shown in Pattern ② in [reference], , such as Figure 2 the pattern in [reference] shown, or or , such as Figure 2 as shown in Pattern ④ in [reference], or or , such as Figure 2 as shown in Pattern ⑤ in [reference], , such as Figure 2 as shown in Pattern ⑥ in [reference].
[0044] Taking co-authorship of papers as an example to illustrate Figure 2 the meanings represented by the six patterns in [reference]: Suppose there are three authors A, B, and C. One case of Pattern ① is that A and B have collaborated, B and C have collaborated, but A and C have not collaborated, and A, B, and C have never collaborated simultaneously; Pattern ② means that A and B have collaborated, B and C have collaborated, A and C have collaborated, but A, B, and C have never collaborated simultaneously; Pattern ③ means that A, B, and C have only published when they collaborate simultaneously, and there is no situation where A and B collaborate and publish alone, B and C collaborate and publish alone, or A and C collaborate and publish alone; Pattern ④ means that A, B, and C have published when they collaborate simultaneously, and it also includes the situation of two-person solo collaborations (A + B or A + C or B + C); Pattern ⑤ means that A, B, and C have published when they collaborate simultaneously, and it also includes the situation of two groups of two-person solo collaborations; Pattern ⑥ means that A, B, and C have works published when they collaborate simultaneously, and at the same time, there are also works published when each pair of the three collaborate.
[0045] At the same time, in the study of network structures, ordinary graphs can only describe pairwise interactions between nodes. For example, for a connected ordinary graph with 3 nodes, there are only two induced patterns: tree-like paths and cyclic triangles. However, hypergraphs can depict higher-order interactions, and hyperedges can connect any number (≥2) of nodes, which makes the number of its induced sub-hypergraphs increase rapidly with the number of nodes. Taking k = 3 as an example, the number of its connected induced hypergraphs can reach 6, while for 4 nodes, it can reach 171.
[0046] For a hypergraph containing k nodes, it is very difficult to directly solve its analytical expression. Instead, we can consider estimating the upper and lower bounds of the number m of all possible induced connected sub-hypergraphs it contains:
[0047] When estimating the upper bound, we first relax the "non-isomorphism" and "connectivity" constraints. Among k nodes, when only considering hyperedges with a node count of at least 2, their sum is , which is obtained by subtracting the sets composed of single nodes and the empty set from the total number of subsets of k nodes . When considering the existence of each hyperedge (i.e., whether it is labeled), the total number of sub-hypergraphs at this time is , and this quantity is the upper bound of m.
[0048] When estimating the lower bound, we consider the "non-isomorphism" and "connectivity" constraints. Among k nodes, first use k - 1 connecting edges to ensure the connectivity between nodes. At this time, there may be hyperedges formed. Further considering the existence of each hyperedge (i.e., whether it is labeled), at least sub-hypergraphs can be formed. Since each sub-hypergraph corresponds to at most label permutation ways, and different label permutation ways form isomorphic sub-hypergraphs. At this time, the lower bound of m is .
[0049] When calculating the non-isomorphic connected sub-hypergraphs formed by the final k nodes, in addition to considering the case of high-order interactions, it is also necessary to take into account low-order interaction patterns, including the pairwise pairing interaction patterns of ordinary graphs. Taking 3 nodes as an example, it contains 4 high-order interaction patterns as shown in ③, ④, ⑤, and ⑥ in Figure 2 and 2 low-order interaction patterns as shown in ① and ② in Figure 2 . Among them, the number 2 of low-order interaction patterns is obtained by subtracting the sets composed of single nodes and the empty set from the total number of subsets of 3 nodes and ensuring the connectivity of k - 1 = 3 - 1 = 2 connecting edges. Finally, it is calculated according to ; while the number 4 of high-order interaction patterns is based on a hyperedge and is obtained by combining isomorphic patterns according to the pairwise pairing ( ) situation among the three nodes.
[0050] Therefore, to enumerate all sub-hypergraphs of size k in each institutional cooperation hypergraph, the specific steps include:
[0051] 1. Enumerate all hyperedges of size k. These hyperedges exactly contain k nodes and can directly form a k-order motif, which can be regarded as the basic induced sub-hypergraph of a k-order motif. At the same time, for these k nodes, it is also necessary to consider their low-order connected graphs of order less than k, and its calculation formula is . Add the total number of sub-hypergraphs obtained in these two cases to get the number of basic induced sub-hypergraphs of the k-order motif.
[0052] 2. Consider other hyperedges with the number of included nodes less than k. Since the number of nodes in these hyperedges is less than , a k-order motif cannot be directly formed. Therefore, neighbor nodes need to be added to the node sets included in these hyperedges until there are k nodes in these node sets. Among them, if only one node needs to be added, this node must belong to the union of the exclusive neighborhood of the nodes included in the original hyperedge and the adjacent hyperedges of the original hyperedge; if 2 or more nodes need to be added, the added nodes must be selected from the adjacent hyperedges of the original hyperedge, and the latter added node needs to be the exclusive neighborhood of the previous added node.
[0053] 3. After all hyperedges with the number of nodes less than k also contain k nodes, generate the power set of these k nodes, indicating that any combination of these nodes may generate hyperedges, and retain the part of these hyperedges that exists in the original hypergraph . These parts are the new induced sub-hypergraphs. Then, merge the isomorphic types among them and combine them with the sub-hypergraph obtained in 1 to get all non-isomorphic connected sub-hypergraphs of the entire hypergraph.
[0054] Still taking the above hypergraph as an example, assuming k = 3, in this hypergraph, only the hyperedge contains 3 nodes. Therefore, it directly forms a 3-order motif. Let , or or ( Figure 2 Pattern ①). At the same time, as can be seen from the above, 3 nodes can form a total of six non-isomorphic connected sub-hypergraphs, which are or or , , , or or , or or , . Therefore, the non-isomorphic connected sub-hypergraphs of this 3-order motif are , , , , , .
[0055] Next, consider the hyperedges that contain less than 3 nodes, and , they need to be extended to 3 nodes by adding neighbor nodes. For the hyperedge , select as the newly added node because is and 's common neighbor and satisfies the exclusive neighborhood condition; for the hyperedge , select as the added node because is and 's common neighbor and satisfies the exclusive neighborhood condition; for the hyperedge , select as the added node because satisfies the adjacent hyperedge and exclusive neighborhood conditions. Therefore, for the hyperedges and , the same node set is finally extended; for the hyperedge , the finally extended node set . For the node set , its power set , among these combinations, exists in , exists in , is just 's part and does not form an independent hyperedge. Therefore, the finally extended 3 - order connected sub - hypergraph , , , that is is isomorphic to ; for the node set , its power set , among these combinations, only the hyperedge completely exists in in . Therefore, the finally extended 3 - order connected sub - hypergraph , 。
[0056] That is to say, for the hypergraph , the node set , the hyperedge set , it includes 6 non - isomorphic 3 - order connected sub - hypergraphs, where the number of pattern ① is 3, the number of pattern ② is 1, and the number of pattern ③ is 3 ( + + ), the number of Pattern ④ is 6 ( + + + ), the number of Pattern ⑤ is 4 ( + ), and the number of Pattern ⑥ is 1.
[0057] Meanwhile, corresponding to each institutional collaboration hypergraph, n random hypergraphs are generated as null models, where n is an integer greater than 1. The null model is a hypergraph that retains some key structural features of the original hypergraph (such as degree distribution, hyperedge size distribution), but is generated by randomizing other features. It is used as a control to help identify non-random and statistically significant structural patterns in the original hypergraph. In this embodiment, the Chodrow configuration model that retains the degree distribution in the original hypergraph and the node number distribution in the hyperedges is used to generate random hypergraphs. Then, the number of k - order non - isomorphic connected sub - hypergraphs in each random hypergraph is determined by the enumeration method, and the average number of k - order non - isomorphic connected sub - hypergraphs in these n random hypergraphs is calculated , n is an integer greater than 1. In this embodiment, n takes 1000.
[0058] In this embodiment, n null - model random hypergraphs corresponding to each of the above 6 patterns are respectively generated using the Chodrow configuration model, and then the average values of the 3 - order non - isomorphic connected sub - hypergraphs in the n null - model random hypergraphs corresponding to these 6 patterns are calculated respectively.
[0059] Furthermore, according to the total number of k - order non - isomorphic connected sub - hypergraphs in each institutional collaboration hypergraph and its corresponding average value, the abundance of k - order non - isomorphic connected sub - hypergraphs in each institutional collaboration hypergraph is calculated and a significance profile is generated, including: Calculating the difference and sum between each total number and its corresponding average value respectively; Taking the sum obtained plus a preset adjustment factor as the denominator and the difference as the numerator to calculate the abundance of k - order non - isomorphic connected sub - hypergraphs in the corresponding institutional collaboration hypergraph; Normalizing the abundance to obtain its significance profile.
[0060] According to the total number of k - order non - isomorphic connected sub - hypergraphs in the institutional collaboration hypergraph and the average number of k - order non - isomorphic connected sub - hypergraphs in its corresponding n random hypergraphs, the statistical significance of the k - order non - isomorphic connected sub - hypergraphs is calculated, and its statistical significance is represented by the abundance and the calculation formula is: .
[0061] Among them, denotes the number of non-isomorphic connected sub-hypergraphs of order k in the institutional cooperation hypergraph. In this embodiment, k = 3, is an adjustment factor, usually set to 4. According to different requirements, it is also possible to further calculate the number and type of motifs participated by each node, i.e., institution, as well as the number and type of motifs contained in each type of node.
[0062] Since different networks usually have different scales and degree sequences, it is impossible to directly compare the structural similarities of different networks. Therefore, it is necessary to normalize to obtain the significance profile of each type of induced sub-hypergraph . By normalization, the influence of network size and node degree distribution differences on the motif analysis results is eliminated, enabling effective comparison between different networks based on their motif compositions.
[0063] Finally, different networks are compared according to the significance profiles of non-isomorphic connected sub-hypergraphs of order k. Specifically, the calculation results of the significance profiles are presented in the form of a line chart or a correlation coefficient matrix, which can be used to compare the similarities and differences in the innovation system models presented by different countries in different non-isomorphic connected sub-hypergraphs of order k. In this embodiment, it is mainly used to reveal the influence of different innovation system models on innovation achievements and industrial stability. For example, through analysis, it is found that certain motif types, such as triangular motifs, i.e., three nodes connected pairwise, can promote information sharing and technological exchanges, thus driving the innovation process. At the same time, subjective insights from domain experts are introduced to interpret the results. For example, experts can provide insights on why certain motif types are more common or rare based on factors such as industry characteristics and historical events. Then, based on expert interpretations and data analysis results, targeted policy recommendations are proposed. For example, encourage the formation of more intermediary nodes to enhance knowledge diffusion; or support the establishment of redundant supply chains to improve industrial resilience.
[0064] In this embodiment, by introducing network motifs, the interaction patterns among multiple innovation entities in the entire industry can be intuitively characterized, and then the main innovation systems that drive future industrial technological breakthroughs can be determined. For example, for a motif in the shape of a closed triangle, and the three nodes correspond to the government, industry, and academia respectively, it can be inferred that government-industry-academia collaborative innovation is an important model in the national innovation system that drives technological breakthroughs in this industry of this country.
[0065] As Figure 3 shown, this embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above-mentioned international comparison method for future industrial innovation systems based on hypergraph motifs.
[0066] The computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes the steps of an international comparison method for the future industrial innovation system based on hypergraph motifs.
[0067] This embodiment also provides a computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction is executed by the processor, it realizes the steps of an international comparison method for the future industrial innovation system based on hypergraph motifs as described above.
[0068] This embodiment also provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by the processor, it realizes the steps of an international comparison method for the future industrial innovation system based on hypergraph motifs as described above.
[0069] These computer-readable programs / instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, thereby producing a machine. When these instructions are executed by the processor of the computer or other programmable data processing devices, a device is produced that realizes the functions / actions specified in one or more boxes in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium. These instructions cause the computer, the programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article, which includes instructions for realizing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0070] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A method for international comparison of future industrial innovation systems based on hypergraph motifs, characterized by: The following steps are involved: Collecting original data sets representing the cooperative relationships between different institutions in the target industry, and preprocessing the original data sets to obtain several institutional cooperation data sets divided by country and the attributes of each institution; Based on the institutional cooperation dataset, an institutional cooperation hypergraph is constructed for each country, wherein nodes in the institutional cooperation hypergraph represent institutions and are accompanied by institutional attributes, and hyperedges represent cross-institutional cooperation events; For each institutional cooperation hypergraph, the total number of k-order non-isomorphic connected sub-hypergraphs is determined by enumeration method, where k is an integer greater than 1, and multiple random null models are generated for it, and the average value of the k-order non-isomorphic connected sub-hypergraphs in the multiple random null models is calculated; According to each of the totals and their corresponding average values, the abundance of the k-order non-isomorphic connected sub-hypergraphs in each institutional cooperation hypergraph is calculated and a significance profile is generated to compare the target industry innovation systems of different countries.
2. According to claim 1, the international comparison method of future industrial innovation system based on hypergraph motif is characterized by: The preprocessing of the original data set obtains a number of institutional cooperation data sets divided by country and the attributes of each institution, including: Extract all the names of institutions contained in the original dataset and classify all institutions by country; Disambiguate all institution names and annotate each institution’s attributes according to the research institution registry; The original data set is divided into a number of institutional cooperation data sets differentiated by country according to the classified institutions and the disambiguated institution names, and the institutional cooperation data sets contain attributes of each institution.
3. The international comparison method of future industrial innovation system based on hypergraph motif according to claim 1 is characterized in that: The method of using enumeration to determine the total number of k-order non-isomorphic connected sub-hypergraphs includes: Enumerate all hyperedges containing k nodes in the organization cooperation hypergraph to obtain the number of basic non-isomorphic connected sub-hypergraphs of the k-order motif; For hyperedges with less than k nodes, expand their number of nodes to k, and generate an extended k-order non-isomorphic connected sub-hypergraph based on the expanded node set; The number of basic non-isomorphic connected sub-hypergraphs of the k-order motif is added to the number of the extended k-order non-isomorphic connected sub-hypergraphs to obtain the total number of k-order non-isomorphic connected sub-hypergraphs.
4. The international comparison method of future industrial innovation system based on hypergraph motif according to claim 3 is characterized in that: The method expands the number of nodes to k, including: If a node is added to the node set contained in a hyperedge with less than k nodes, the number of nodes will be expanded to k, and the added node exists in the union of the exclusive neighborhood of the node contained in the hyperedge and its adjacent hyperedge; If at least two nodes are added to the node set included in a hyperedge with less than k nodes, the number of nodes is expanded to k, then the added nodes belong to its adjacent hyperedge, and the nodes added later are the exclusive neighbors of the nodes added earlier.
5. According to claim 4, the international comparison method of future industrial innovation system based on hypergraph motif is characterized in that: The step of generating a k-order non-isomorphic connected sub-hypergraph according to the expanded node set includes: Generate a power set of k nodes included in the expanded node set, determine the combinations in the power set that exist in the corresponding organization cooperation hypergraph, and generate an extended k-order non-isomorphic connected sub-hypergraph based on this.
6. The international comparison method of future industrial innovation system based on hypergraph motif according to claim 1 is characterized in that: Before calculating the average values of the k-order non-isomorphic connected sub-hypergraphs in the multiple random zero models, the method further includes: The number of k-order non-isomorphic connected sub-hypergraphs in each random null model is determined by the enumeration method described in claims 3-5.
7. The international comparison method of future industrial innovation system based on hypergraph motif according to claim 1 is characterized in that: The method of calculating the abundance of the k-order non-isomorphic connected sub-hypergraphs in each institution cooperation hypergraph according to each total number and its corresponding average value and generating a significant profile includes: Calculate the difference and sum of each total and its corresponding average value respectively; The obtained sum is added to the preset adjustment factor as the denominator, and the difference is used as the numerator to calculate the abundance of the k-order non-isomorphic connected sub-hypergraph in the corresponding institution cooperation hypergraph; The abundances were normalized to obtain their significance profiles.
8. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement an international comparison method for future industrial innovation systems based on a hypergraph model as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that: The computer program enables the computer to implement an international comparison method for future industrial innovation systems based on a hypergraph model as described in any one of claims 1 to 7 when executed.
10. A computer program product, comprising a computer program or instructions, which, when executed by a processor, implements an international comparison method for future industrial innovation systems based on a hypergraph model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Complex network key motif mining method based on multi-attribute decision
CN113486217A
Network similarity evaluation method based on network motif
CN115766476A
Future industry innovation direction identification method and system based on deep learning
CN116842182A
Method and device for predicting future industry direction by using hybrid intelligence
CN118656713A