A research and development partner recommendation method based on a technology innovation knowledge situation super network
By constructing a hypernetwork model of technological innovation knowledge context, the problems of knowledge dispersion and unreasonable selection of R&D partners in engineering construction projects are solved, and efficient R&D partner recommendation and knowledge sharing are realized.
Patent Information
- Application Number
- CN202310128431.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-02-16
AI Technical Summary
In engineering construction projects, existing technologies are unable to effectively collect and analyze scattered technological innovation knowledge, lack attention to the knowledge context, and lead to unreasonable selection of R&D partners, making it impossible to quickly respond to engineering problems.
We construct a hypernetwork model of knowledge context for technological innovation, and recommend suitable R&D partners by extracting multidimensional knowledge context information and using an improved hypernetwork Bayesian inference method.
It improved the accuracy and efficiency of R&D partner recommendations, solved the problems of knowledge dispersion and low utilization efficiency, and promoted the sharing and reuse of engineering and technological innovation knowledge.
Smart Images

Figure CN116304308B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information recommendation, and particularly relates to a research and development partner recommendation method based on a technical innovation knowledge situation super network. BACKGROUND
[0002] With the development of knowledge economy, more and more engineering construction enterprises realize the importance of knowledge and technical innovation. In the process of technical innovation, engineering construction enterprises can produce new knowledge and new technology with the help of the knowledge and experience of research and development partners. Since engineering construction enterprises are mostly project-oriented, a large number of technical innovation activities exist in engineering construction projects. The characteristics of engineering construction projects lead to the fact that these technical innovation activities often have the characteristics of shorter research and development cycle and higher application requirement of achievements. It can be seen that when urgent engineering problems occur in the project, the enterprise needs to quickly select appropriate research and development partners to solve the problems. At this time, the enterprise needs to evaluate the experience and ability of the research and development institutions to solve certain engineering problems according to their past accumulation of technical innovation knowledge. Therefore, the analysis and utilization of past technical innovation knowledge is a key factor for the selection of appropriate research and development partners, and directly affects the success or failure of technical innovation activities.
[0003] However, engineering construction projects are temporary, and the information of the architecture, engineering and construction (AEC) industry is scattered. The knowledge generated by a large number of technical innovation activities is scattered in various information entities of multiple sources, multiple subjects and multiple organizations, and it is difficult to collect and uniformly analyze the knowledge of each research and development institution. At the same time, the knowledge generated by engineering technical innovation is generated in a specific situation, and if the knowledge situation is not paid attention to, the related enterprises or individuals cannot fully understand the knowledge, and it is difficult to capture and reuse the knowledge. In addition, the demand of enterprises for knowledge is dynamic and related to the situation of tasks and problems. Therefore, the ability of research and development institutions to solve the same type of scientific research demand in different situations is different, and the knowledge cannot be directly and simply migrated and reused. In summary, when selecting research and development partners, not only the matching degree of their experience knowledge needs to be considered, but also the technical innovation knowledge situation and the engineering problem situation need to be combined. In this way, it is evaluated whether the experience knowledge of the research and development institution can be used for the current engineering situation, and whether the institution is suitable for undertaking the technical innovation work. However, at present, the selection of research partners often relies on experience, and there is no suitable method to reasonably evaluate and recommend research partners.
[0004] Prior art one (CN201410476082.7) discloses a method and device for matching projects and professionals, which analyzes the correlation between different keywords based on the keyword data of projects and professionals, establishes a keyword correlation network, quantifies the correlation degree of projects and professionals, quantifies the correlation degree of projects and professionals that are difficult to contact, and can realize the recommendation of different matching such as cultivation experts and review experts by customizing the weight of direct and indirect contact.
[0005] The method has the following disadvantages:
[0006] 1. The project information is not fully analyzed: the method traverses the project record, calculates the number of all keywords contained in the field corresponding to the field keyword, and only uses the intersection of the field keywords in the project information to measure the field correlation, which easily loses the semantic information and context information in the project.
[0007] 2. The characterization of professionals is not reasonable: the method uses expert degree, expert level and expert index to characterize experts, and uses H-index and PageRank index. These two indexes are mainly used to characterize the academic level of experts, but this index can only reflect the overall academic level of experts, and cannot deeply reflect their ability in a certain field. Therefore, some experts with many academic achievements may be frequently recommended, while some experts with less overall achievements but expertise in a certain field cannot be identified.
[0008] Prior art two (CN202010914996.2) is a method, device and equipment for recommending experts based on paper data analysis, and a storage medium. The invention calculates the recommendation score of candidate experts in three dimensions of text similarity, author contribution rate and paper impact factor. The higher the score, the more likely it is that the candidate expert is the most needed technical expert of the demand side, and finally realizes expert recommendation, significantly improving the recommendation accuracy and efficiency of expert recommendation.
[0009] The method has the following disadvantages:
[0010] 1. The recommendation method is not perfect: the method proposes that the recommendation score of experts is determined by three dimensions of text similarity, author contribution rate and paper impact factor. Enterprises often have complex processes and situations when encountering specific technical problems. This method only considers the similarity of the paper abstract information and a specific project or problem text, as well as the impact factor and paper contribution of the expert's papers, ignoring the research and development situation of these experts in some projects, and cannot depict the deep relationship between experts and research needs. The technical experts recommended may not be the most suitable. SUMMARY
[0011] In view of the above deficiencies in the prior art, the present application provides a research and development partner recommendation method based on a technology innovation knowledge situation hypernetwork.
[0012] In order to achieve the above-mentioned purposes, the technical scheme adopted by the present application is as follows:
[0013] A research and development partner recommendation method based on a technology innovation knowledge situation hypernetwork comprises the following steps:
[0014] S1, acquiring technology innovation project data;
[0015] S2, constructing a technology innovation knowledge situation model and determining technology innovation knowledge situation dimensions;
[0016] S3, based on the technology innovation knowledge situation dimensions, extracting multi-dimensional knowledge situation information from the technology innovation project data;
[0017] S4, based on the multi-dimensional knowledge situation information, constructing a technology innovation knowledge situation hypernetwork model and determining node features and hyperedge features;
[0018] S5, according to the node features and the hyperedge features, generating a research and development partner recommendation result by using an improved hypernetwork Bayesian inference method.
[0019] The present application has the following beneficial effects:
[0020] (1) The present application designs an automatic extraction method of engineering technology innovation knowledge situations, which can automatically extract the knowledge situations of engineering technology innovation activities from massive unstructured documents. This method can effectively improve the utilization efficiency of knowledge, solve the problems of scattered knowledge and knowledge flooding, and is conducive to the sharing and reuse of engineering technology innovation knowledge.
[0021] (2) The present application maps the knowledge situation information into a hypernetwork model, constructs a four-layer hypernetwork of field-technology innovation-expert-research and development institution, and designs an improved hypernetwork Bayesian algorithm to recommend research and development institutions. In this algorithm, a node similarity calculation method combining field ontology, standardized mutual information and deep learning algorithm is designed, and a node influence weight index is proposed, which effectively excavates the deep relationships between knowledge situations in the hypernetwork, thereby significantly improving the recommendation accuracy and efficiency of research and development partner recommendation. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 FIG. 1 is a flowchart of a research and development partner recommendation method based on a technology innovation knowledge situation hypernetwork according to an embodiment of the present application;
[0023] Figure 2 FIG. 2 is a framework diagram of a research and development partner recommendation method based on a technology innovation knowledge situation hypernetwork according to an embodiment of the present application;
[0024] Figure 3 A framework diagram for extracting the knowledge context dimension in the embodiment of the present application is shown in the figure.
[0025] Figure 4 A technical innovation knowledge context super network model diagram in the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0026] The specific embodiments of the present application are described below to facilitate the understanding of the present application by those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, any changes that are obvious within the spirit and scope of the present application as defined and determined by the appended claims are within the scope of the present application.
[0027] As shown in Figure 1 and Figure 2 The embodiment of the present application provides a research and development partner recommendation method based on a technical innovation knowledge context super network, which includes the following steps S1 to S5:
[0028] S1, acquiring technical innovation project data;
[0029] In an optional embodiment of the present application, the recording methods of the engineering field technical innovation knowledge context are relatively diverse, mainly including published papers, applied patents and fund projects, and internal engineering case documents of enterprises. Therefore, the embodiment collects the technical innovation knowledge of related enterprises, including papers, patents, fund projects, and engineering cases, to obtain the technical innovation project data.
[0030] S2, constructing a technical innovation knowledge context model and determining the technical innovation knowledge context dimension;
[0031] In an optional embodiment of the present application, the embodiment combines the characteristics of the specific engineering field technical innovation knowledge, decomposes the business process of engineering technical innovation, sorts out the logical process of knowledge output and records, and finally determines the dimension of the engineering field technical innovation knowledge context. The context here refers to information that can represent the environment, state, and situation of an entity, where the entity can be a person or a resource itself, or a related event or project. The knowledge context can be understood as a body of information that explains or gives meaning to words, thoughts, and actions.
[0032] In the engineering field, the generation and application of knowledge have specific backgrounds and environments; that is, different knowledge has different contexts. The purpose of knowledge context modeling is to establish a scalable knowledge description model for accurately sharing and applying knowledge in complex and ever-changing engineering environments. Experts solve technical problems encountered during engineering construction through technological innovation. When relevant personnel encounter similar knowledge contexts, they can search a knowledge context database to find matching knowledge resources. This invention, combining the characteristics of technological innovation knowledge in the engineering field, identifies eight dimensions of engineering technological innovation knowledge contexts: domain, object, process, result, expert, institution, location, and time. Formalized as:
[0033] KC={Domain,Subject,Process,Results,
[0034] Expert,Institution,Location,Time}
[0035] Where Results =<Name,Category> This consists of the title and type of the output. The meaning of each dimension is explained below:
[0036]
[0037] S3. Based on the knowledge context dimension of technological innovation, extract multi-dimensional knowledge context information from technological innovation project data;
[0038] In an optional embodiment of the present invention, the knowledge generated by technological innovation is produced within a specific context. Without attention to the knowledge context, a few keywords alone are insufficient to fully analyze project information. Furthermore, enterprises' knowledge needs are dynamic and context-dependent. Therefore, R&D institutions have varying capabilities to address the same type of research needs in different contexts, and knowledge cannot be directly and simply transferred and reused. Thus, this invention introduces the theory of knowledge context and designs an automated method for extracting technological innovation knowledge context, using knowledge context to represent the knowledge information of technological innovation projects.
[0039] This embodiment first preprocesses and standardizes the acquired raw document data, then uses named entity recognition methods to extract relevant dimensions of knowledge context from the text, and finally stores these dimension data into the knowledge context database.
[0040] Step S3 specifically includes:
[0041] Data on technological innovation projects is categorized into structured information, unstructured text, and semi-structured text.
[0042] The structured information is directly mapped to the achievement dimension, the expert dimension, the institution dimension, and the time dimension of the technical innovation knowledge situation model;
[0043] The unstructured text is preprocessed by word segmentation and part-of-speech tagging, then encoded by BIOES to annotate the corpus, and then the preprocessed text is embedded by the Word2Vec model to obtain the vector expression of words. Then the BiLSTM-CRF model is used to automatically extract the domain, object, and location entities. Finally, the extracted domain, object, and location entities are mapped to the domain dimension, object dimension, and location dimension of the technical innovation knowledge situation model.
[0044] The semi-structured text is segmented and word segmented, then data cleaning is performed, and finally the processed text is mapped to the process dimension of the technical innovation knowledge situation model.
[0045] The specific extraction process is as follows:
[0046] 1. Structured information extraction: The structured information in the source data, such as name, type, time, expert, research institution, etc., is directly mapped to the corresponding dimension of the knowledge situation;
[0047] 2. Unstructured text extraction: For unstructured text such as technical background and article abstract, the main extraction process includes (1) data preprocessing: using Chinese word segmentation tool for word segmentation and part-of-speech tagging. (2) corpus annotation: using BIOES encoding to annotate the corpus, and annotating the domain, object, and location entities in the unstructured text. (3) word embedding: using the Word2Vec model to embed the preprocessed text to obtain the vector expression of words. (4) entity extraction: using the BiLSTM-CRF model to automatically extract the three types of entities (domain, object, and location), and then mapping the extracted three types of entities to the dimensions of the knowledge situation, which are represented by nouns or noun sets.
[0048] 3. Technical text content extraction: For the technical text content in the semi-structured text, including the knowledge generation process, solution, research process, and process flow, the main extraction process includes (1) word segmentation: using Chinese word segmentation tool for word segmentation. (2) data cleaning: using natural language processing toolkit to remove html tags, abnormal characters, redundant characters, useless symbols, and irrelevant numbers in the text. (3) mapping: mapping the processed text directly to the process dimension of the knowledge situation.
[0049] 4. Knowledge situation dimension extraction: According to the above four steps, the complete knowledge situation information of the eight dimensions of achievement, time, expert, institution, domain, object, location, and process is obtained, and the knowledge situation dimension extraction is completed, such as Figure 3as shown.
[0050] S4, constructing a technical innovation knowledge situation super network model based on the multi-dimensional knowledge situation information, determining node features and super edge features;
[0051] In an optional embodiment of the present application, the characterization of the research institutions and experts cannot rely only on the paper text information and some field keywords, and needs to mine their technical capabilities in different situations from the historical technical innovation activities and project information. The knowledge generated by the research institutions and experts in a large number of technical innovation activities is scattered in various information entities of multiple sources, multiple subjects and multiple organizations, and it is difficult to collect and uniformly analyze the knowledge of each research institution and expert. Therefore, the present application uses the project information in which the research institutions have participated to extract the situation entities of each dimension as their historical situations, so as to characterize their ability to solve a certain problem in a certain situation. On this basis, the present application proposes a knowledge situation super network model and designs four-layer sub-networks to store the knowledge situation information by using a multi-layer heterogeneous network.
[0052] Step S4 specifically includes:
[0053] constructing a technical innovation knowledge situation super network model including a field sub-network, a technical innovation sub-network, a research institution sub-network and an expert sub-network with a four-layer structure;
[0054] The field sub-network takes each field entity of the knowledge situation as a node, and determines the association relationship between the field entities according to the field entity concept tree;
[0055] The technical innovation sub-network constructs a technical innovation node for each knowledge situation, in which the achievements, objects, processes, locations and times are taken as the attributes of the technical innovation node, and the semantic association relationship of the technical innovation knowledge output in different situations is taken as the undirected edge;
[0056] The research institution sub-network takes each research institution as a node, and takes the cooperation relationship between the research institutions as the undirected edge;
[0057] The expert sub-network takes each expert as a node, and takes the cooperation relationship between the experts as the undirected edge.
[0058] The node influence weight in the technical innovation knowledge situation super network model is calculated respectively, and the weighted node super degree and the multi-node super degree are calculated according to the node influence weight.
[0059] Specifically, the present embodiment converts the multi-dimensional knowledge situation information in the knowledge situation library into a super network based on the topological structure and super edge characteristics of the super network, that is, a super edge and the nodes connected by the super edge are used to represent a piece of situation information. Further, the related features of the constructed super network are calculated.
[0060] Since hyperedges in a hypergraph can include any number of nodes, hypernetworks based on hypergraphs can reflect the coexistence relationships between multiple nodes and better represent the mutual influence and interaction between nodes. To reflect the hierarchical heterogeneous relationships of knowledge context information and to incorporate dimensions as nodes into the computation, this invention selects hypernetworks as the storage method for knowledge context. This invention constructs a hypernetwork for engineering and technological innovation knowledge context, comprising four subnetworks: domain subnetwork, technological innovation subnetwork, R&D institution subnetwork, and expert subnetwork, through analysis of the dimensions of engineering and technological innovation knowledge context. Here, the definition of a hypergraph is: Let V = {v1, v2, ..., v...} n Let} be a finite set, and let where i∈(1,m) and Let E = {E1, E2, ..., E} n If H = (V, E) is a finite hypergraph, then V is the set of nodes in the hypergraph, and E is the set of hyperedges in the hypergraph, where each element is a node of the hypergraph. Hypernetworks have various types; this invention mainly describes hypernetworks based on hypergraphs. A hypernetwork based on a hypergraph is defined as follows: Given a finite hypergraph H = (V, E) and G is a mapping from [0, +∞) to H, then G(t) = (V(t), E(t)) is called a finite hypergraph, where t ≥ 0.
[0061] This invention uses the domain entity V of each knowledge context. d As nodes, the relationships between domain entities are determined based on the domain entity concept tree. d Establish a domain subnet. Domain subnet S D It is the domain entity V d and its related relationships E d The network formed is represented as
[0062] S D =(V d E d )
[0063] For the relationship E between domain entities dThe present application selects the domain ontology concept similarity as a measurement index of the association relationship. The domain ontology is a standardized concept representation, which can describe the structural relationship between knowledge and formalize the knowledge. Taking the railway tunnel construction in the engineering field as an example, the present application selects the Chinese national standard "GB / T 16566-2018 Terms for railway tunnel". The standard defines the basic vocabulary of railway tunnel and its definition, and is applicable to the planning, survey, design, construction, operation, scientific research and teaching of railway tunnel engineering. According to the hierarchical relationship and the superior-inferior relationship of the terms in the standard, an ontology library is constructed. The present application combines the characteristics of the railway tunnel domain ontology concept tree, and proposes a calculation method combining depth and path, and the association degree calculation formula of the domain node and is:
[0064]
[0065] Wherein, depth(d lcs ) represents the depth of the common parent node of and in the concept hierarchy tree, depth(d) represents the depth of V d in the concept hierarchy tree; shortlength(d i ,d j ) represents the number of edges of the shortest path between and .
[0066] The node in the technology innovation subnetwork mainly refers to the core knowledge content generated by each technology innovation activity, which contains multiple dimensions of knowledge context. The present application constructs a technology innovation node V r for each piece of knowledge context, wherein the achievement, object, process, location and time are the attributes of the node V r , which is represented as:
[0067] V r =<Result,Subject,Process,Location,Time>
[0068] At the same time, the semantic association relationship E r of the technology innovation knowledge output under different contexts is taken as a directed edge, and the technology innovation subnetwork is established. The technology innovation subnetwork S T is a directed network composed of technology innovation nodes V r and the association relationship E r of the knowledge specific content, which is represented as:
[0069] S T =(V r ,Er )
[0070] For semantic association relationship E r , only the semantic similarity of the knowledge content in the two pieces of knowledge context information needs to be calculated, that is, the semantic similarity of the achievements, objects, processes, and place dimensions containing the specific technical innovation content needs to be calculated.
[0071] For the similarity of the object and place dimensions, the innovation object and the application place are both concept sets composed of a small number of words, and the method based on normalized mutual information (NMI) is used to calculate the semantic similarity of the word sets, and the calculation formula is as follows:
[0072]
[0073]
[0074] wherein, denotes the similarity of the innovation object in the technical innovation nodes and , denotes the similarity of the application place in the technical innovation nodes and , denotes the word concept set of the innovation object in the technical innovation node , denotes the word concept set of the application place in the technical innovation node .
[0075] The calculation method of Sim_NMI is as follows:
[0076]
[0077] wherein, NMI(i,j) denotes the similarity of the words i and j, p(i,j) denotes, p(i) denotes, p(j) denotes, and D denotes. The formula is used to calculate the similarity of two words, and when the two words are different in meaning, the similarity is calculated by using the NMI method, otherwise the similarity is 1.
[0078]
[0079] wherein, Sim_NMI(word i ,word j ) denotes the similarity of the word set word i and the word set word j , and num i denotes the number of words in the word set word i .
[0080] For the similarity of the process, result dimension, the technical solutions of the related results and the process are stored in the process dimension, which are usually non-structured text, long text and complex semantics. The Sentence-BERT method is selected to calculate the similarity of the research process, and the calculation formula is:
[0081]
[0082] Wherein, orocess i represents.
[0083] Firstly, the multilingual pre-training model distiluse-base-multilingual-cased-v1 for text similarity calculation is adopted as the model of the experiment; then the orocess i and orocess j two texts are transmitted into the model; two fixed-dimensional embedding vectors are generated through the double-layer Bert model, and the features are extracted through the Pooling layer; finally, the cosine function is used to calculate the text similarity.
[0084] For the result dimension, the text similarity of the result name (title) content is mainly used to calculate the similarity of the dimension. Similar to the process dimension, the Sentence-BERT method is selected to calculate the similarity of the result dimension, and the calculation formula is:
[0085]
[0086] Wherein, result i represents.
[0087] The similarity of the above four dimensions is considered comprehensively, and the semantic correlation relationship E r is weighted to obtain the semantic correlation calculation formula of the technical innovation node and .
[0088]
[0089] The research and development institution is an enterprise entity that undertakes engineering and technical innovation activities, and gathers experts, results and other factors. The research and development institution and the construction unit realize the technical transfer between each other through cooperative innovation. In the present application, each knowledge context contains one or more research and development institutions participating in technical innovation. Taking each research and development institution V i as a node, the cooperation relationship E i between the research and development institutions as a non-directional edge, a research and development institution subnetwork (R&D Partner Subnetwork) is established.I is a directed network composed of R&D institution nodes V i and their associated relationships E i , denoted as:
[0090] S I =(V i , E i )
[0091] where the cooperation between R&D institutions is denoted as:
[0092]
[0093] Experts are the core participants in technological innovation activities and the finishers of innovation achievements. Similar to the R&D institution subnetwork, the expert subnetwork is established with each expert V e as a node and the cooperation between experts E e as an undirected edge. The expert subnetwork S E is an undirected network composed of scientific research experts V e and their cooperation E e , denoted as:
[0094] S E =(V e , E e )
[0095] The cooperation between experts ESim is denoted as:
[0096]
[0097] Based on the four-layer subnetworks of field, technological innovation, R&D institution, and expert, the present application constructs a technological innovation knowledge context super-network model KCHE, denoted as:
[0098] KCHE=(V, E)=(S D , S T , S E , S I , HE)
[0099] V=(V d , V r , V e , V i )
[0100] E=(E d , E r , E e , E i , HE)
[0101] Here, KCHE represents the hypernetwork model of technological innovation knowledge contexts, V is the set of nodes, E is the set of edges, and HE is the hyperedge connecting the four subnetworks. Each hyperedge HE completely represents a knowledge context, such as... Figure 4 As shown.
[0102] In this embodiment, the network feature indicators of the technology innovation knowledge context hypernetwork model include node influence weight, weighted node superdegree, and multi-node superdegree.
[0103] Node influence refers to a node's contribution to a hyperedge. For texts such as patents, papers, and funded projects, research institutions and experts are strictly ranked according to their contribution to the scientific and technological achievements. This means that different research institutions or experts have varying influences within a shared knowledge context, and the same research institution or expert plays different roles in different knowledge contexts. Therefore, for this type of text, this invention assigns a value to node V. e and node V i Different weights are set for different hyperedges. Taking a research and development institution as an example... HE at the edge f The influence weight in the data is:
[0104]
[0105] in, Indicates the subnet node of the research and development institution HE for superedge f Influence weight, Indicates the superedge HE f Total number of subnet nodes of Chinese R&D institutions Represents the subnet nodes of the research and development institution In the super-edge HE f In the sorting, m represents the counting parameter;
[0106] For texts that do not explicitly differentiate the contribution levels of participating R&D institutions and experts, such as engineering case studies, it is assumed that each R&D institution and expert has the same super-edge influence weight. In this case, the R&D institution... The weight of the super-edge influence is:
[0107]
[0108] in, Represents the subnet nodes of the research and development institution HE for super-edge f Influence weight, Indicates the superedge HE f Total number of subnet nodes of R&D institutions in China;
[0109] For node V of the domain subnet dand the nodes of the technology innovation subnetwork V r , without considering its influence on different hyperedges, so its hyperedge influence weight is:
[0110]
[0111] where HEW f (V d ) represents the influence weight of the field subnetwork node V d on the hyperedge HE f .
[0112]
[0113] where HEW f (V r ) represents the influence weight of the field subnetwork node V r on the hyperedge HE f .
[0114] The hyperdegree of a node is generally defined as the number of hyperedges containing the node. After the introduction of node influence weight in the present application, the weighted node hyperdegree of node V is obtained, and the calculation method is:
[0115]
[0116] where wd(V) represents the weighted node hyperdegree of node V in the technology innovation knowledge situation hypernetwork model, HE represents a hyperedge in the technology innovation knowledge situation hypernetwork model, and HEW f (V) represents the influence weight of node V on the hyperedge HE f in the technology innovation knowledge situation hypernetwork model.
[0117] The multi-node hyperdegree represents the hyperedge association degree of multiple nodes of different subnetworks, and the calculation method is:
[0118]
[0119] where, represents the hyperedge hyperdegree containing the research institution subnetwork node the field subnetwork node and the technology innovation subnetwork node , represents the influence weight of the research institution subnetwork node on the hyperedge HE f , represents the influence weight of the field subnetwork node on the hyperedge HE f , represents the influence weight of the field subnetwork node on the hyperedge HE fthe influence weight of the knowledge situation, and HE represents the hyperedge in the hypernetwork model of the knowledge situation of technical innovation. The greater the value is, the greater the correlation between the R&D institutions and the research field is, and the higher the possibility of solving the research problems in the related field is.
[0120] S5, generating the R&D partner recommendation result by using the improved hypernetwork Bayesian inference method according to the node characteristics and the hyperedge characteristics.
[0121] In an optional embodiment of the present application, the technical innovation activity is complex, and it is one-sided to find the experts or R&D partners only by using the keywords, and the most suitable R&D partners cannot be found. Therefore, the present application proposes a new R&D partner recommendation method based on the hypernetwork of the knowledge situation of technical innovation by using the hierarchical structure and the hyperedge to represent the more complex relationship, and using the feature indicators in the hypernetwork and the improved Bayesian algorithm.
[0122] Step S5 specifically includes:
[0123] obtaining the current problem situation data, and extracting the field entity, the object entity and the location entity from the current problem situation data;
[0124] matching the extracted field entity with the field subnetwork nodes in the hypernetwork model of the knowledge situation of technical innovation, and storing the associated nodes with the correlation greater than zero in the set
[0125] adding the technical innovation information of the current problem situation as a new node to the technical innovation subnetwork, storing all the nodes in the hypernetwork model of the knowledge situation of technical innovation with the correlation greater than the set correlation threshold alpha with the node in the set in the set traversing all the nodes in the set in the set
[0126] generating the initial recommendation result of the R&D institutions and the experts by using the improved hypernetwork Bayesian inference method according to the field entity and the technical innovation entity of the current problem situation;
[0127] storing the initial recommendation result of the R&D institutions in the candidate set CAND i , and storing the initial recommendation result of the experts in the candidate set CANDe The candidate set CAND e The research and development institutions to which the experts belong are stored in the candidate set CAND e-i The candidate set CAND i The union set CAND e-i of the candidate set CAND i_final The research and development institutions in the candidate set CAND i_final are comprehensively evaluated, and the comprehensive evaluation results are normalized and sorted to obtain the final research and development partner recommendation results.
[0128] Specifically, to recommend suitable research and development institutions to solve an engineering problem or complete a scientific research project, the problem situation in the current engineering construction needs to be matched with the existing knowledge situation, and the research and development institutions most suitable for the research task are found according to the similarity between them. Therefore, based on the Bayesian inference method based on the super network, the influence weight of the node pair super edge is considered, the engineering technology innovation knowledge situation is fused, and a research and development partner recommendation method based on the knowledge situation super network is proposed. The specific recommendation process is as follows:
[0129] 1. Engineering problem entity extraction
[0130] The field, object and location are extracted from the engineering problem by using the related method of knowledge situation extraction.
[0131] 2. Field subnetwork matching
[0132] In order to match the field involved in the engineering problem with the field concept in the field subnetwork of the super network, the extracted research field is matched with the node with the same semantics in the field subnetwork by searching the nodes in the field subnetwork, and the associated nodes of the node are stored in the set for subsequent calculation.
[0133] 3. Add a new technology innovation subnetwork
[0134] In order to compare the similarity between the engineering problem and the technology innovation nodes in the super network, the technology innovation information of the engineering problem is taken as a new node and added to the technology innovation subnetwork. The node of the newly added technology innovation subnetwork takes the object and location extracted in the front as the values of the Subject and Location attributes in the node , and the values of the Result, Process and Time attributes are empty. Set V r The correlation threshold of the node is α, and all nodes in the super network with a correlation greater than α with the node are stored in the set Since newly added nodes only have two attributes, Subject and Location, some related nodes may be missed when performing similarity matching with nodes in the hypernetwork. Therefore, further traversal is necessary. All nodes in the hypernetwork will be connected to the set Nodes with a correlation degree greater than α are stored in a set.
[0135] 4. Preliminary Recommendation Based on Bayesian Algorithm
[0136] After extracting the engineering problem, an improved Bayesian inference method was used to make initial recommendations for experts and institutions, resulting in TOPK research institutions and TOPM experts. When the engineering problem relates to the research field... and technology innovation entities At that time, recommend research and development institutions The probability can be expressed as The calculation formula is:
[0137]
[0138] KCHE represents the technology innovation knowledge context hypernetwork model. This indicates that the node contains a research institution. Hyperedge and Domain Subnet Nodes Technology Innovation Subnet Node The probability of association, This indicates that the node contains a research institution. The probability of the occurrence of a superedge. Represents a domain subnet node With domain subnet nodes The degree of correlation, Represents the technology innovation subnet node With technology innovation subnet nodes The degree of correlation, This indicates that it includes subnet nodes of research and development institutions. Domain subnet nodes Technology Innovation Subnet Node The transcendence of the edge, Indicates the subnet node of the research and development institution Weighted node over-scaling, Indicates the subnet node of the research and development institution Weighted node over-scaling, Represents a domain subnet node With technology innovation subnet nodes The transcendence of the edge, Represents a domain subnet node Weighted node over-scaling, Represents the technology innovation subnet node weighted node hyperdegree of the expert subnetwork node
[0139] Similarly, when the engineering problem involves research fields and technology innovation entities , the probability of recommending the scientific research expert can be represented as The calculation formula is:
[0140]
[0141] wherein, represents the association probability of the hyperedge containing the expert subnetwork node and the field subnetwork node and the technology innovation subnetwork node , represents the occurrence probability of the hyperedge containing the expert subnetwork node , represents the hyperdegree of the hyperedge containing the expert subnetwork node the field subnetwork node and the technology innovation subnetwork node , represents the weighted node hyperdegree of the expert subnetwork node , and represents the weighted node hyperdegree of the expert subnetwork node . 5. Comprehensive evaluation and recommendation By recommending the experts, the recommendation results of the research and development institutions can be corrected and improved. For example, when solving some problems, a certain expert plays a great role in different knowledge situations, but the research and development institution to which the expert belongs ranks at the back in these knowledge situations. The recommendation effect is poor if only the research and development institutions are recommended, which cannot identify this information. Therefore, the preliminary recommendation results of the institutions are stored in the candidate set CAND i , the preliminary recommendation results of the experts are stored in the candidate set CAND e , and the research and development institutions to which the experts in the expert candidate set belong are stored in the candidate set CAND e-i . Then, the union set CAND i of CAND e-i and CAND i_final is taken, and the research and development institutions in the candidate set CAND i_final are comprehensively evaluated, specifically as follows:
[0142]
[0143]
[0144]
[0145] wherein, represents the research and development institution subnetwork node in the candidate set CAND i_final . The probability of recommendation Representing the R&D institution subnet nodes in the hypernetwork model of technological innovation knowledge context The probability of recommendation Expert subnet nodes in the hypernetwork model representing knowledge context of technological innovation The probability of recommendation This refers to the set of recommended experts from research and development institutions.
[0146] The results of the comprehensive evaluation of research institutions are normalized and ranked to obtain the final recommendation, denoted as:
[0147]
[0148] Among them, P final (V) max Candidate set of research institutions (CAND) i_final The highest recommendation probability, P final (V) min Candidate set of research institutions (CAND) i_final The lowest recommendation probability.
[0149] In summary, technological innovation knowledge is a crucial resource for engineering construction projects, and its effective utilization can solve many engineering problems and meet research needs. However, this knowledge is scattered across various information entities from multiple sources, involving multiple stakeholders and organizations, making it difficult for enterprises to extract useful knowledge from massive amounts of data to guide the search for the most suitable R&D partners for specific technological innovations. This invention constructs an engineering technology innovation knowledge context model and proposes a hypernetwork recommendation framework based on engineering technology innovation knowledge context, providing methods and tools to address this problem. On one hand, this invention designs an automatic extraction method for engineering technology innovation knowledge context, which can automatically extract the knowledge context of engineering technology innovation activities from massive amounts of unstructured documents. This method can effectively improve the efficiency of knowledge utilization, solve problems such as knowledge dispersion and knowledge overload, and facilitate the sharing and reuse of engineering technology innovation knowledge. On the other hand, this invention maps knowledge context information into a hypernetwork model, constructing a four-layer hypernetwork of domain-technology innovation-expert-R&D institution, and designs an improved hypernetwork Bayesian algorithm to recommend R&D institutions. This invention designs a node similarity calculation method that combines domain ontology, standardized mutual information, and deep learning algorithms, and proposes a node influence weight index, which effectively explores the deep relationships between knowledge contexts in hypernetworks.
[0150] The present application is described in reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the flow Figure 1 one or more flows and / or blocks. Figure 1 one or more blocks.
[0151] These computer program instructions can also be stored in a computer readable memory that can direct the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction devices that implement the flow Figure 1 one or more flows and / or blocks. Figure 1 one or more blocks.
[0152] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the flow Figure 1 one or more flows and / or blocks. Figure 1 one or more blocks.
[0153] The principles and implementation modes of the present application are described in the specific embodiments in the present application, and the above embodiment description is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed, and the above description should not be understood as the limitation of the present application.
[0154] Those skilled in the art will realize that the embodiments described herein are for the purpose of helping the reader to understand the principles of the present application, and should be understood as the protection scope of the present application is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the protection scope of the present application.
Claims
1. A method for recommending R&D partners based on a technology innovation knowledge situation supernetwork, characterized in that, The method comprises the following steps: S1, acquiring technical innovation project data; S2, constructing a technical innovation knowledge context model and determining technical innovation knowledge context dimensions; S3, extracting multi-dimensional knowledge context information from the technical innovation project data based on the technical innovation knowledge context dimensions; S4, constructing a technical innovation knowledge context super network model based on the multi-dimensional knowledge context information, and determining node features and super edge features, specifically including: constructing a technical innovation knowledge context super network model including a field subnetwork, a technical innovation subnetwork, an R&D institution subnetwork, and an expert subnetwork with a four-layer structure; wherein the field subnetwork takes each field entity of a knowledge context as a node, and determines the association relationship between the field entities according to a field entity concept tree; the technical innovation subnetwork constructs a technical innovation node for each knowledge context, wherein the attributes of the technical innovation node are achievements, objects, processes, locations, and times, and the semantic association relationship of the technical innovation knowledge output under different contexts is taken as a non-directional edge; the R&D institution subnetwork takes each R&D institution as a node, and takes the cooperation relationship between the R&D institutions as a non-directional edge; the expert subnetwork takes each expert as a node, and takes the cooperation relationship between the experts as a non-directional edge; respectively calculating the node influence weight in the technical innovation knowledge context super network model, and respectively calculating the weighted node super degree and the multi-node super degree according to the node influence weight, wherein the calculation method of the node influence weight in the technical innovation knowledge context super network model is: for the text that clearly distinguishes the contribution degree of the participating R&D institutions and experts, calculating the influence weight of the R&D institution subnetwork node in the super edge as: in, Indicates the subnet node of the research and development institution For superedge Influence weight, Indicates the superedge Total number of subnet nodes of Chinese R&D institutions Indicates the subnet node of the research and development institution In the super-edge The sorting in m Indicates the counting parameter; for the text that does not clearly distinguish the contribution degree of the participating R&D institutions and experts, calculating the influence weight of the R&D institution subnetwork node in the super edge as: wherein, represents the R&D institute subnetwork nodes the influence weight of the super edge , wherein, represents the total number of R&D institute subnetwork nodes in the super edge . for the field subnetwork node, calculating the influence weight thereof in the super edge as: wherein, representing domain subnet nodes influence weight of a super-edge ; for the technical innovation subnetwork node, calculating the influence weight thereof in the super edge as: wherein, representing a domain subnet node influence weight of a super-edge ; S5, generating an R&D partner recommendation result by using an improved super network Bayesian inference method according to the node features and the super edge features. 2.The R&D partner recommendation method based on a technical innovation knowledge situation super network according to claim 1, characterized in that, The technical innovation knowledge context model constructed in step S2 specifically includes: a field dimension, an object dimension, a process dimension, an achievement dimension, an expert dimension, an institution dimension, a location dimension, and a time dimension. 3.The R&D partner recommendation method based on a technical innovation knowledge situation super network according to claim 1, characterized in that, Step S3 specifically includes: dividing the technical innovation project data into structured information, unstructured text, and semi-structured text; directly mapping the structured information to the achievement dimension, the expert dimension, the institution dimension, and the time dimension of the technical innovation knowledge context model; preprocessing the unstructured text by performing word segmentation and part-of-speech tagging; then performing corpus annotation by using BIOES encoding; then performing word vector embedding on the preprocessed text by using a Word2Vec model to obtain a vector expression of the words; then automatically extracting three types of entities of fields, objects, and locations by using a BiLSTM-CRF model; and finally mapping the extracted three types of entities of fields, objects, and locations to the field dimension, the object dimension, and the location dimension of the technical innovation knowledge context model; performing sentence segmentation and word segmentation on the semi-structured text, then performing data cleaning, and finally mapping the processed text to the process dimension of the technical innovation knowledge context model.
4. The method of claim 1, wherein the method is characterized by, The calculation manner of the weighted node super-degree is: wherein, denotes the weighted node hyperdegree of a node in the hypernetwork model of the knowledge context of technological innovation, denotes a hyperedge in the hypernetwork model of the knowledge context of technological innovation, denotes the influence weight of a node on a hyperedge in the hypernetwork model of the knowledge context of technological innovation.
5. The method of claim 1, wherein the method is characterized by, The calculation manner of the multi-node super-degree is: wherein, represents a super-edge containing R&D institution sub-network nodes , field sub-network nodes and technology innovation sub-network nodes , represents the influence weight of R&D institution sub-network nodes on super-edges , represents the influence weight of field sub-network nodes on super-edges , represents the influence weight of field sub-network nodes on super-edges , represents a super-edge in the technology innovation knowledge situation super-network model.
6. The method of claim 1, wherein the method is characterized by, The step S5 specifically comprises: Obtaining current problem context data, and extracting domain entities, object entities and place entities from the current problem context data; The extracted domain entity is matched with the domain subnetwork node in the technology innovation knowledge situation super network model, and the associated nodes with the correlation degree greater than zero are stored in the set into the set ; The technical innovation information of the current problem context is taken as a new node All nodes in the technical innovation sub-network whose association degree with the node in the technical innovation knowledge context super-network model is greater than a set association degree threshold value are stored in a set ; all nodes in the set are traversed again, and all nodes in the technical innovation knowledge context super-network model whose association degree with the nodes in the set is greater than a set association degree threshold value are stored in a set ; An improved hypernetwork Bayesian inference method is used to generate initial recommendation results of research institutions and experts according to domain entities of current problem situations and technical innovation entities Storing the initial recommendation result of the research and development institution into the candidate set Storing the initial recommendation result of the expert into the candidate set Storing the candidate set Storing the research and development institution to which the expert belongs into the candidate set Taking the candidate set The union of the candidate set The union of the candidate set Comprehensively evaluating the research and development institutions in the candidate set , and normalizing the comprehensive evaluation result to obtain the final research partner recommendation result.
7. The method according to claim 6, wherein, Using the improved hypernetwork Bayesian inference method, the domain entity and technical innovation entity in the current problem context is generated, The recommendation probability calculation formula of the research and development institution is: wherein, represents the technical innovation knowledge situation super network model, represents the association probability of the super edge containing the scientific research institution node and the field subnetwork node and the technical innovation subnetwork node , represents the appearance probability of the super edge containing the scientific research institution node , represents the association degree of the field subnetwork node and the field subnetwork node , represents the association degree of the technical innovation subnetwork node and the technical innovation subnetwork node , represents the super degree of the super edge containing the research and development institution subnetwork node , the field subnetwork node and the technical innovation subnetwork node , represents the weighted node super degree of the research and development institution subnetwork node , represents the weighted node super degree of the research and development institution subnetwork node , represents the super degree of the super edge of the field subnetwork node and the technical innovation subnetwork node , represents the weighted node super degree of the field subnetwork node , represents the weighted node super degree of the technical innovation subnetwork node ; The recommendation probability calculation formula of the expert is: wherein, denotes the association probability of a super-edge containing an expert sub-net node a domain sub-net node and a technology innovation sub-net node , denotes the occurrence probability of a super-edge containing an expert sub-net node , denotes the super-degree of a super-edge containing an expert sub-net node a domain sub-net node and a technology innovation sub-net node , denotes the weighted node super-degree of an expert sub-net node , denotes the weighted node super-degree of an expert sub-net node .
8. The R&D partner recommendation method based on the technical innovation knowledge situation super network according to claim 6, characterized in that, For candidate set The comprehensive evaluation of the research and development institutions is specifically as follows: wherein, representing the candidate set the recommended probability of the R&D institute sub-network node , representing the recommended probability of the R&D institute sub-network node in the technology innovation knowledge situation super-network model , representing the recommended probability of the expert sub-network node in the technology innovation knowledge situation super-network model , , representing the set of recommended experts of the R&D institute.
Citation Information
Patent Citations
Method and device for matching item and professionals
CN104700190A
Expert recommendation method and device based on paper data analysis, equipment and storage medium
CN112100470A
Expert recommendation method based on super-network model
CN111737451A
KR20210150103A