A knowledge graph-based research and development project intelligent review decision method

CN122819982APending Publication Date: 2026-09-25SHENZHEN RUNDIAN INFORMATION TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610913744.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]本发明要解决的技术问题是:针对现有研发项目评审方法中知识图谱构建维度单一、评审决策缺乏多准则融合机制、决策过程不可追溯、缺乏反馈学习与动态调整能力的技术缺陷,提供一种基于知识图谱的研发项目智能评审决策方法,实现多维度的项目语义表征、多准则融合的智能决策、可追溯的推理路径以及持续反馈优化的闭环机制

Benefits of technology

1、通过构建涵盖项目、技术、人员、机构、成果五类实体及其多维语义关系的研发项目知识图谱,并引入技术演化关系边,实现了比现有技术更全面的项目语义表征,有效解决了信息利用不充分导致的评审偏差问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819982A_ABST
    Figure CN122819982A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on knowledge graph's research and development project intelligent review decision-making method, belong to intelligent review and decision support technical field.The method includes the following steps: obtaining the multi-modal project data of the research and development project to be reviewed, the multi-modal project data includes project declaration text, technical scheme document, project team history data and historical project data;Entity, relationship and attribute are extracted from the multi-modal project data, and a multidimensional research and development project knowledge graph with the project as the core node is constructed;The pre-trained heterogeneous graph neural network model is used to perform semantic coding on the multidimensional research and development project knowledge graph, to generate node embedding representation and graph-level representation;The present application effectively solves the technical problems of insufficient information utilization, non-uniform evaluation criteria and poor decision-making interpretability in traditional evaluation methods, significantly improving the objectivity, accuracy and traceability of research and development project evaluation and decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent review and decision support technology, specifically to an intelligent review and decision-making method for R&D projects based on knowledge graphs. Background Technology

[0002] Research and development (R&D) project review is a core component of science and technology management, corporate innovation, and the allocation of research funds. The quality of the review directly impacts the efficiency of science and technology resource allocation, the selection of innovation directions, and the return on R&D investment. With the accelerating pace of global technological innovation and the continuous growth in R&D investment, the number of R&D projects requiring review each year has increased dramatically, posing a dual challenge to both efficiency and quality in the review process.

[0003] Traditional R&D project reviews primarily rely on manual review and meeting assessments by domain experts, which have significant technical limitations. First, from an information utilization perspective, manual review struggles to comprehensively correlate and compare the technical characteristics of the project under review with historical projects, easily leading to misjudgments of technological novelty—a "missed recall" problem—the failure to retrieve and utilize relevant historical project information. Simultaneously, review experts, limited by their individual knowledge, are prone to "recall errors"—the retrieved knowledge being irrelevant to the actual review needs, or "recall failures"—completely omitting crucial comparative information. Second, from a review consistency perspective, differences in review standards and weighting preferences among experts result in a lack of objective and consistent evaluation benchmarks. Third, from a decision interpretability perspective, traditional review processes typically only provide a final score or grade, lacking detailed traceability of the review decision-making basis, making effective verification and supervision of review results difficult.

[0004] In recent years, some researchers have attempted to introduce artificial intelligence technology into the field of project review. For example, a publicly available technology, "Project Evaluation and Review Method and System Integrating Natural Language Processing" (Publication No. CN 118780767A), has been published. This method constructs a knowledge graph of the projects to be reviewed, uses a heterogeneous graph neural network to obtain node embedding representations, obtains historical similar projects through a spectral clustering algorithm, and finally generates review opinions. Another publicly available technology, "Project Decision Optimization Control Method, Device, and Equipment Based on Knowledge Graph" (Authorization Announcement No. CN 120509686B), has been published. This method constructs a project knowledge graph and performs consistency verification, performs semantic analysis based on a multi-objective performance function, performs compression and filtering through a graph neural network, and finally generates executable control instructions.

[0005] However, the aforementioned existing technical solutions still have the following technical problems: (1) The knowledge graph construction has a single dimension. The knowledge graphs constructed by existing technologies mainly revolve around the two dimensions of the project itself and the review experts. They do not fully consider the deep semantic relationships between multiple entities such as technical entities, team members, institutional background, and output, as well as the technological evolution relationships across projects (such as technology inheritance and technology differentiation), resulting in an incomplete semantic representation of R&D projects.

[0006] (2) The review and decision-making process lacks a multi-criteria integration mechanism. Most existing technologies rely on single-dimensional similarity comparison or classification models for review and judgment, lacking a systematic integration of multi-dimensional evaluation criteria such as technological novelty, team capability, technological feasibility, and project risk, and failing to resolve potential conflicts and contradictions between evaluation results from different dimensions.

[0007] (3) The decision-making process is not traceable. The review results generated by existing technologies lack an interpretable reasoning path, and review experts and decision-makers cannot trace the source of the scores, making it difficult to conduct effective review and supervision.

[0008] (4) Lack of feedback learning and dynamic adjustment capabilities. The existing technology methods have fixed model parameters after deployment, and cannot learn online and continuously optimize based on the feedback of review experts, thus failing to adapt to the dynamic changes in review standards and focus. Summary of the Invention

[0009] The technical problem this invention aims to solve is to address the shortcomings of existing R&D project review methods, such as single-dimensional knowledge graph construction, lack of multi-criteria fusion mechanism in review decisions, untraceable decision-making process, and lack of feedback learning and dynamic adjustment capabilities. This invention provides a knowledge graph-based intelligent review and decision-making method for R&D projects, which achieves multi-dimensional project semantic representation, intelligent decision-making through multi-criteria fusion, traceable reasoning paths, and a closed-loop mechanism for continuous feedback optimization.

[0010] The technical solution of this invention is: an intelligent review and decision-making method for R&D projects based on knowledge graphs, comprising the following steps: S1. Obtain multimodal project data of the R&D projects to be reviewed. The multimodal project data includes project application text, technical solution documents, project team resume data, and historical project data. S2. Extract entities, relationships, and attributes from the multimodal project data to construct a multidimensional R&D project knowledge graph with projects as the core nodes. The multidimensional R&D project knowledge graph includes project entity nodes, technology entity nodes, personnel entity nodes, organization entity nodes, achievement entity nodes, and semantic relationship edges between each node. S3. Use a pre-trained heterogeneous graph neural network model to perform semantic encoding on the multi-dimensional R&D project knowledge graph, generate node embedding representations of each entity node, and aggregate the node embedding representations through graph pooling to obtain the graph-level representation of the multi-dimensional R&D project knowledge graph. S4. Based on the graph-level representation, calculate the multi-dimensional semantic similarity between the R&D project to be reviewed and the previously reviewed projects, and generate a technology novelty score, a team capability score, a technology feasibility score, and a risk level score based on the multi-dimensional semantic similarity. S5. Input the technology novelty score, team capability score, technology feasibility score, and risk level score into the multi-criteria fusion decision model based on the attention mechanism. Calculate the attention weights of each scoring dimension through adaptive weight learning, and perform weighted fusion of each scoring dimension based on the attention weights to generate a comprehensive review decision result. The comprehensive review decision result includes: a binary decision of review pass / fail, and when the review is passed, a priority ranking of the R&D project to be reviewed among multiple projects to be reviewed based on the comprehensive score.

[0011] Further, step S2 involves constructing a multi-dimensional R&D project knowledge graph, specifically including: S21, using a pre-trained natural language processing model to perform named entity recognition and relation extraction on the project application text and technical solution documents to obtain project entities, technical entities, and their technical relationships; S22, extracting personnel entities and their attribute information from the project team's resume data, including technical expertise, research achievements, project experience, and academic influence indicators, and constructing capability matching relationships between the personnel entities and the technical entities; S23, extracting existing project entities and their review result labels from the historical project data; S24, constructing technical evolution relationship edges between projects; S25, performing consistency verification on the constructed initial knowledge graph, detecting and eliminating entity conflicts and relationship conflicts, and using a neighbor-based entity aggregation-based imputation strategy for missing attribute values.

[0012] Furthermore, the pre-trained heterogeneous graph neural network model in step S3 includes: a node type-aware linear transformation layer, a multi-layer heterogeneous graph attention network, inter-layer residual connections and layer normalization, and an attention-weighted graph-level readout layer. The pre-training process of this model employs a joint training strategy of supervised learning and contrastive learning.

[0013] Furthermore, in step S4, when calculating multi-dimensional semantic similarity, the overall similarity and the similarity of each subgraph dimension are calculated separately, and the K most similar historical items are selected to form a set of similar items. The scores of each dimension are generated based on the weighted average.

[0014] Furthermore, the multi-criteria fusion decision model in step S5 includes a scoring embedding layer, a self-attention encoding layer, a cross-attention layer, and a decision output layer (containing a parallel binary classifier and a ranking regressor). The training process employs joint optimization of the cross-entropy loss function and the pairwise ranking loss function.

[0015] Furthermore, after step S5, the process also includes: S6, generating a traceable review decision explanation report; S7, obtaining feedback and adjustment instructions from review experts; S8, converting the revised opinions into knowledge graph update triples; and S9, fine-tuning the model online and dynamically updating the model parameters.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. By constructing a knowledge graph of R&D projects that covers five types of entities—projects, technologies, personnel, institutions, and results—and their multidimensional semantic relationships, and by introducing technological evolution relationship edges, a more comprehensive semantic representation of projects than existing technologies has been achieved, effectively solving the problem of review bias caused by insufficient information utilization.

[0017] 2. By working together with a heterogeneous graph neural network and a multi-criteria fusion decision model based on an attention mechanism, the adaptive fusion of multi-dimensional review criteria was achieved, which resolved the potential conflicts and contradictions between evaluation results of different dimensions and improved the accuracy and robustness of review decisions.

[0018] 3. By generating traceable review decision explanation reports, including contribution distribution of each dimension, retrospective analysis of similar historical projects, and visualization of semantic association paths, the problem of black box in the decision-making process in existing technologies is solved.

[0019] 4. Through the incremental update mechanism of knowledge graph driven by expert feedback and the online fine-tuning mechanism of the model, the review system has achieved continuous self-optimization and can adapt to the dynamic changes of review standards in different fields and at different times. Attached Figure Description

[0020] Figure 1 A flowchart illustrating an intelligent review and decision-making method for R&D projects based on knowledge graphs, provided as an embodiment of the present invention; Figure 2 This is a flowchart of the pre-training process of the heterogeneous graph neural network model in an embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] Example 1; This embodiment details the specific implementation method of constructing a multi-dimensional R&D project knowledge graph in step S2.

[0023] In this embodiment, the R&D projects to be reviewed originate from the annual application of a national-level scientific research fund project. The multimodal project data includes the following four categories: Project proposal text: A structured document in PDF format, containing sections such as project name, research objectives, research content, technical approach, expected results, and budget. This example uses an OCR-based text extraction tool to convert the PDF file into plain text format and segments it into paragraphs according to the chapter structure.

[0024] Technical solution document: A detailed technical solution in Word format, including technical background, detailed description of the technical solution, key technical indicators, explanation of innovation points, feasibility analysis, etc. This embodiment uses a document parser to extract the main text and table data.

[0025] Project team resume data: Structured data in JSON format, including basic information of each team member (name, title, affiliation), a list of representative papers (including title, journal, citation count), a list of patent applications, a list of projects led or participated in, award records, etc.

[0026] Historical project data: Historical project records from the project review database, including basic information, application summary, technical keywords, review scores (out of 100, with scores for each dimension), review conclusions (pass / fail), and review comments text for 1,247 reviewed projects in the past 5 years.

[0027] The above data was preprocessed, including: removing special characters and formatting marks from the text; unifying the encoding format to UTF-8; and standardizing the team resume data by converting the number of paper citations into log-normalized values ​​and the amount of project funding into normalized values.

[0028] This embodiment uses a pre-trained BERT-BiLSTM-CRF sequence labeling model for named entity recognition. This model was fine-tuned on a corpus of 50,000 labeled research project data and is able to identify entity types in 5 major categories and 23 subcategories.

[0029] The specific identification categories are shown in the table below:

[0030] Relation extraction employs an entity pair classification method, utilizing a pre-trained language model to encode the context of two entities, and then classifying the relation type through a fully connected layer. This embodiment defines the following eight relation types: (Project-Technology): The project involves a certain technology. (Project-Technology) Proposal: The project proposes new technological methods. (Personnel-Role): A person assumes a certain role in the project. Expertise (Personnel-Technology): A person is skilled in a certain technology. Belongs to (personnel-organization): The person belongs to a certain organization. Output (Project-Result): The project is expected to produce a specific result. Similar to a previous project (project-to-project): The current project is similar to a previous project. Differentiation from previous projects (project-to-project): The technical differentiation of the current project based on previous projects. The extraction rules for the relationship "propose" are as follows: if the technical entity does not have a matching record in the existing technology database of the knowledge graph, and the application text contains related modifiers such as "innovation", "first proposal", and "originality", it is determined to be a "propose" relationship; if the technical entity has a matching record in the existing technology database, it is determined to be a "belong to" relationship.

[0031] This embodiment specifically introduces technology evolution relationship edges to characterize the technological development trajectory across projects. Technology evolution relationships include two types: "technology inheritance" and "technology differentiation".

[0032] For constructing the technology inheritance relationship: Calculate the set of technical keywords T = {t1, t2, ..., t} for the current project under review. m The set of technical keywords for historical projects is T'={t'1, t'2, ..., t'}. n The Jaccard similarity between the items is calculated. When the Jaccard similarity meets a preset threshold θ1 (e.g., θ1=0.4), and the current item's time is later than the historical item's time, a "technology inheritance" relationship edge is established.

[0033] For the construction of technology differentiation relationships: when two historical projects A and B both have an inheritance relationship with the current project C in terms of technology keyword set, but the Jaccard similarity between A and B is lower than the threshold θ2 (e.g., θ2=0.3), C is determined to be the "technology differentiation" node of A and B, and "technology differentiation" relationship edges are established between C and A, and between C and B respectively.

[0034] Perform consistency checks on the constructed initial knowledge graph, including: Entity conflict detection: Detects whether entities with the same or highly similar names (edit distance less than 2) point to different entity types or have contradictory attribute values. If a conflict is detected, the source with the higher confidence level is used (in priority: team resume data > technical solution documents > project proposal text).

[0035] Relationship conflict detection: Detects whether there are contradictory semantic relationships. For example, a person entity cannot have a "belong to" relationship with two organization entities at the same time (unless the person has multiple jobs, in which case two relationships can be established and marked as "primary unit" and "part-time unit").

[0036] Attribute missing imputation: For entities lacking attributes, an imputation strategy based on neighbor entity aggregation is adopted. Specifically, suppose entity e is missing the value of attribute a. Collect all k similar neighbor entities N(e) of e. Use the weighted average of the values ​​of the neighbor entities on attribute a as the estimated value of attribute a of e, with the weight being the semantic similarity between e and the neighbor entities.

[0037] After the above steps, the final multi-dimensional R&D project knowledge graph contains entities and relationships of the following scale:

[0038] Example 2; This embodiment details the specific implementation of steps S3 and S4.

[0039] The heterogeneous graph neural network model used in this embodiment is based on the Heterogeneous Graph Transformer (HGT) architecture and has been improved.

[0040] Layer 1: Node Type-Aware Linear Transformation Layer For different types of nodes in a knowledge graph, the initial feature vectors have different dimensions and semantic spaces. For example, the initial features of a technology entity node are obtained by averaging the word vectors of technical keywords (300 dimensions), the initial features of a personnel entity node are composed of statistical vectors of resume information (64 dimensions), and the initial features of a project entity node are composed of BERT representations of the application text (768 dimensions).

[0041] This layer sets an independent linear transformation matrix for each node type τ. , the initial feature vector Mapped to a unified Hidden space (in this embodiment) =256): in Represents a node The node type.

[0042] Layers 2 through 4: Heterogeneous graph attention network layers This embodiment employs a 3-layer heterogeneous graph attention network for information aggregation. In the l-th layer (l=1,2,3), the embedding representation of node v is updated using the following formula: First, for nodes For each type of relation edge, compute the node. For neighboring nodes Attention weights: Where Q and K are the query matrix and key matrix, respectively, and different parameter matrices are used for different types of source nodes and different types of relation edges.

[0043] Then, the new embedding representation of node v is obtained by weighted summation of the value vectors of all neighboring nodes and followed by nonlinear activation: Where V is the value matrix and σ is the GELU nonlinear activation function.

[0044] Layer 5: Graph-level readout layer After completing the three-layer message passing, the final set of embedded representations of all nodes in the knowledge graph is obtained. The graph-level readout layer uses attention-weighted global pooling to generate the graph-level representation g.

[0045] Where u is a learnable global query vector. The importance weight of node v.

[0046] (II) Pre-training of Heterogeneous Graph Neural Network Models In this embodiment, the pre-training of the heterogeneous graph neural network model adopts a strategy of joint training of supervised learning and contrastive learning.

[0047] Supervised learning task: For each historical item i, construct its corresponding knowledge graph. Graph-level representation is obtained through a heterogeneous graph neural network model. The fully connected layer maps the result to the predicted comprehensive review score. Based on historical actual review scores For the monitoring signal, the mean square error loss function is used:

[0048] Where N is the number of training samples.

[0049] Contrastive Learning Auxiliary Task: Building upon supervised learning, a contrastive learning auxiliary task is introduced to enhance the model's ability to distinguish similar items. For items within each batch, they are divided into positive and negative sample sets based on their actual review labels (pass / fail). The contrastive loss is defined as:

[0050] Where P represents the set of positive sample pairs (item pairs with the same label), N represents the set of negative sample pairs (item pairs with different labels), s(·,·) is the cosine similarity function, and τ is the temperature parameter (τ=0.07 in this embodiment).

[0051] Total training loss:

[0052] Where λ is the equilibrium hyperparameter (λ=0.3 in this embodiment).

[0053] Training configuration: The AdamW optimizer was used with an initial learning rate of 0.001, a cosine annealing learning rate decay strategy, a batch size of 32, 200 training epochs, and an early stopping strategy (patience=20) on the validation set.

[0054] (III) Multi-dimensional semantic similarity calculation and score generation This embodiment details the process of multi-dimensional semantic similarity calculation and score generation.

[0055] Overall similarity calculation: Projects pending review Graph-level representation With each historical item P in the knowledge graph i Graph-level representation of (i=1,...,M, M=1,247) Perform cosine similarity calculation:

[0056] Subgraph dimensional similarity calculation: In addition to overall similarity, this embodiment also extracts three key subgraphs from the knowledge graph: The technology subgraph G_tech contains project nodes and their associated technology entity nodes, technology indicator nodes, and technology evolution relationship edges. The technology subgraph is extracted from the knowledge graph according to predefined subgraph extraction rules and then input into a heterogeneous graph neural network model to obtain the technology subgraph representation g_tech.

[0057] The team subgraph G_team contains project nodes and their associated personnel entity nodes, organization entity nodes, and capability matching edges. The team subgraph is extracted from the knowledge graph according to predefined subgraph extraction rules and then input into a heterogeneous graph neural network model to obtain the team subgraph representation g_team.

[0058] The achievement subgraph G_achievement contains project nodes, their associated achievement entity nodes, and their attributes. The achievement subgraph is extracted from the knowledge graph according to predefined subgraph extraction rules and then input into a heterogeneous graph neural network model to obtain the achievement subgraph representation g_achievement.

[0059] Calculate the similarity scores for the technology dimension (sim_tech), team dimension (sim_team), and achievement dimension (sim_achievement) respectively: The similarity calculation method for team dimension and outcome dimension is the same, and the corresponding subgraph representation is used for calculation.

[0060] Building a collection of similar projects: Select the K historical projects that have the highest similarity to the project to be reviewed to form a similar project set S_K (K=50 in this example):

[0061] Multi-dimensional rating generation: Based on the review result labels of each project in the similar project set S_K, a weighted average method is used to calculate the score of the project to be reviewed in each dimension. Weight The normalized value of the overall similarity between the project to be reviewed and its corresponding historical projects:

[0062] The novelty score S_novelty is calculated as a weighted average of the historical novelty scores of the top 10 projects with the highest technical similarity among similar projects. If the technical subgraph similarity between a similar project and the project under review exceeds 0.85, the novelty score of the project under review will be further reduced (high technical overlap).

[0063] Team capability score S_team: Based on the weighted average of the historical team capability scores of the top 10 projects with the highest team dimension similarity among similar projects, and also directly assesses the comprehensive resume indicators of the team members of the project to be reviewed.

[0064] Technical feasibility score S_feasibility: Based on the weighted average of the historical scores of all K similar projects in the technical feasibility dimension, and adjusted by introducing technical route complexity adjustment factors and technical maturity adjustment factors.

[0065] Risk level score S_risk: The comprehensive risk level score is calculated by comprehensively assessing three sub-dimensions: technical risk, team risk, and schedule risk, and using a weighted summation method.

[0066] The example scores generated in this embodiment are as follows (score range: 0-100):

[0067] Example 3; This embodiment details the specific implementation of steps S5 to S9.

[0068] The scoring embedding layer maps the technological novelty score S_novelty, team capability score S_team, technological feasibility score S_feasibility, and risk level score S_risk to 128-dimensional scoring embedding vectors e_novelty, e_team, e_feasibility, and e_risk, respectively, using four independent linear projection matrices. To preserve the order of the scores, sinusoidal positional encoding is added to the embedding vectors, with position indices fixedly assigned according to the scoring dimension (novelty → 0, team → 1, feasibility → 2, risk → 3).

[0070] Self-attention encoding layer: concatenates the four rating embedding vectors into a sequence. The interaction relationships between the rating dimensions are calculated by a two-layer multi-head self-attention encoder (head=4), generating a context-aware rating representation sequence E'=[e'_novelty; e'_team; e'_feasibility; e'_risk].

[0071] Cross-Attention Layer: This layer is the key innovation of this model. It uses the context-aware rating representation sequence E' as the query and the graph-level representation g output by the heterogeneous graph neural network model as the key and value. Through a cross-attention mechanism, the weights of the rating dimension can learn to perceive the deep semantic information contained in the knowledge graph.

[0072] Before being used as inputs to K and V, g is first mapped to a 128-dimensional representation vector through a linear projection matrix.

[0073] Decision output layer: Binary classifier: The fused score representation is obtained by mean pooling to get a 128-dimensional fused vector h_fuse, which is then passed through a fully connected layer and a sigmoid activation function to output the probability p_pass of passing the review.

[0074]

[0075] When p_pass ≥ 0.5, it is judged as "passed"; when p_pass < 0.5, it is judged as "failed".

[0076] Ranking Regressor: The fused vector h_fuse is fed into a ranking score prediction network in parallel, and the network outputs a priority score s_priority (0-100) for each project. All projects that pass the review are ranked from highest to lowest according to their s_priority scores, which is used to determine funding priorities in budget-constrained scenarios.

[0077] Reasoning example: The input rating vector in this embodiment is (S_novelty=78.6, S_team=82.3, S_feasibility=71.5, S_risk=35.2), and the following results are obtained through model inference:

[0078] The attention weights show that, in this project, the team capability dimension (weight 0.312) and the technology novelty dimension (weight 0.283) have the greatest impact on the decision-making results, indicating that the model believes these two are the key factors to focus on when reviewing the current project.

[0079] In this embodiment, the multi-criteria fusion decision model is trained using a joint optimization strategy of cross-entropy loss function and pairwise ranking loss function. The training dataset contains multi-dimensional scoring data for 1,247 historical items and corresponding actual review conclusions.

[0080] Binary classifier training: A weighted cross-entropy loss function is used, with different weights assigned to positive samples (passed items) and negative samples (failed items) to handle class imbalance (pass rate was approximately 45% in historical data).

[0081] Where y i The labels are real (1 = pass, 0 = fail). This represents the predicted probability of passage from the model. =1 / (2×pass rate)=1.11, =1 / (2×failure rate)=0.91.

[0082] Ranking regressor training: A pairwise ranking loss function is used. Two items, a and b, are randomly sampled from the passed items set. If the actual overall score of a is higher than that of b, a positive sample pair (a, b) is constructed, with the expected predicted priority score of a higher than that of b. One item, c, and one item, d, are sampled from the passed items and the failed items respectively, and a constraint pair (c, d) is constructed, with the expected predicted priority score of c significantly higher than that of d. The pairwise ranking loss is:

[0083] Where margin is a preset safety interval (margin=0.3 in this embodiment), s i s j These are the prediction priority scores for items i and j, respectively.

[0084] Total training loss: Total loss = classification loss + ranking loss. The Adam optimizer is used with a learning rate of 0.0005, a batch size of 64, 100 training epochs, and an early stopping strategy (patience=15).

[0085] In this embodiment, after generating the comprehensive review and decision results, the system automatically generates a structured review and decision explanation report.

[0086] Dimensional Contribution Analysis: Based on the attention weights calculated during the reasoning process of the multi-criteria fusion decision-making model, the contribution distribution of each dimension to the decision result is generated. The contribution distribution includes the contribution of technological novelty dimension (28.3%), team capability dimension (31.2%), technological feasibility dimension (22.1%), and risk level dimension (18.4%).

[0087] Similar historical project retrospective: Select the 5 historical projects that are most similar to the project to be reviewed and have the greatest impact on the technical novelty score from the similar project set S_K, and output their project names, review conclusions and key similar technical features in tabular form.

[0088] Semantic Association Path Visualization: Generates a local subgraph of the project under review within a multi-dimensional R&D project knowledge graph, visually representing the semantic association paths between the project under review and related projects, technologies, and personnel. Nodes and edges in the local subgraph are distinguished by different colors.

[0089] Key Findings and Recommendations: The natural language generation module automatically integrates review results and key evidence, outputting review summaries and recommendations. For example: "Recommendation passed. This project has strong novelty in its technical solution, the team possesses the core technical capabilities required to execute the project, the technical route is feasible, and the risks are controllable. It is recommended to focus on the implementation plan for the training data acquisition and annotation of the deep learning model in the technical solution." In this embodiment, the system provides an expert feedback interface, which supports review experts to manually review and correct the automatically generated review results.

[0090] Feedback and Adjustment Instructions Received: After viewing the automatic review results and explanation report through the system interface, review experts can choose from three feedback options: "Agree with Automatic Review Results," "Revise Review Conclusion," or "Revise Score." When selecting "Revise Review Conclusion," experts must fill in the revised review conclusion and the reasons for the revision.

[0091] Incremental knowledge graph updates: Expert feedback is transformed into knowledge graph update triplets. For example, if an expert points out that "the novelty of this project should be lowered because a highly similar historical project was not correctly identified," the system adds the "similarity to a previous project" relationship between that historical project and the project to be reviewed to the knowledge graph.

[0092] Online model fine-tuning: The system triggers an online fine-tuning process, using the corrected review results and original feature data to construct fine-tuning samples. The heterogeneous graph neural network model and the multi-criteria fusion decision model are fine-tuned in a small number of rounds (usually 5-10 rounds) with a small learning rate (one-tenth of the original training learning rate) to update the model parameters.

[0093] Version management and rollback mechanism: After model fine-tuning, the system performs performance verification on the updated model, evaluating model performance metrics (AUC, accuracy, etc.) on the validation set. If the model performance degrades beyond a preset threshold after fine-tuning (e.g., AUC drops by more than 0.02), the system automatically rolls back to the model version before fine-tuning and records relevant exception information in the system log.

[0094] The knowledge graph-based intelligent review and decision-making method for R&D projects provided by this invention can be widely applied to the following scenarios: application and review of national and local scientific research fund projects; project approval and priority ranking of internal R&D projects within enterprises; phase review and completion acceptance of science and technology program projects; and evaluation and decision-making for technology introduction and technology cooperation projects. This method can effectively improve the objectivity, consistency, and efficiency of R&D project review, reduce the human and time costs of review, and has significant industrial practical value.

Claims

1. A knowledge graph-based intelligent review and decision-making method for R&D projects, characterized in that, Includes the following steps: S1. Obtain multimodal project data of the R&D projects to be reviewed. The multimodal project data includes project application text, technical solution documents, project team resume data, and historical project data. S2. Extract entities, relationships, and attributes from the multimodal project data to construct a multidimensional R&D project knowledge graph with projects as the core nodes. The multidimensional R&D project knowledge graph includes project entity nodes, technology entity nodes, personnel entity nodes, organization entity nodes, achievement entity nodes, and semantic relationship edges between each node. S3. Use a pre-trained heterogeneous graph neural network model to perform semantic encoding on the multi-dimensional R&D project knowledge graph, generate node embedding representations of each entity node, and aggregate the node embedding representations through graph pooling to obtain the graph-level representation of the multi-dimensional R&D project knowledge graph. S4. Based on the graph-level representation, calculate the multi-dimensional semantic similarity between the R&D project to be reviewed and the previously reviewed projects, and generate a technology novelty score, a team capability score, a technology feasibility score, and a risk level score based on the multi-dimensional semantic similarity. S5. Input the technology novelty score, team capability score, technology feasibility score and risk level score into the multi-criteria fusion decision model based on the attention mechanism, calculate the attention weight of each scoring dimension through adaptive weight learning, and perform weighted fusion of each scoring dimension based on the attention weight to generate a comprehensive review decision result. The comprehensive review decision results include: a binary decision of review pass / fail, and when the review is passed, the priority ranking of the R&D project to be reviewed among multiple projects to be reviewed based on the comprehensive score.

2. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 1, characterized in that, The construction of a multi-dimensional R&D project knowledge graph in step S2 specifically includes: S21. Based on a pre-trained natural language processing model, perform named entity recognition and relation extraction on the project application text and technical solution document to obtain project entities, technical entities and their technical relationships. S22. Extract personnel entities and their attribute information from the project team resume data. The attribute information includes technical expertise, research results, project experience and academic influence indicators, and construct a capability matching relationship between the personnel entities and the technical entities. S23. Extract existing project entities and their review result tags from the historical project data. The review result tags include review scores, pass status, and technical evaluation text. S24. Construct cross-project technology evolution relationship edges. Based on the semantic similarity and chronological order of the sets of technical keywords of each project, establish technology inheritance and technology differentiation relationships between projects. S25. Perform consistency verification on the constructed initial knowledge graph, detect and eliminate entity conflicts and relationship conflicts, and use a neighbor entity aggregation-based filling strategy for missing attribute values ​​to generate the multi-dimensional R&D project knowledge graph.

3. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 1, characterized in that, The pre-trained heterogeneous graph neural network model in step S3 includes: A node type-aware linear transformation layer is used to map the initial feature vectors of different types of nodes to a unified latent space. The node types include project type, technology type, personnel type, organization type, and achievement type. A multi-layer heterogeneous graph attention network is used to aggregate the feature information of multiple types of neighbor nodes of each node based on the attention mechanism in each layer and update the embedding representation of the node. In this network, an independent attention parameter matrix is ​​used for different types of relation edges. Interlayer residual connections and layer normalization are used to alleviate the gradient vanishing problem in deep networks and accelerate model convergence; The graph-level readout layer employs attention-weighted global pooling to assign different importance weights to each node in the knowledge graph, aggregating the embedded representations of all nodes into the graph-level representation.

4. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 3, characterized in that, The pre-training process of the heterogeneous graph neural network model includes: Obtain multimodal project data and corresponding review result data of historically reviewed projects as training samples; For each training sample, a corresponding historical project knowledge graph is constructed, and the comprehensive score in the review result data is used as the sample label. Based on the historical project knowledge graph and the sample labels, the heterogeneous graph neural network model is trained under supervised conditions using the mean squared error loss function to optimize the model parameters; and Contrastive learning is used to enhance the model's ability to distinguish similar items. During training, the consistency between graph representations of similar items within the same batch is maximized, while the consistency between graph representations of dissimilar items is minimized.

5. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 1, characterized in that, Step S4, which calculates the multi-dimensional semantic similarity between the R&D project to be reviewed and previously reviewed projects, specifically includes: The cosine similarity of the graph-level representation of the R&D project to be reviewed with the graph-level representation of each historically reviewed project is calculated to obtain the overall similarity. The technology subgraph, team subgraph, and achievement subgraph are extracted from the multi-dimensional R&D project knowledge graph. The subgraph-level representation of each subgraph is calculated, and the similarity of the technology dimension, team dimension, and achievement dimension is calculated based on the subgraph-level representation. Select the K historically reviewed projects that have the highest similarity to the overall R&D project to be reviewed to form a set of similar projects; Based on the review result labels of each project in the set of similar projects, the technological novelty score, team capability score, technological feasibility score, and risk level score are calculated using a weighted average method, with the weight being the normalized value of the overall similarity between the R&D project to be reviewed and the corresponding historical project.

6. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 1, characterized in that, The multi-criteria fusion decision model based on the attention mechanism in step S5 includes: A scoring embedding layer is used to map the technology novelty score, team capability score, technology feasibility score, and risk level score into a high-dimensional scoring embedding vector. A self-attention encoding layer is used to perform self-attention operations on the high-dimensional rating embedding vector, capture the interaction relationships between each rating dimension, and generate a context-aware rating representation. A cross-attention layer is used to perform cross-attention fusion between the context-aware rating representation and the graph-level representation, so that the weight learning of the rating dimension can perceive the deep semantic information contained in the knowledge graph. The decision output layer includes a parallel binary classifier and a ranking regressor. The binary classifier outputs the probability values ​​of passing / failing the review, and the ranking regressor outputs the priority score of the project.

7. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 6, characterized in that, The training process of the multi-criteria fusion decision model includes: Collect historical review data, including the scoring data of each historical project under multiple dimensions and the corresponding actual review conclusions, and construct a training dataset; Using the multi-dimensional scores and actual review conclusions of the historical projects as supervision signals, the binary classifier is trained using the cross-entropy loss function, and the ranking regressor is trained using the pairwise ranking loss function. The pairwise ranking loss function is constructed based on the superiority-inferiority relationship between each pair of items in the actual review conclusions. Positive sample pairs and negative sample pairs are sampled from the set of passed items and the set of failed items, respectively. During training, the attention weights are automatically learned through gradient backpropagation, eliminating the need for manual setting of prior values ​​for each dimension's weights.

8. A knowledge graph-based intelligent review and decision-making method for R&D projects according to any one of claims 1 to 7, characterized in that, The process after step S5 also includes: S6. Generate a traceable review decision explanation report, including: Based on the attention weights of each scoring dimension learned in the multi-criteria fusion decision model, a contribution distribution map of each dimension is generated. Tracing back the calculation path of the technology novelty score, output the historical projects and key technical features that have the greatest impact on the technology novelty of the R&D project to be reviewed; Generate a partial subgraph of the R&D project to be reviewed in the multi-dimensional R&D project knowledge graph, and visualize the semantic association path between the R&D project to be reviewed and related projects, technologies, and personnel.

9. The intelligent review and decision-making method for R&D projects based on knowledge graphs according to claim 8, characterized in that, Following step S6, the following is also included: S7. Obtain feedback and adjustment instructions from the review experts, wherein the feedback and adjustment instructions include the proposed corrections and reasons for the comprehensive review decision results; S8. Convert the proposed corrections into knowledge graph update triples and perform incremental update operations on the knowledge graph. S9. Based on the revised review results and the original feature data, construct fine-tuning samples, and perform online fine-tuning of the heterogeneous graph neural network model and the multi-criteria fusion decision model, dynamically updating the model parameters.

10. A knowledge graph-based intelligent review and decision-making system for R&D projects, characterized in that, include: The data acquisition module is used to acquire multimodal project data of the R&D projects to be reviewed. The multimodal project data includes project application text, technical solution documents, project team resume data, and historical project data. The knowledge graph construction module is used to extract entities, relationships and attributes from the multimodal project data and construct a multi-dimensional R&D project knowledge graph with projects as the core nodes. The graph semantic encoding module is used to perform semantic encoding on the multi-dimensional R&D project knowledge graph using a pre-trained heterogeneous graph neural network model, generating node embedding representations and graph-level representations; The multi-dimensional scoring module is used to calculate the multi-dimensional semantic similarity between the R&D project to be reviewed and the previously reviewed projects based on the graph-level representation, and to generate a technology novelty score, a team capability score, a technology feasibility score, and a risk level score. The fusion decision module is used to perform weighted fusion of the multi-dimensional scores by the multi-criteria fusion decision model based on the attention mechanism, and generate a comprehensive review decision result, including a binary decision of review pass / fail and priority ranking. The system is used to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Project decision optimization control method, device and equipment based on knowledge graph

    CN120509686B