A Method and System for Matching Research Directions with Professional and Technical Personnel Based on Knowledge Graphs

CN122196575BActive Publication Date: 2026-08-14BEIJING SCI & TECH PATENT OFFICE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-11
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

人工经验匹配虽然在一定程度上能结合实际情况,但受限于个人的知识储备和认知偏差,难以全面、客观地评估人才与项目的适配度,且效率低下,难以应对大规模的匹配需求

Benefits of technology

[0008]结合上述任一方面,通过构建科研知识图谱网络,将科研项目和专业技术人才以节点的形式呈现,并利用科研攻关关联边建立二者之间的联系,直观地展示了科研生态中的复杂关系,对科研知识图谱网络进行结构特征提取,得到项目节点和人才节点的结构嵌入特征集合,能够从网络结构的角度挖掘科研项目和专业技术人才的潜在特征,调用预训练的特征提取器对描述文本进行语义编码,获取项目需求和人才能力的文本语义特征,跨模态特征融合匹配处理综合考虑了结构特征和语义特征,生成多维匹配度度量参数集合,评估了科研项目单元与专业技术人才单元之间的适配程度,基于该参数集合筛选目标人才并生成人才匹配推荐列表,提高了科研资源配置的精准度和效率,有助于提升科研成果质量和创新性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122196575B_ABST
    Figure CN122196575B_ABST
Patent Text Reader

Abstract

This invention provides a knowledge graph-based method and system for matching research directions with professional and technical personnel, relating to the field of scientific research management technology. It aims to address the difficulty of accurately matching research projects with talent capabilities in the context of tackling key core technologies. First, it acquires a set of information on research projects and professional and technical personnel, constructing a research knowledge graph network containing research project nodes, professional and technical personnel nodes, and research-related edges. Next, it extracts the graph's structural features, generating a set of structural embedding features for project nodes and talent nodes. Then, it uses a pre-trained feature extractor to obtain the textual semantic features of project requirements and talent capabilities. Finally, it performs cross-modal feature fusion matching to generate a multi-dimensional matching degree measurement parameter set, filtering out the matched target professional and technical personnel units and generating a talent matching recommendation list. This invention significantly improves the accuracy and efficiency of scientific research resource allocation, providing decision support for organized scientific research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of scientific research management technology, and more specifically, to a method and system for matching scientific research directions with professional and technical personnel based on knowledge graphs. Background Technology

[0002] In the field of scientific research management, the rational and efficient allocation of scientific research resources and the precise and effective matching of research directions with professional and technical personnel are crucial to the smooth progress of scientific research. With the increasing complexity and diversity of information regarding research projects, technological achievements, and professional directions, and the continuous expansion of the professional and technical personnel workforce, how to scientifically utilize advanced technological means to further improve the matching degree and accuracy between research projects and professional and technical personnel is an urgent problem to be solved in scientific research management and a powerful guarantee for tackling key core technologies.

[0003] Traditional methods primarily rely on human experience and simple keyword matching. While human experience-based matching can incorporate practical considerations to some extent, it is limited by individual knowledge and cognitive biases, making it difficult to comprehensively and objectively assess the suitability of talent and projects. Furthermore, it is inefficient and struggles to handle large-scale matching demands. Simple keyword matching, on the other hand, is too mechanical, failing to deeply understand the semantic information in the text describing research projects and professional technical personnel. It easily overlooks key features and potential connections, leading to inaccurate matching results that cannot meet the precise matching requirements of scientific research. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a method and system for matching scientific research directions with professional and technical personnel based on knowledge graphs.

[0005] In conjunction with the first aspect of this application, a method for matching research directions with professional and technical personnel based on knowledge graphs is provided, which is applied to a knowledge graph-based research direction and professional and technical personnel matching system. The method includes: Acquire a set of scientific research project information and a set of professional and technical personnel information, wherein the set of scientific research project information includes multiple scientific research project units and the set of professional and technical personnel information includes multiple professional and technical personnel units; A scientific research knowledge graph network is constructed based on the project tackling direction description text in the scientific research project information set and the professional and technical personnel ability description text in the professional and technical personnel information set. The scientific research knowledge graph network includes a set of scientific research project nodes, a set of professional and technical personnel nodes, and a set of scientific research tackling association edges connecting the scientific research project nodes and the professional and technical personnel nodes. The scientific research knowledge graph network is subjected to graph structure feature extraction processing to generate the project node structure embedding feature set of the scientific research project node in the scientific research knowledge graph network and the talent node structure embedding feature set of the professional and technical talent node in the scientific research knowledge graph network. The pre-trained project requirement feature extractor is invoked to perform semantic encoding processing on the text describing the project's key research direction of the scientific research project unit, and the resulting text semantic feature vector is used as the semantic feature of the project requirement text. At the same time, the pre-trained talent ability feature extractor is invoked to perform semantic encoding processing on the text describing the professional and technical talent ability of the professional and technical talent unit, and the resulting text semantic feature vector is used as the semantic feature of the talent ability text. Cross-modal feature fusion and matching processing is performed on the project node structure embedding feature set, the talent node structure embedding feature set, the project requirement text semantic features, and the talent ability text semantic features to generate a multi-dimensional matching degree measurement parameter set between the scientific research project unit and the professional and technical talent unit. Based on the multi-dimensional matching degree measurement parameter set, target professional and technical talent units that match the scientific research project unit are selected from the professional and technical talent information set, and a talent matching recommendation list containing the correspondence between the identifier of the target professional and technical talent unit and the identifier of the scientific research project unit is generated.

[0006] In conjunction with the second aspect of this application, a knowledge graph-based research direction and professional technical personnel matching system is provided. The knowledge graph-based research direction and professional technical personnel matching system includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the knowledge graph-based research direction and professional technical personnel matching system implements the aforementioned knowledge graph-based research direction and professional technical personnel matching method.

[0007] In conjunction with the third aspect of this application, a computer-readable storage medium is provided, wherein computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed, the aforementioned method for matching research directions with professional and technical personnel based on knowledge graphs is implemented.

[0008] Combining any of the above aspects, by constructing a scientific research knowledge graph network, scientific research projects and professional and technical personnel are presented in the form of nodes, and the connections between the two are established using scientific research breakthroughs. This intuitively demonstrates the complex relationships in the scientific research ecosystem. Structural feature extraction is performed on the scientific research knowledge graph network to obtain the structural embedding feature sets of project nodes and talent nodes. This allows for the mining of potential features of scientific research projects and professional and technical personnel from the perspective of network structure. A pre-trained feature extractor is used to semantically encode the descriptive text to obtain the textual semantic features of project requirements and talent capabilities. Cross-modal feature fusion matching processing comprehensively considers structural and semantic features to generate a multi-dimensional matching degree measurement parameter set. This evaluates the fit between scientific research project units and professional and technical personnel units. Based on this parameter set, target talents are screened and a talent matching recommendation list is generated, which improves the accuracy and efficiency of scientific research resource allocation and helps to improve the quality and innovation of scientific research results. Attached Figure Description

[0009] Figure 1 This application provides a flowchart illustrating the method for matching research directions with professional and technical personnel based on knowledge graphs. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0013] Figure 1 This paper illustrates a flowchart of a knowledge graph-based method for matching research directions with professional and technical personnel, as provided in an embodiment of this application. The method includes, in detail: Step S110: Obtain a set of scientific research project information and a set of professional and technical personnel information. The set of scientific research project information includes multiple scientific research project units, and the set of professional and technical personnel information includes multiple professional and technical personnel units.

[0014] In this embodiment, firstly, basic information on all research projects in the application, ongoing, or completed stages is extracted from the research project management database to construct a research project information set. This set contains multiple research project units, each corresponding to an independent research project. The data structure of each research project unit is defined as including attribute fields such as a unique project identifier (represented by the symbol P_ID), project name, description of the project's research direction (represented by the symbol P_Desc), project leader, project start and end dates, and project funding amount. Simultaneously, basic information on all currently employed or collaborating professional and technical personnel is extracted from the talent resource management system to construct a professional and technical personnel information set. This set contains multiple professional and technical personnel units, each corresponding to an independent talent. The data structure of each talent unit is defined as including attribute fields such as a unique talent identifier (represented by the symbol T_ID), talent name, description of professional and technical personnel capabilities (represented by the symbol T_Desc), affiliated institution, research direction, and historical project participation records. The construction of these two sets provides the basic data source for subsequent knowledge graph construction and matching analysis. During the data extraction process, sensitive information involving personal privacy, such as contact information and ID numbers of talents, is anonymized using desensitization techniques, retaining only non-sensitive attributes used for matching analysis, to ensure that the data processing process complies with laws, regulations, and ethical requirements.

[0015] Step S120: Construct a scientific research knowledge graph network based on the project research direction description text in the scientific research project information set and the professional and technical personnel ability description text in the professional and technical personnel information set. The scientific research knowledge graph network includes a set of scientific research project nodes, a set of professional and technical personnel nodes, and a set of scientific research research association edges connecting the scientific research project nodes and the professional and technical personnel nodes.

[0016] Next, the research project information set and the professional and technical personnel information set obtained in step S110 are processed for knowledge graph construction. The core of this step is to transform unstructured text descriptions into structured relationships between nodes and edges. First, each research project unit in the research project information set is abstracted as a node in the research knowledge graph network, with the node type being "research project" and the node attributes including the project's P_ID and P_Desc. Similarly, each talent unit in the professional and technical personnel information set is abstracted as a node, with the node type being "professional and technical personnel" and the node attributes including the talent's T_ID and T_Desc. Based on this, by analyzing the semantic relationship between P_Desc and T_Desc, edges connecting the two types of nodes are constructed, namely, research and development relationship edges. These relationship edges are not generated out of thin air, but are specifically implemented through the following series of sub-steps.

[0017] Step S121: Extract key technical phrases from the description text of the research project unit's research direction to obtain a set of key technical phrases for the research project unit.

[0018] In this embodiment, for any scientific research project unit, its project research direction description text P_Desc is taken as input, and the following sub-steps are performed to extract key technical phrases that can represent the core technical research direction of the project.

[0019] Step S1211: Perform text sentence segmentation on the project tackling direction description text of the scientific research project unit to obtain the project tackling direction description sentence sequence contained in the project tackling direction description text.

[0020] First, the project's key objectives description text P_Desc is input into the sentence segmentation module. This module segments the text based on a preset set of sentence boundary punctuation marks, including periods, question marks, exclamation marks, and semicolons. Specifically, the module iterates through each character in the text, and when it detects one of the aforementioned sentence boundary punctuation marks, it identifies the continuous character sequence preceding that position as a complete sentence. After segmentation, the original project key objectives description text P_Desc is converted into an ordered sentence sequence, denoted as Sents_P={Sent_P1, Sent_P2, ..., Sent_Pm}, where m is the total number of sentences segmented from the text. For example, if P_Desc contains three natural sentences, a sentence sequence containing three elements is generated.

[0021] Step S1212: Perform part-of-speech tagging on each project problem description statement in the project problem problem description statement sequence to identify the noun phrase units and verb phrase units in the project problem problem description statements.

[0022] Next, for each sentence Sent_Pi in the sentence sequence Sents_P generated in step S1211, it is input into the part-of-speech tagging module. This module uses a sequence tagging model based on conditional random fields. The model takes each lexical unit in the sentence as input and predicts a part-of-speech tag for each lexical unit using parameters pre-trained on a large-scale corpus. The tagging system uses a general part-of-speech tagging set. For example, the tag set includes "noun", "verb", "adjective", "adverb", etc. The output of the part-of-speech tagging module is a tag sequence of the same length as the input sentence. Based on this tag sequence, noun phrase units and verb phrase units are identified from the sentence. A noun phrase unit consists of one or more consecutive lexical units labeled as nouns or adjectives; a verb phrase unit consists of one or more consecutive lexical units labeled as verbs. Each identified phrase unit contains its text content and its start and end position indices in the original sentence.

[0023] Step S1213: Perform phrase combination pattern matching on the noun phrase unit and the verb phrase unit according to the preset key technology phrase word formation pattern library to generate candidate key technology phrases that conform to the key technology phrase word formation pattern library.

[0024] A pre-defined key technology phrase formation pattern library is invoked. This library is a set of predefined rules containing phrase formation patterns that frequently appear in technical literature. For example, the library contains patterns such as "noun + noun", "adjective + noun", "verb + noun", and "noun + verb". For all noun phrase units and verb phrase units identified in step S1212, they are combined in the order they appear in the original sentence. Specifically, the algorithm iterates through adjacent phrase units in the sentence, checking whether their part-of-speech sequence matches any pattern in the pattern library. For example, if a noun phrase is immediately followed by another noun phrase, the combined phrase matches the pattern "noun + noun", and this combined phrase is extracted as a candidate key technology phrase. If an adjective phrase is immediately followed by a noun phrase, the combined phrase matches the pattern "adjective + noun" and is also extracted. In this way, the original simple phrase units are combined into semantically richer, more representative compound phrases, forming the initial candidate key technology phrase list Cand_Phrases_P.

[0025] Step S1214: Statistical processing is performed on the frequency of occurrence of the candidate key technology phrases in the description text of the project's key technology direction to obtain the project's frequency of occurrence parameter for each candidate key technology phrase.

[0026] The number of times each phrase in the candidate key technology phrase list Cand_Phrases_P generated in step S1213 appears in the original project's key technology direction description text P_Desc is counted. The count uses string matching, considering complete phrase matches while ignoring case sensitivity and partial word form variations. The occurrence count of each candidate phrase KP_Cand_i is recorded as its frequency parameter within the project, denoted as Freq_InProj_i. This parameter reflects the importance of the phrase in describing the project's own needs.

[0027] Step S1215: Calculate the inverse document frequency parameter of the candidate key technology phrase at the set level of all scientific research project units to obtain the global scarcity weight parameter of each candidate key technology phrase.

[0028] To measure the scarcity of a candidate phrase within the entire research project set, its inverse document frequency (IDF) needs to be calculated. First, the project focus description texts of all research project units in the entire research project information set are collected, constructing a global document set, Corpus_Proj. For each candidate key technology phrase KP_Cand_i, the number of documents containing this phrase in Corpus_Proj is counted, denoted as DocFreq_i. The total number of global documents is denoted as TotalDocs_Proj. The global scarcity weight parameter of this phrase, i.e., the inverse document frequency, is calculated using the formula IDF_Global_i = log(TotalDocs_Proj / (DocFreq_i + 1)). The denominator is incremented by 1 to avoid division by zero errors when the phrase does not appear in the document set. A larger IDF_Global_i value indicates a scarcer phrase and stronger discriminative power.

[0029] Step S1216: Based on the product calculation result of the frequency parameter in the project and the global scarcity weight parameter, calculate the importance score of each candidate key technology phrase. According to the preset importance score threshold or the preset selection quantity, select the K candidate key technology phrases whose importance scores exceed the importance score threshold or whose importance scores rank first, and form the set of key technology phrases for the project.

[0030] For each candidate key technology phrase KP_Cand_i, its frequency parameter Freq_InProj_i within the project is multiplied by the global scarcity weight parameter IDF_Global_i to obtain the importance score Score_Phrase_i = Freq_InProj_i * IDF_Global_i. This product combines the phrase's local importance within the project and its global discriminative power across the entire corpus. After calculating the scores of all candidate phrases, a preset filtering strategy is adopted. If a threshold strategy is adopted, an importance score threshold Thresh_Imp is set, and all phrases with Score_Phrase_i greater than Thresh_Imp are retained. If a quantity strategy is adopted, a desired number of key technology phrases K is set, and all candidate phrases are sorted from high to low according to Score_Phrase_i, and the top K phrases are selected. The selected phrases constitute the set of key technology phrases for this project unit, denoted as KP_P = {KP_P1, KP_P2, ..., KP_PK}.

[0031] Step S1217: Perform phrase standardization processing on each key technology phrase in the set of key technology phrases for project breakthroughs, and merge key technology phrases with different expressions having the same semantics into a unified standard key technology phrase identifier.

[0032] Finally, the initially selected set of key technical phrases, KP_P, is normalized to eliminate synonyms and variants. A thesaurus of key technical phrases is established, which pre-collects common technical phrases in the research field and their standard expressions. For example, the dictionary defines "machine learning" and "MachineLearning" as mapped to the standard identifier "machine learning," and defines "convolutional neural network" and "CNN" as mapped to "convolutional neural network." Each phrase in KP_P is traversed, and its corresponding standard identifier is searched in the dictionary. If found, the original phrase is replaced with the standard identifier; otherwise, the original phrase is retained. This step ensures that the same technical concept appears only once in the set, providing an accurate basis for subsequent semantic mapping and similarity calculation.

[0033] Step S122: Extract key technical phrases of talent expertise from the text describing the professional and technical talent capabilities of the professional and technical talent unit to obtain a set of key technical phrases of talent expertise for the professional and technical talent unit.

[0034] For any professional and technical personnel unit, take its professional and technical personnel capability description text T_Desc as input, and extract key technical phrases that can represent the core technical capabilities of the personnel through the following sub-steps.

[0035] Step S1221: Perform text segmentation processing on the professional and technical personnel ability description text of the professional and technical personnel unit to obtain the personnel ability description paragraph units contained in the professional and technical personnel ability description text.

[0036] First, the text T_Desc describing the professional and technical personnel's abilities is input into the paragraph segmentation module. This module segments the text into multiple independent paragraph units describing the personnel's abilities based on natural paragraph markers, i.e., two consecutive line breaks. Each paragraph typically describes the personnel's experience or ability in a certain area, such as "educational background," "work experience," "research projects," "representative achievements," etc. After segmentation, T_Desc is converted into an ordered sequence of paragraphs, denoted as Pars_T={Par_T1, Par_T2, ..., Par_Tn}, where n is the total number of paragraphs segmented from the text.

[0037] Step S1222: Perform sentence boundary detection processing on the talent ability description paragraph unit, and segment each talent ability description paragraph unit into a talent ability description sentence sequence.

[0038] For each paragraph Par_Tj in the paragraph sequence Pars_T generated in step S1221, it is input into the sentence boundary detection module. This module uses a punctuation-based segmentation algorithm similar to that in step S1211, but performs segmentation within each paragraph. Specifically, the module traverses the paragraph text and, based on sentence-ending punctuation such as periods, question marks, and exclamation marks, segments the paragraph text into one or more sentences, forming a sequence of sentences describing the talent abilities of that paragraph. For example, for paragraph Par_Tj, after segmentation, the sentence sequence Sents_Tj={Sent_Tj1, Sent_Tj2, ..., Sent_Tjmj} is obtained, where mj is the number of sentences segmented from the paragraph.

[0039] Step S1223: Perform syntactic dependency analysis on each statement in the sequence of statements describing talent ability to identify the core verb units and core noun units that have a dependency relationship with the core verb units.

[0040] For each statement `Sent_Tji`, it is input into a syntactic dependency parsing model. This model employs a dependency parser based on a graph neural network, capable of analyzing the grammatical dependencies between lexical units in the sentence. The model output is a dependency tree, where each node is a lexical unit, and each directed edge represents the type of dependency relationship between two words, such as "subject-verb," ​​"verb-object," or "attributive-head." By traversing and analyzing the dependency tree, the core verb units in the sentence are identified. The core verb is usually the predicate of the sentence, the center of the action or state. Then, core noun units with specific dependency relationships to the core verb are identified. For example, nouns with a "verb-object" relationship or a "subject-verb" relationship with the core verb are found. These combinations of core verbs and core nouns often form core phrases describing technical capabilities. The identified core verb units are denoted as `Verb_Core`, and the core noun units are denoted as `Noun_Core`.

[0041] Step S1224: Based on the dependency relationship type between the core verb unit and the core noun unit, combine the core verb unit and the core noun unit into candidate proficient technical phrases.

[0042] Based on the core verb Verb_Core and core noun Noun_Core identified in step S1223, and the dependency relationships between them, phrase combinations are performed. If there is a "verb-object relationship" between Verb_Core and Noun_Core, the combination form is "Verb_Core + Noun_Core", for example, "researching deep learning". If there is a "subject-verb relationship", the combination form is "Noun_Core + Verb_Core", for example, "algorithm optimization". If there are other related relationships such as "prepositional phrase relationship", the combination is performed according to a preset combination template. For each statement, one or more candidate technical phrases composed of verbs and nouns may be extracted to form the candidate phrase list Cand_Phrases_T for that talent.

[0043] Step S1225: Perform statistical processing on the distribution density of the candidate technical skill phrases in the professional and technical personnel ability description text to obtain the paragraph coverage rate parameter of each candidate technical skill phrase in the professional and technical personnel ability description text.

[0044] For each phrase KP_Cand_Ti in the candidate phrase list Cand_Phrases_T generated in step S1224, count the number of paragraphs in all paragraphs Pars_T divided in step S1221 that it appears in. If a phrase appears in any sentence within a paragraph, it is considered that the paragraph covers the phrase. Count the number of paragraphs that cover the phrase, denoted as ParaCover_i. The total number of paragraphs is n. Then, the paragraph coverage parameter Coverage_Para_i = ParaCover_i / n. This parameter reflects whether the ability described by the phrase is reflected in different stages or aspects of the talent's experience, thus determining whether it is the core technical direction that the talent has been engaged in for a long time.

[0045] Step S1226: Based on the paragraph coverage parameter, the candidate proficiency technical phrases are filtered, and the candidate proficiency technical phrases whose paragraph coverage parameter exceeds the preset coverage threshold are retained as the set of key technical phrases proficient in the talent.

[0046] A preset coverage threshold, Thresh_Coverage, is set, for example, 0.5. This means that if a phrase appears in more than half of the paragraphs, it is considered a core competency of the talent. All candidate phrases KP_Cand_Ti are iterated through, and their paragraph coverage parameter Coverage_Para_i is checked against Thresh_Coverage. If it is greater, the phrase is retained; if less, it is filtered out, as it may be a non-core or incidentally mentioned technology. The phrases retained after filtering constitute the set of key competency phrases for this professional technical talent unit, denoted as KP_T = {KP_T1, KP_T2, ..., KP_TL}, where L is the final number of selected phrases.

[0047] Step S123: Based on the semantic vector space similarity calculation results between the semantic vector of the key technology phrase in the project and the semantic vector of the key technology phrase in the talent’s expertise, construct the initial scientific research connection edge between the scientific research project node and the professional and technical talent node.

[0048] After obtaining the set of key technical phrases KP_P on the project side and the set of key technical phrases KP_T on the talent side, the semantic relationship between the two is quantified through the following sub-steps, and the associated edges in the graph are constructed accordingly.

[0049] Step S1231: Perform a vector dot product operation on the semantic vectors of the key technology phrases for tackling key problems in the project and the semantic vectors of the key technology phrases for which the talent excels, to obtain a phrase-level semantic similarity matrix between the key technology phrases for tackling key problems in the project and the key technology phrases for which the talent excels.

[0050] The same pre-trained language model as in step S123 is used as the phrase encoder. First, each phrase KP_Pi from the project-side phrase set KP_P is input into the encoder to obtain its corresponding semantic vector V_Pi, with dimension d. Similarly, each phrase KP_Tj from the talent-side phrase set KP_T is input into the encoder to obtain the semantic vector V_Tj, with dimension d. Then, the vector dot product between all V_Pi and all V_Tj is calculated. Specifically, a matrix Sim_Matrix with shape K rows and L columns is constructed. For the element Sim_ij in the i-th row and j-th column of the matrix, its value is the dot product of V_Pi and V_Tj, i.e., Sim_ij=Σ(V_Pi_k*V_Tj_k), where the summation ranges from k=1 to d. This matrix Sim_Matrix is ​​the phrase-level semantic similarity matrix, which quantifies the degree of matching between each key technology phrase in the project and each key technology phrase in the talent capabilities.

[0051] Step S1232: Perform row-direction maximum pooling on the phrase-level semantic similarity matrix, and extract the maximum similarity value between the key technology phrases for each project and the key technology phrases for all talents, as the maximum similarity vector on the project side.

[0052] Perform row-wise operations on the Sim_Matrix. For each row of the matrix, i.e., for each item phrase KP_Pi, find the maximum value among the L elements in that row. This maximum value, Max_Pi = max(Sim_i1, Sim_i2, ..., Sim_iL), represents the best match between that specific item phrase and all talent ability phrases. Collect the maximum values ​​from all K rows to form a vector of length K, Max_P = [Max_P1, Max_P2, ..., Max_PK]. This vector is called the item-side maximum similarity vector.

[0053] Step S1233: Perform column-direction maximum pooling on the phrase-level semantic similarity matrix, and extract the maximum similarity value between the key technology phrases that each talent excels in and the key technology phrases that all projects tackle, as the maximum similarity vector on the talent side.

[0054] Similarly, column-wise operations are performed on the Sim_Matrix. For each column of the matrix, i.e., for each talent phrase KP_Tj, find the maximum value among the K elements in that column. This maximum value, Max_Tj = max(Sim_1j, Sim_2j, ..., Sim_Kj), represents the best match between that specific talent ability and all the requirement phrases of the project. The maximum values ​​from all L columns are collected to form a vector of length L, Max_T = [Max_T1, Max_T2, ..., Max_TL], which is called the maximum similarity vector on the talent side.

[0055] Step S1234: Perform mean calculation on the maximum similarity vector on the project side to obtain the average semantic similarity parameter between the scientific research project unit and the professional and technical personnel unit in the project-to-talent direction.

[0056] Calculate the arithmetic mean of all elements in the maximum similarity vector Max_P on the project side. That is, Avg_Sim_PtoT = (Max_P1 + Max_P2 + ... + Max_PK) / K. This parameter Avg_Sim_PtoT represents the best average degree of matching that can be found in the talent pool for all key technical requirements from the project's perspective, reflecting the overall level to which project requirements are met by talent.

[0057] Step S1235: Perform mean calculation on the maximum similarity vector on the talent side to obtain the average semantic similarity parameter between the scientific research project unit and the professional and technical talent unit in the talent-to-project direction.

[0058] Calculate the arithmetic mean of all elements in the maximum similarity vector Max_T on the talent side. That is, Avg_Sim_TtoP = (Max_T1 + Max_T2 + ... + Max_TL) / L. This parameter Avg_Sim_TtoT represents the best average degree of matching that all of the talent's core technical capabilities can find in the project's requirement set from the talent's perspective, reflecting the overall level at which the talent's potential is utilized in the project.

[0059] Step S1236: Based on preset weight coefficients, perform weighted fusion processing on the average semantic similarity parameter from project to talent and the average semantic similarity parameter from talent to project to generate a comprehensive semantic matching parameter between the scientific research project unit and the professional and technical talent unit.

[0060] To achieve a comprehensive and balanced match, weighting coefficients α and β are introduced, representing the importance of the project-side and talent-side perspectives, respectively. Typically, α + β = 1. The comprehensive semantic match score parameter Match_Score_Initial is calculated as α * Avg_Sim_PtoT + β * Avg_Sim_TtoP. This parameter integrates bidirectional matching information, avoiding biases that might arise from a one-sided perspective.

[0061] Step S1237: Based on the comprehensive semantic matching degree parameter, construct the initial scientific research breakthrough association edge carrying the comprehensive semantic matching degree parameter between the scientific research project node and the professional and technical personnel node.

[0062] Set an initial edge construction threshold, Thresh_Initial_Edge. Iterate through all possible project-talent pairs. For the current project node P_x and talent node T_y, if their comprehensive semantic matching degree parameter Match_Score_Initial exceeds Thresh_Initial_Edge, create an edge in the knowledge graph pointing from P_x to T_y (or bidirectionally). This edge is labeled as an "Initial Scientific Research Collaboration Edge," and its attributes store the calculated Match_Score_Initial value as the initial association strength weight. This edge signifies a potential collaboration possibility between the two based on core technology semantics.

[0063] Step S124: Based on the semantic vector space similarity calculation results between the semantic vector of the key technology phrase in the project and the semantic vector of the key technology phrase in the talent’s expertise, construct the initial scientific research connection edge between the scientific research project node and the professional and technical talent node.

[0064] After obtaining the vector sets V_Pi and V_Tj, the association between research project nodes and talent nodes is constructed. Specifically, firstly, the vector dot product between all V_Pi and all V_Tj is calculated, resulting in a phrase-level semantic similarity matrix of dimension K rows and L columns, denoted as Sim_Matrix, where each element Sim_ij = V_Pi·V_Tj. Then, this matrix is ​​subjected to row-wise max pooling, i.e., the maximum value of each row is taken, resulting in a project-side maximum similarity vector Max_P of length K, where each element Max_Pi represents the highest similarity between the i-th phrase of the project and all talent phrases. Similarly, column-wise max pooling is performed, resulting in a talent-side maximum similarity vector Max_T of length L. Next, the arithmetic mean of all elements in Max_P is calculated, yielding the average semantic similarity parameter Avg_Sim_PtoT from project to talent, which reflects the average best matching degree between the key technologies and talent capabilities from the project's perspective. Similarly, the mean of Max_T is calculated to obtain the average semantic similarity parameter Avg_Sim_TtoP from talent to project direction. To obtain a comprehensive matching degree, preset weight coefficients α and β are introduced to weight and fuse the two average similarities mentioned above, generating a comprehensive semantic matching degree parameter Match_Score_Initial=α*Avg_Sim_PtoT+β*Avg_Sim_TtoP. If this Match_Score_Initial exceeds a preset edge construction threshold, an initial scientific research connection edge is constructed between the corresponding scientific research project node P_ID and the professional and technical talent node T_ID. This edge carries Match_Score_Initial as its initial connection strength weight, representing the direct correlation between the two based on textual semantics.

[0065] Step S125: Obtain a set of historical research collaboration records, which includes the historical collaboration correspondence between project identifiers of multiple completed research projects and talent identifiers of professional and technical personnel who participated in the completed research projects.

[0066] To improve the accuracy of the knowledge graph, in addition to text mining-based associations, it is necessary to incorporate real, past collaborations. Therefore, a set of historical research collaboration records is obtained from a research project completion database. Each record is a binary tuple in the form (P_ID_History, T_ID_History), indicating that talent T_ID_History participated in and completed project P_ID_History. This set provides a reliable basis for connections within the knowledge graph.

[0067] Step S126: Based on the historical cooperation correspondence, supplement the historical cooperation association edges between the scientific research project nodes and professional and technical personnel nodes in the scientific research knowledge graph network.

[0068] Based on the historical cooperation records obtained in step S125, the nodes in the current graph are traversed. If there exists a project node whose P_ID matches the P_ID_History in the historical records, and a talent node whose T_ID matches the T_ID_History, and no edge has been established between these two nodes, then a historical cooperation connection edge is directly constructed between these two nodes. If an initial scientific research collaboration connection edge generated in step S124 already exists between these two nodes, then the attributes of this edge are updated, a "historical cooperation" tag is added, and its corresponding historical project ID is recorded.

[0069] Step S127: Merge the initial scientific research breakthrough association edges with the historical cooperation association edges to generate the set of scientific research breakthrough association edges in the scientific research knowledge graph network. Each scientific research breakthrough association edge in the set of scientific research breakthrough association edges contains a weighted fusion result of the initial association strength weight and the historical cooperation strength weight as a comprehensive association strength parameter.

[0070] At this point, the edges in the graph may have two sources: edges based on pure semantic matching and edges based on historical collaboration. To unify the measurement of edge strength, a fusion process is needed. For any pair of project nodes P and talent nodes T, if only an initial research collaboration connection edge exists, then the comprehensive association strength parameter Col_Strength of that edge is equal to its initial association strength weight Match_Score_Initial. If only a historical collaboration connection edge exists, then it is assigned a preset base strength value Base_Strength_Coop representing the collaboration fact. If both types of edges exist, then its comprehensive association strength parameter Col_Strength = γ * Match_Score_Initial + δ * Base_Strength_Coop is calculated, where γ and δ are weight coefficients, and γ + δ = 1. After this step, all edges connecting project nodes and talent nodes uniformly carry the comprehensive association strength parameter Col_Strength, forming a complete set of research collaboration connection edges.

[0071] Step S128: Optimize the network topology of the scientific research knowledge graph network. Based on the comprehensive association strength parameter, prune the scientific research problem association edges that are below the preset strength threshold to obtain the optimized scientific research knowledge graph network.

[0072] Finally, to reduce network noise and complexity, network topology optimization is performed. A global strength threshold, Thresh_Prune, is set. All edges related to scientific research projects are traversed; if the overall association strength parameter Col_Strength of an edge is less than Thresh_Prune, the association is considered weak or unreliable and is removed from the network. After pruning, the remaining edges form a more concise and reliable scientific knowledge graph network. The nodes and edges in this network provide high-quality input for subsequent structural feature extraction.

[0073] Step S130: Perform graph structure feature extraction processing on the scientific research knowledge graph network to generate the project node structure embedding feature set of the scientific research project node in the scientific research knowledge graph network and the talent node structure embedding feature set of the professional and technical talent node in the scientific research knowledge graph network.

[0074] Structural features are extracted from the optimized scientific knowledge graph network to encode the topological location information of nodes into low-dimensional dense vectors. This embodiment employs a graph neural network-based approach.

[0075] Step S131: Perform node neighborhood sampling processing on the scientific research knowledge graph network, and collect a set of first-order neighbor nodes that are directly connected to the scientific research project node through the scientific research breakthrough association edge for each scientific research project node.

[0076] Specifically, for any target scientific research project node P_x in the graph, first obtain all professional and technical personnel nodes that are directly connected to it through scientific research breakthroughs. These nodes constitute the first-order neighbor node set of P_x, denoted as Nbr1_Px={T_y1,T_y2,...}.

[0077] Step S132: Perform recursive sampling on each professional and technical talent node in the first-order neighbor node set, and collect other scientific research project nodes that are directly connected to each professional and technical talent node through scientific research breakthroughs as the second-order neighbor node set.

[0078] Next, for each professional and technical talent node T_y in Nbr1_Px, the above neighbor sampling process is repeated to obtain all research project nodes directly connected to T_y (except for the starting node P_x itself). These nodes constitute the set of second-order neighbor nodes of P_x, denoted as Nbr2_Px={P_z1, P_z2, ...}.

[0079] Step S133: Construct the local subgraph structure of the scientific research project node in the scientific research knowledge graph network based on the first-order neighbor node set and the second-order neighbor node set.

[0080] The target node P_x, its first-order neighbor set Nbr1_Px, and its second-order neighbor set Nbr2_Px, along with all the edges connecting them related to scientific research, are extracted to form a local subgraph G_Px centered on P_x. This subgraph fully describes the topological structure within two hops of P_x.

[0081] Step S134: Perform graph convolution propagation processing on the local subgraph structure of the research project node, and aggregate the feature information of the neighboring nodes in the local subgraph structure onto the research project node to generate the project node structure embedding feature of the research project node.

[0082] Perform graph convolution operations on the local subgraph G_Px. First, each node in the subgraph needs to be assigned initial features. Since the goal of this step is to extract structural features, the node's basic attributes (such as project or talent IDs) can be one-hot encoded or uniformly initialized as a constant vector. Then, information is propagated through a graph convolutional network layer. In the graph convolutional layer, for the target node P_x, its updated feature vector Emb_Px_new is calculated by aggregating the feature vectors of all its first-order neighbors T_y and combining them with its own features. The aggregation function can be summing or averaging the neighbor feature vectors, or using a more complex attention mechanism. For example, using a weighted summation approach, `Emb_Px_new=σ(W_self*Emb_Px_old+Σ(W_neigh*Emb_Ty_old*Col_Strength_PxTy))`, where `W_self` and `W_neigh` are trainable weight matrices, `Col_Strength_PxTy` is the comprehensive association strength parameter on the edges used as weights, and `σ` is a non-linear activation function. After one or more layers of graph convolution, the final vector representation of `P_x` is the item node structural embedding feature that incorporates its local topological structure, denoted as `Struct_Px`.

[0083] Step S135: Repeat the node neighborhood sampling process and local subgraph structure construction process for each professional and technical talent node in the scientific research knowledge graph network, and perform graph convolution propagation processing on the local subgraph structure of each professional and technical talent node to generate the talent node structure embedding feature of each professional and technical talent node.

[0084] For each professional and technical talent node T_x in the graph, repeat the operations that are completely symmetrical to steps S131 to S134. First, construct a local subgraph centered on T_x, containing its first-order neighbors (project nodes) and second-order neighbors (other talent nodes connected through project nodes). Then, using the same graph convolutional network parameters, perform information propagation and aggregation on this local subgraph to finally obtain the talent node structure embedding feature of T_x, denoted as Struct_Tx.

[0085] Step S136: Combine the project node structure embedding features of all scientific research project nodes into the project node structure embedding feature set, and combine the talent node structure embedding features of all professional and technical talent nodes into the talent node structure embedding feature set.

[0086] All calculated Struct_Px vectors are indexed and stored according to their corresponding project identifiers P_ID, forming the project node structure embedding feature set {Struct_P}. Similarly, all Struct_Tx vectors are indexed according to their talent identifiers T_ID, forming the talent node structure embedding feature set {Struct_T}. These two sets will serve as one of the key inputs for subsequent cross-modal fusion matching.

[0087] Step S140: Call the pre-trained project requirement feature extractor to perform semantic encoding processing on the text describing the project's key research direction of the scientific research project unit, and obtain the text semantic feature vector of the scientific research project unit as the semantic feature of the project requirement text. At the same time, call the pre-trained talent ability feature extractor to perform semantic encoding processing on the text describing the professional and technical talent ability of the professional and technical talent unit, and obtain the text semantic feature vector of the professional and technical talent unit as the semantic feature of the talent ability text.

[0088] Simultaneously, the textual descriptions of research project units are independently semantically encoded to capture their content-level requirements. This embodiment employs a project requirement feature extractor pre-trained on a large-scale corpus, the core of which is a Transformer-based deep neural network.

[0089] Step S141: Normalize the text length of the project tackling direction description text of the scientific research project unit, and truncate or fill the project tackling direction description text to a preset fixed text length.

[0090] First, the input P_Desc is preprocessed. Since pre-trained models typically require a fixed input length, a standard length L_seq is set. For text longer than L_seq, it is truncated, retaining only the first L_seq characters or vocabulary units. For text shorter than L_seq, special padding is added to the end until the length reaches L_seq.

[0091] Step S142: Input the normalized length description text of the project's key points into the word embedding layer of the pre-trained project requirement feature extractor, and map each lexical unit in the description text of the project's key points into a sequence of lexical embedding vectors.

[0092] The normalized text is input into the first layer of the feature extractor, namely the word embedding layer. This layer maintains a mapping matrix from the vocabulary to a high-dimensional vector space. Each word unit (which can be a character or a word) in the text is used to find the corresponding dense vector from the mapping matrix according to its index, ultimately resulting in a sequence of word embedding vectors, denoted as Seq_WordEmb={Emb_Word1, Emb_Word2, ..., Emb_WordL_seq}.

[0093] Step S143: Input the word embedding vector sequence into the multi-head self-attention layer of the pre-trained project requirement feature extractor to model the semantic dependencies between word units in the word embedding vector sequence and generate a context-aware word vector sequence carrying contextual semantic information.

[0094] Seq_WordEmb is then fed into a multi-head self-attention module with multiple stacks. In each self-attention layer, for each position in the sequence, attention weights are computed with respect to all other positions in the sequence (including itself). These weights reflect the semantic relevance of the current word to other words. Through the multi-head mechanism, the model can learn various types of dependencies from different representation subspaces. The outputs of each head are concatenated and linearly transformed. After multiple layers of the above processing, the vector at each position is aggregated with global contextual information, resulting in a context-aware word vector sequence, denoted as Seq_Context={Emb_Ctx1, Emb_Ctx2, ..., Emb_CtxL_seq}.

[0095] Step S144: Input the context-aware vocabulary vector sequence into the pooling layer of the pre-trained project requirement feature extractor, and perform average pooling on all vocabulary vectors in the context-aware vocabulary vector sequence to obtain sentence-level semantic vectors of fixed dimensions.

[0096] To convert the variable-length context vector sequence Seq_Context into a fixed-length vector, a pooling layer is required. Here, average pooling is used, which calculates the average of all L_seq vectors in Seq_Context across each dimension, ultimately resulting in a single, fixed-length vector, denoted as Sent_Vec_P. This vector represents the overall semantics of the text describing the project's key objectives.

[0097] Step S145: Input the sentence-level semantic vector into the fully connected mapping layer of the pre-trained project requirement feature extractor, perform dimensionality transformation on the sentence-level semantic vector, and output the text semantic feature vector of the scientific research project unit as the project requirement text semantic feature.

[0098] Finally, Sent_Vec_P is fed into a mapping network consisting of one or more fully connected layers. This network maps the sentence-level semantic vector to the same dimensional space as other modal features (such as structural embedding features) in subsequent steps, facilitating fusion and comparison. After the nonlinear transformation by the fully connected layers, the final output vector is the semantic feature of the project requirement text for this research project unit, denoted as Text_Px.

[0099] Similarly, semantic encoding is performed on the text T_Desc describing the abilities of professional and technical personnel. However, considering that the text describing the abilities of personnel is often longer and more diverse in structure (including educational background, work experience, project achievements, etc.), this embodiment uses a model based on a bidirectional long short-term memory network as the feature extractor for the abilities of personnel.

[0100] Step S146: Perform text segmentation processing on the professional and technical personnel ability description text of the professional and technical personnel unit, and divide the professional and technical personnel ability description text into multiple description segment units.

[0101] First, based on the natural paragraph markers (such as newlines) in T_Desc, it is divided into a sequence of paragraph descriptor units, denoted as Pars_T={Par_T1, Par_T2, ..., Par_Tn}.

[0102] Step S147: Perform text length normalization and word embedding mapping on each descriptive paragraph unit to generate a paragraph word embedding vector sequence for each descriptive paragraph unit.

[0103] For each paragraph Par_Tj in Pars_T, a process similar to steps S141 and S142 is performed independently. First, its length is normalized to the preset paragraph length L_para. Then, through a word embedding layer, each lexical unit in the paragraph is mapped to an embedding vector, resulting in the lexical embedding vector sequence Seq_ParaEmb_j for that paragraph.

[0104] Step S148: Input the paragraph word embedding vector sequence of each paragraph description unit into the bidirectional long short-term memory network layer of the pre-trained talent ability feature extractor, and perform forward and reverse sequence encoding processing on each paragraph word embedding vector sequence to obtain the forward hidden state sequence and the reverse hidden state sequence of each paragraph description unit.

[0105] The lexical embedding sequence Seq_ParaEmb_j for each paragraph is fed into a bidirectional long short-term memory (LSM) network layer. This layer contains a forward LSM network that reads the sequence sequentially (from the first word to the last word) and generates a forward hidden state h_f_t at each time step t, which encodes the current word and its preceding context. Simultaneously, a backward LSM network reads the sequence in reverse order (from the last word to the first word) and generates a backward hidden state h_b_t, which encodes the current word and its following context. For a paragraph of length L_para, after processing, the resulting forward hidden state sequence H_f_j = {h_f_1, h_f_2, ..., h_f_L_para} and backward hidden state sequence H_b_j = {h_b_1, h_b_2, ..., h_b_L_para} are obtained.

[0106] Step S149: Concatenate the last time step hidden state of the forward hidden state sequence of each paragraph unit with the first time step hidden state of the reverse hidden state sequence to obtain the paragraph semantic vector of each paragraph unit.

[0107] To obtain the semantic representation of the entire paragraph, the boundary hidden states output by the bidirectional long short-term memory network are usually concatenated. That is, the output h_f_L_para of the last time step of the forward hidden state sequence and the output h_b_1 of the first time step of the backward hidden state sequence are taken and concatenated along the feature dimension to form a vector that can represent the semantics of the entire paragraph, denoted as Para_Vec_j=[h_f_L_para;h_b_1].

[0108] Step S1410: Perform weighted summation on the paragraph semantic vectors of all descriptive paragraph units, and assign different weight coefficients to each descriptive paragraph unit according to its position order in the professional and technical personnel capability description text, thereby generating a document-level semantic vector that integrates information from multiple paragraphs.

[0109] Talent descriptions typically consist of multiple paragraphs, with varying levels of importance (e.g., the first paragraph is usually an overview, while the last might be a future outlook). Therefore, a position-based attention mechanism is employed to fuse the semantic vectors of all paragraphs. Each position j is assigned a learnable or pre-defined weight coefficient w_j, which may vary with position (e.g., middle paragraphs have higher weights). The document-level semantic vector Doc_Vec_T is obtained by weighted summation of all paragraph vectors: Doc_Vec_T = Σ(w_j * Para_Vec_j), where j ranges from 1 to n.

[0110] Step S1411: Input the document-level semantic vector into the output layer of the pre-trained talent ability feature extractor, and output the text semantic feature vector of the professional and technical talent unit as the talent ability text semantic feature.

[0111] Finally, the document-level semantic vector Doc_Vec_T, which integrates information from multiple paragraphs, is transformed in dimension through an output layer (which can be a simple linear layer) to make it have the same dimension as the project requirement text semantic feature Text_Px and the structural embedding feature, thus obtaining the talent capability text semantic feature, denoted as Text_Ty.

[0112] Step S150: Perform cross-modal feature fusion matching processing on the project node structure embedding feature set, the talent node structure embedding feature set, the project requirement text semantic features, and the talent ability text semantic features to generate a multi-dimensional matching degree measurement parameter set between the scientific research project unit and the professional and technical talent unit. Based on the multi-dimensional matching degree measurement parameter set, select target professional and technical talent units that match the scientific research project unit from the professional and technical talent information set, and generate a talent matching recommendation list containing the correspondence between the target professional and technical talent unit identifier and the scientific research project unit identifier.

[0113] At this point, for any research project P_x and any professional and technical personnel T_y, four different modalities of features have been obtained: project structure feature Struct_Px, personnel structure feature Struct_Ty, project text feature Text_Px, and personnel text feature Text_Ty. The goal of step S150 is to deeply fuse these four features to calculate a comprehensive score that can fully reflect the degree of matching between the two.

[0114] Step S151: Perform vector concatenation processing on each project node structure embedding feature in the project node structure embedding feature set and each talent node structure embedding feature in the talent node structure embedding feature set to generate a project-talent structure feature pair vector.

[0115] For a specific pair (P_x, T_y), their structural feature vectors are first concatenated to form a joint vector representing the structural relationship between the two. For example, if Struct_Px and Struct_Ty are both d-dimensional vectors, the concatenation results in a 2d-dimensional vector Concat_Struct_xy=[Struct_Px; Struct_Ty].

[0116] Step S152: Perform vector concatenation processing on the semantic features of the project requirement text and the semantic features of the talent ability text to generate a project-talent text feature pair vector.

[0117] Similarly, the project text feature Text_Px and the talent text feature Text_Ty are concatenated to obtain a 2d-dimensional vector Concat_Text_xy=[Text_Px;Text_Ty].

[0118] Step S153: Input the project-talent structure feature pair vector into the first matching degree calculation branch, and perform nonlinear transformation processing on the project-talent structure feature pair vector through the first multilayer perceptron network to output the structure dimension matching degree parameter.

[0119] The Concat_Struct_xy is fed into the first multilayer perceptron network. This network consists of multiple fully connected layers and non-linear activation functions (such as ReLU) stacked together. The input vector is first transformed by the first layer to map it to the hidden layer dimension, then passed through the activation function, and then fed into the next layer. After several layers of transformation, the final output layer of the network is a neuron with only one node, and its output value is a scalar representing the structural dimension matching degree parameter measured from the perspective of topological structure, denoted as Score_Struct_xy.

[0120] Step S154: Input the project-talent text feature pair vector into the second matching degree calculation branch, and perform nonlinear transformation processing on the project-talent text feature pair vector through the second multilayer perceptron network to output the text dimension matching degree parameter.

[0121] Similarly, Concat_Text_xy is fed into a second multilayer perceptron network with a potentially different structure (but also composed of fully connected layers). After nonlinear mapping, it finally outputs a scalar, representing the text dimension matching degree parameter measured from the perspective of text content, denoted as Score_Text_xy.

[0122] Step S155: Perform feature cross processing on the embedded features of the project node structure and the semantic features of the project requirement text to generate a project-side fused feature vector.

[0123] To capture the interaction information between the project's internal structural features and text features, Struct_Px and Text_Px are cross-processed. An effective cross-processing method is to perform a vector outer product to obtain a d*d matrix, and then flatten this matrix into a d^2 dimensional vector. Alternatively, a simpler approach can be used: concatenate the vectors again to generate the project-side fused feature vector Fuse_Px=[Struct_Px; Text_Px].

[0124] Step S156: Perform feature cross processing on the embedded features of the talent node structure and the semantic features of the talent ability text to generate a talent-side fusion feature vector.

[0125] Similarly, the same cross-processing is performed on the talent side, concatenating Struct_Ty and Text_Ty to generate the talent side fusion feature vector Fuse_Ty=[Struct_Ty;Text_Ty].

[0126] Step S157: Perform attention interaction processing on the project-side fusion feature vector and the talent-side fusion feature vector, calculate the attention weight distribution of the project-side fusion feature vector to the talent-side fusion feature vector, and perform weighted aggregation processing on the talent-side fusion feature vector according to the attention weight distribution to generate interaction attention matching degree parameters.

[0127] This step aims to simulate the project's "attention" process to talent capabilities. Consider Fuse_Px as the query vector and Fuse_Ty as the key vector. First, Fuse_Px is mapped to the query vector Q = W_Q * Fuse_Px using a trainable weight matrix W_Q. Then, Fuse_Ty is mapped to the key vector K = W_K * Fuse_Ty using another weight matrix W_K. Next, the dot product between Q and K is calculated to obtain a raw attention score representing the project's "attention" to various dimensions of talent characteristics. This score is normalized using the Softmax function to obtain an attention weight distribution, which is a vector with the same dimensions as Fuse_Ty, where each element represents the importance of the corresponding dimension. Finally, this attention weight distribution is multiplied element-wise with Fuse_Ty (i.e., attention-weighted aggregation) to obtain a new, "purified" talent representation vector. This vector is then passed through a simple linear layer or its elements are averaged to obtain a scalar, which serves as the interaction attention matching parameter, denoted as Score_Attn_xy. It reflects the core matching degree elicited by the project requirements focusing on the comprehensive capabilities of talent.

[0128] Step S158: Perform a weighted summation of the structural dimension matching parameter, the text dimension matching parameter, and the interaction attention matching parameter to generate a comprehensive matching score between the research project unit and the professional and technical talent unit. Based on the comprehensive matching score, select target professional and technical talent units that match the research project unit from the professional and technical talent information set, and generate a talent matching recommendation list containing the identifier of the target professional and technical talent unit, the identifier of the research project unit, and the corresponding comprehensive matching score.

[0129] Finally, three hyperparameters λ, μ, and ν are introduced, representing the importance weights of the matching scores in the three dimensions of structure, text, and interaction attention, respectively. The overall matching score Total_Score_xy is calculated as Total_Score_xy = λ * Score_Struct_xy + μ * Score_Text_xy + ν * Score_Attn_xy, where λ + μ + ν = 1. For a given project P_x, its Total_Score_xy is calculated for all talents T_y in the professional and technical talent information set. Then, all talents are sorted from highest to lowest according to Total_Score_xy. The top N talents, or those with scores exceeding the preset admission threshold Thresh_Rec, are selected as the target professional and technical talent units matching the project. Finally, a talent matching recommendation list is generated for project P_x. This list is a data structure containing multiple records, each containing at least a project identifier P_ID, a target talent identifier T_ID, and the corresponding overall matching score Total_Score.

[0130] For example, the method may further include: step S160: obtaining a set of historical project results texts corresponding to each target professional and technical talent unit in the talent matching recommendation list, wherein the set of historical project results texts contains description documents of results from multiple completed scientific research projects.

[0131] After obtaining the initial talent matching recommendation list, to further verify and refine the rationality of the matching, especially from the perspective of technical roadmap evolution and team collaboration, the following supplementary steps are performed. First, based on the talent identifier T_ID of each target talent unit in the recommendation list, retrieve all completed research project achievement description documents from the research project achievement database. Collect these documents to form a historical project achievement text set for that talent, denoted as Docs_T={Doc1, Doc2, ...}.

[0132] Step S161: Perform key technology breakthrough point mining processing on the set of historical project results texts to obtain a set of key technology breakthrough point phrases for each target professional and technical talent unit.

[0133] For each achievement description document in Docs_T, a phrase extraction process similar to steps S121 and S122 is performed. Specifically, each document undergoes structural parsing to identify key sections such as "technical contributions" and "technical innovations," and sentence fragments from these sections are extracted. Then, key technical phrase extraction (including part-of-speech tagging, syntactic analysis, pattern matching, and term frequency-inverse document frequency calculation and filtering) is performed on these sentence fragments. Finally, a set of key technical breakthrough phrases extracted from all the talent's achievement documents, after deduplication and normalization, is obtained, denoted as KP_T_Achieve={KP_A1, KP_A2, ..., KP_AM}.

[0134] Step S162: Input the set of key technology breakthrough phrases of the achievement and the set of key technology phrases of the project unit into the pre-trained technical route fit analysis model. Through the feature cross layer of the technical route fit analysis model, perform phrase-level interactive modeling processing on the set of key technology breakthrough phrases of the achievement and the set of key technology phrases of the project unit to generate a breakthrough point overlap matrix and a breakthrough point complementarity matrix.

[0135] To analyze the deep relationship between past achievements of talent and future project needs, a pre-trained technical roadmap fit analysis model is introduced. The core of this model is to evaluate the relationship between two sets of phrases.

[0136] Step S1621: Retrieve the completed research projects in which each target professional and technical personnel unit participated from the research project results database, obtain the result description document corresponding to each completed research project, and combine all result description documents into the historical project result text set.

[0137] This step is a detailed explanation of step S160, clarifying that the data source is the scientific research project results database, thus ensuring the traceability of the data.

[0138] Step S1622: Perform document structure parsing processing on each achievement description document, identify the technical contribution description section and the technical innovation point description section in the achievement description document, and extract the core technical statement fragments of the achievement from the technical contribution description section and the technical innovation point description section.

[0139] For each achievement description document, a pre-defined keyword library of chapter titles (such as "technical contribution," "main innovation points," "research results," etc.) is used to locate the chapters containing core technical information. Then, continuous sentence text is extracted from these chapters as core technical sentence fragments of the achievement, while non-core content such as background introductions and experimental data is filtered out.

[0140] Step S1623: Extract key technical phrases from the core technical statement fragments of the achievement to obtain a candidate set of key technical phrases for each core technical statement fragment. Perform deduplication and normalization on the candidate set of key technical phrases to generate a set of key technical breakthrough phrases for the target professional and technical talent unit.

[0141] The extracted sentence fragments are processed using the same phrase extraction algorithm as in step S122 to obtain a candidate set. Then, the candidate set is normalized by merging synonymous phrases (such as merging "convolutional neural network" and "CNN") to form the final set of key technical breakthrough phrases, KP_T_Achieve.

[0142] Step S1624: Input each key technology breakthrough phrase in the set of key technology breakthrough phrases into the phrase encoding layer of the technology route fit analysis model, and map each key technology breakthrough phrase into a semantic vector of breakthrough phrase through the phrase encoding layer.

[0143] The technical route fit analysis model includes a phrase encoding layer, which can be the same pre-trained language model as in step S123 (such as a bidirectional encoder representation model). Each phrase KP_Ai in KP_T_Achieve is input into this encoding layer to obtain the corresponding breakthrough point phrase semantic vector V_Ai.

[0144] Step S1625: Input each key technology phrase in the set of key technology phrases for the research project unit into the phrase encoding layer of the technology route fit analysis model, and map each key technology phrase for the research project unit into a semantic vector of project requirement phrases through the phrase encoding layer.

[0145] Similarly, each phrase KP_Pi in the original set of key technical phrases for project breakthroughs KP_P is input into the same phrase encoding layer to obtain the corresponding project requirement phrase semantic vector V_Pi (here V_Pi has the same meaning as V_Pi in step S123, because the same or parameter-shared encoder is used).

[0146] Step S1626: Input the semantic vectors of the achievement breakthrough phrases and the semantic vectors of the project requirement phrases into the feature cross layer of the technical route fit analysis model, and calculate the cosine similarity between each achievement breakthrough phrase semantic vector and each project requirement phrase semantic vector through the feature cross layer to obtain the initial similarity matrix.

[0147] The feature cross layer receives two groups of vectors. For each pair (V_Ai, V_Pj), their cosine similarity is calculated: CosSim_ij = (V_Ai·V_Pj) / (||V_Ai||*||V_Pj||). All the calculation results form an initial similarity matrix Sim_Achieve_Proj with M rows (number of achievement phrases) and K columns (number of project phrases).

[0148] Step S1627: Perform a double-threshold segmentation process on the initial similarity matrix. Set the coincidence degree determination threshold and the complementarity degree determination threshold. Mark the positions of the elements in the initial similarity matrix whose values are greater than the coincidence degree determination threshold as the elements of the coincidence degree matrix, and mark the positions of the elements in the initial similarity matrix whose values are less than the complementarity degree determination threshold and greater than 0 as the candidate elements of the complementarity degree matrix.

[0149] Set two thresholds: the coincidence degree threshold Thresh_Overlap (higher, such as 0.8) and the complementarity degree threshold Thresh_Comp (lower, such as 0.3). Traverse the Sim_Achieve_Proj matrix. If the element CosSim_ij > Thresh_Overlap, it is considered that the achievement phrase KP_Ai and the project phrase KP_Pj describe highly coincident technical points, and record their position (i, j) as a coincidence point. If 0 < CosSim_ij < Thresh_Comp, it is considered that there is only a weak semantic association between these two phrases, which may represent related but different technical directions. Therefore, mark the position (i, j) as a candidate point for the complementarity degree.

[0150] Step S1628: Perform a semantic direction difference analysis process on the semantic vectors of the achievement breakthrough point phrases and the semantic vectors of the project requirement phrases corresponding to the candidate elements of the complementarity degree matrix. Calculate the orthogonal projection component of the semantic vector of the achievement breakthrough point phrase and the semantic vector of the project requirement phrase in the semantic space. If the modulus value of the orthogonal projection component exceeds the preset orthogonality threshold, determine the candidate elements of the complementarity degree matrix as the elements of the complementarity degree matrix.

[0151] For each candidate complementarity point (i, j), a more rigorous test is needed to confirm whether it is truly "complementary" rather than simply noise. Projecting the vector V_Ai onto the direction of V_Pj yields the projection vector Proj = ((V_Ai·V_Pj) / (||V_Pj||^2))*V_Pj. The component orthogonal to V_Pj is then Ortho = V_Ai - Proj. The magnitude ||Ortho|| of the orthogonal component Ortho is calculated. If ||Ortho|| is large (exceeding the preset orthogonality threshold Thresh_Ortho), it indicates that V_Ai contains a large number of unique information directions unrelated to V_Pj, which embodies the meaning of "complementarity." Only candidate complementarity points that meet this condition are ultimately determined as elements of the complementarity matrix.

[0152] Step S1629: Construct the breakthrough point overlap degree matrix based on the overlap degree matrix elements, and construct the breakthrough point complementarity degree matrix based on the complementarity degree matrix elements. The rows of the breakthrough point overlap degree matrix correspond to the key technology breakthrough point phrases of the achievement, the columns of the breakthrough point overlap degree matrix correspond to the key technology phrases of the project, the rows of the breakthrough point complementarity degree matrix correspond to the key technology breakthrough point phrases of the achievement, and the columns of the breakthrough point complementarity degree matrix correspond to the key technology phrases of the project.

[0153] Finally, based on the above labeling results, two sparse matrices are constructed. In the overlap matrix (Overlap_Mat), only the positions (i, j) marked as overlap points have values, which can be set to 1 or the cosine similarity (CosSim_ij). In the complementarity matrix (Comp_Mat), only the positions (i, j) marked as complementarity points have values, which can be set to 1 or the magnitude of the orthogonal components (||Ortho||). These two matrices quantify the overlap and complementarity between talent achievements and project needs.

[0154] Step S163: Perform matrix norm calculation on the breakthrough point overlap matrix to obtain the comprehensive overlap score parameter. At the same time, perform matrix singular value decomposition on the breakthrough point complementarity matrix to extract the principal singular value vector of the breakthrough point complementarity matrix as the complementary direction feature vector.

[0155] To compress the matrix information into a usable score, the Frobenius norm of Overlap_Mat is calculated, which is the square root of the sum of the squares of all elements in the matrix, yielding the comprehensive overlap score parameter Score_Overlap. A larger value indicates greater overlap between the deliverables and the requirements. Singular value decomposition is performed on Comp_Mat, resulting in Comp_Mat = U * Σ * V^T. Here, Σ is a diagonal matrix, and the elements on the diagonal are singular values, arranged in descending order. The first column (or the first few columns) of the left singular vector U corresponding to the largest one or several singular values ​​is extracted as the complementary direction feature vector Vec_Comp_Direction, representing the most significant complementary technical directions between the deliverables and the requirements.

[0156] Step S164: Based on the complementary directional feature vectors, perform path matching retrieval processing in the preset technology development path map, identify the sequential relationship between the set of key technology breakthrough phrases of the achievement and the set of key technology phrases of the project in the technological evolution context, and calculate the continuity score of the technological evolution path based on the coherence and completeness of the identified successive paths.

[0157] This section introduces a pre-defined technology development path graph, a directed graph where nodes are key technology phrases and edges represent technology evolution relationships (e.g., the evolution from "technology A" to "technology B"). The complementary direction feature vector Vec_Comp_Direction is used as the query to retrieve the most relevant technology node in the graph. Simultaneously, the project's KP_P set is also mapped into the graph. Then, directed paths are searched in the graph from nodes related to achievement KP_T_Achieve (or their complementary direction nodes) to nodes related to project KP_P. Path continuity is measured by the average transition probability between nodes along the path, and completeness is measured by the ratio of path length to the known length of complete paths in the graph. These metrics are combined to calculate a technology evolution path continuity score, Score_Path.

[0158] Step S165: Based on the preset weight configuration, perform weighted fusion processing on the overlap degree comprehensive score parameter and the technology evolution path continuity score to generate a technology route depth fit parameter.

[0159] We define weights ε and ζ, and calculate the depth-fit-score parameter of the technology roadmap: Depth_Fit_Score = ε * Score_Overlap + ζ * Score_Path. This parameter comprehensively considers the degree of direct overlap between existing talent achievements and project needs, as well as their connection in the context of technological development.

[0160] Step S166: Reorder the target professional and technical personnel units in the talent matching recommendation list according to the technical route depth matching parameter to generate a preliminary sorted list arranged in descending order of technical route depth matching.

[0161] For each talent in the preliminary recommendation list generated in step S158, re-sort them according to their corresponding Depth_Fit_Score, with those scoring higher first, to generate a new preliminary sorted list Sorted_List.

[0162] Step S167: Input the preliminary sorting list into the talent collaboration network analysis module, extract the historical collaboration relationship edge set between each target professional and technical talent unit in the talent matching recommendation list, and construct the recommended talent collaboration subgraph based on the historical collaboration relationship edge set.

[0163] Considering the practical needs of team building, this paper analyzes the collaborative relationships among talents in the recommended list. From historical research collaboration records, it extracts whether any of the recommended talents (from Sorted_List) have collaborated before. If talents T_A and T_B have participated in the same project together, a collaborative relationship edge is constructed between them. These edges, along with the talent nodes they connect, together form a recommended talent collaboration subgraph, Subgraph_Collab.

[0164] Step S168: Perform graph clustering analysis on the recommended talent collaboration subgraph to identify the dense collaboration community structure in the recommended talent collaboration subgraph, and calculate the community centrality score of each target professional and technical talent unit in its respective dense collaboration community structure based on the modularity parameter of the dense collaboration community structure.

[0165] Apply graph clustering algorithms, such as Louvain's algorithm or label propagation, to the Subgraph Collab to divide closely connected nodes into communities (i.e., potential existing collaborative teams). Then, calculate the centrality of each node within its community. For example, degree centrality (the number of edges a node connects to within a community) or betweenness centrality (the number of times a node acts as a bridge) can be used. This centrality score is denoted as Centrality_Score_T, which reflects the degree to which talent is central to the existing collaborative network.

[0166] Step S169: Obtain the team formation size constraint parameters of the scientific research project unit, and select the core collaborative talent unit set and the peripheral collaborative talent unit set from the recommended talent collaboration subgraph based on the team formation size constraint parameters and the community centrality score.

[0167] Assume the team size constraint for project P_x is Team_Size. Talent selection begins from the head of the Sorted_List. During the selection process, refer to Centrality_Score_T. Prioritize selecting "core" talent with high centrality scores from different communities (e.g., selecting 1-2 people with the highest centrality from each community) to form the core collaborative talent unit set Core_Team. After the core team is determined, if additional personnel are needed, select from the remaining talent; these individuals may come from peripheral members of the same community, forming the peripheral collaborative talent unit set Periphery_Team. The aim is to promote team diversity and utilize existing collaborative foundations while ensuring the quality of individual talent.

[0168] Step S1610: Perform redundancy analysis on the core collaborative talent unit set, detect whether there are talent unit pairs in the core collaborative talent unit set whose research direction overlap exceeds a preset overlap threshold, perform deduplication on talent unit pairs whose research direction overlap exceeds the preset overlap threshold, merge the core collaborative talent unit set and the peripheral collaborative talent unit set after deduplication, generate a talent combination recommendation list, and label each target professional and technical talent unit in the talent combination recommendation list with its community centrality score and its corresponding technical route depth fit parameter in the preliminary ranking list.

[0169] Finally, the Core_Team is optimized. For any two individuals T_A and T_B in the set, the overlap of their respective sets of key technological breakthrough phrases is calculated, for example, using the Jaccard similarity coefficient: Jaccard = |KP_T_Achieve_A∩KP_T_Achieve_B| / |KP_T_Achieve_A∪KP_T_Achieve_B|. If Jaccard exceeds a preset overlap threshold Thresh_Jaccard, it indicates that the two individuals' research fields highly overlap, potentially indicating redundancy. In this case, based on their technical route depth fit parameter Depth_Fit_Score, the individual with the higher score is retained, and the other is removed from the Core_Team (or downgraded to the Periphery_Team). After deduplication, the optimized Core_Team and Periphery_Team are merged to generate the final talent combination recommendation list. In addition to its identifier T_ID, each talent entry in the list must also include its community centrality score (Centrality_Score_T) and its corresponding technology roadmap depth fit parameter (Depth_Fit_Score) so that the project leader can make the final team building decision based on this information.

[0170] In some embodiments, the knowledge graph-based research direction and professional talent matching system for performing the above methods can be any electronic device with data computing, processing, and storage functions. This knowledge graph-based research direction and professional talent matching system can be used to implement the methods provided in the above embodiments, or the processing methods of text processing models.

[0171] Typically, a knowledge graph-based research direction and professional talent matching system includes a processor and memory. The processor may include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor can be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and coprocessors. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which handles computational operations related to machine learning.

[0172] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory are used to store a computer program configured to be executed by one or more processors to implement the above-described method, or the processing method of the text processing model.

[0173] In an illustrative embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored therein. When executed by a processor of a computer device, the computer program implements the above-described method, or the processing method of a text processing model. Optionally, the above-described computer-readable storage medium may be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (CompactDisc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.

[0174] This application provides a computer program product, which includes computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the content recommendation method provided in this application.

[0175] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the content recommendation method provided in this application.

[0176] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or CD-ROM, etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0177] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0178] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0179] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0180] Finally, it should be noted that the above-disclosed embodiments are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for matching research directions with professional and technical personnel based on knowledge graphs, characterized in that, The method includes: Acquire a set of scientific research project information and a set of professional and technical personnel information, wherein the set of scientific research project information includes multiple scientific research project units and the set of professional and technical personnel information includes multiple professional and technical personnel units; A scientific research knowledge graph network is constructed based on the project tackling direction description text in the scientific research project information set and the professional and technical personnel ability description text in the professional and technical personnel information set. The scientific research knowledge graph network includes a set of scientific research project nodes, a set of professional and technical personnel nodes, and a set of scientific research tackling association edges connecting the scientific research project nodes and the professional and technical personnel nodes. The scientific research knowledge graph network is subjected to graph structure feature extraction processing to generate the project node structure embedding feature set of the scientific research project node in the scientific research knowledge graph network and the talent node structure embedding feature set of the professional and technical talent node in the scientific research knowledge graph network. The pre-trained project requirement feature extractor is invoked to perform semantic encoding processing on the text describing the project's key research direction of the scientific research project unit, and the resulting text semantic feature vector is used as the semantic feature of the project requirement text. At the same time, the pre-trained talent ability feature extractor is invoked to perform semantic encoding processing on the text describing the professional and technical talent ability of the professional and technical talent unit, and the resulting text semantic feature vector is used as the semantic feature of the talent ability text. Cross-modal feature fusion and matching processing is performed on the project node structure embedding feature set, the talent node structure embedding feature set, the project requirement text semantic features, and the talent ability text semantic features to generate a multi-dimensional matching degree measurement parameter set between the scientific research project unit and the professional and technical talent unit. Based on the multi-dimensional matching degree measurement parameter set, target professional and technical talent units that match the scientific research project unit are selected from the professional and technical talent information set, and a talent matching recommendation list containing the correspondence between the target professional and technical talent unit identifier and the scientific research project unit identifier is generated. The construction of a scientific research knowledge graph network based on the project research direction description text in the scientific research project information set and the professional and technical personnel capability description text in the professional and technical personnel information set includes: The key technology phrases for the key technologies of the research project unit are extracted from the description text of the project's research direction. The text describing the professional and technical personnel's abilities in the professional and technical personnel unit is processed by extracting key technical phrases that the personnel are good at, so as to obtain a set of key technical phrases that the personnel are good at in the professional and technical personnel unit. The set of key technology phrases for tackling key problems in the project and the set of key technology phrases for talents are mapped into a semantic space of phrases, so that the set of key technology phrases for tackling key problems in the project is mapped into a semantic vector of key technology phrases for tackling key problems in the project, and the set of key technology phrases for talents are mapped into a semantic vector of key technology phrases for talents. Based on the semantic vector space similarity calculation results between the semantic vector of the key technology phrases in the project and the semantic vector of the key technology phrases in which the talent is good, an initial scientific research connection edge is constructed between the scientific research project node and the professional and technical talent node. Obtain a set of historical research collaboration records, which includes the historical collaboration correspondence between project identifiers of multiple completed research projects and talent identifiers of professional and technical personnel who participated in the completed research projects. Based on the historical cooperation correspondence, historical cooperation association edges are constructed between scientific research project nodes and professional and technical personnel nodes in the scientific research knowledge graph network. By fusing the initial scientific research breakthrough association edges with the historical cooperation association edges, a set of scientific research breakthrough association edges in the scientific research knowledge graph network is generated. Each scientific research breakthrough association edge in the set of scientific research breakthrough association edges contains a weighted fusion result of the initial association strength weight and the historical cooperation strength weight as a comprehensive association strength parameter. The scientific research knowledge graph network is optimized by performing network topology optimization. Based on the comprehensive association strength parameter, the scientific research-related edges that are below the preset strength threshold are pruned to obtain the optimized scientific research knowledge graph network.

2. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The process of extracting key technical phrases from the description text of the research project unit's research direction yields a set of key technical phrases for the research project unit, including: The text describing the research project's key research direction is processed by text sentence segmentation to obtain the sequence of project key research description sentences contained in the text describing the research project's key research direction. Part-of-speech tagging is performed on each project tackling description statement in the project tackling description statement sequence to identify noun phrase units and verb phrase units in the project tackling description statements; Based on a preset key technology phrase word formation pattern library, the noun phrase unit and the verb phrase unit are subjected to phrase combination pattern matching processing to generate candidate key technology phrases that conform to the key technology phrase word formation pattern library; The frequency of occurrence of the candidate key technology phrases in the description text of the project's key research direction is statistically processed to obtain the project's frequency of occurrence parameter for each candidate key technology phrase. The inverse document frequency parameter of the candidate key technology phrases at the set level of all scientific research project units is calculated to obtain the global scarcity weight parameter of each candidate key technology phrase. Based on the product of the frequency parameter and the global scarcity weight parameter within the project, the importance score of each candidate key technology phrase is calculated. According to the preset importance score threshold or the preset selection number, the K candidate key technology phrases with importance scores exceeding the importance score threshold or ranking first in importance score are selected to form the set of key technology phrases for the project. Each key technology phrase in the set of key technology phrases for project breakthroughs is processed for phrase standardization, and key technology phrases with different expressions having the same semantics are merged into a unified standard key technology phrase identifier.

3. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The process of extracting key technical phrases from the description text of the professional and technical personnel's capabilities in the professional and technical personnel unit yields a set of key technical phrases for the professional and technical personnel unit, including: The professional and technical personnel ability description text of the professional and technical personnel unit is processed by text segmentation to obtain the talent ability description paragraph unit contained in the professional and technical personnel ability description text; Sentence boundary detection processing is performed on the talent ability description paragraph unit to segment each talent ability description paragraph unit into a talent ability description sentence sequence; Syntactic dependency analysis is performed on each statement in the sequence of statements describing talent capabilities to identify the core verb units and core noun units that have a dependency relationship with the core verb units in the statements; Based on the dependency relationship type between the core verb unit and the core noun unit, the core verb unit and the core noun unit are combined into candidate technical phrases. The distribution density of the candidate technical skill phrases in the text describing the professional and technical personnel's abilities is statistically processed to obtain the paragraph coverage rate parameter of each candidate technical skill phrase in the text describing the professional and technical personnel's abilities. Based on the paragraph coverage parameter, the candidate proficiency technical phrases are filtered, and the candidate proficiency technical phrases whose paragraph coverage parameter exceeds a preset coverage threshold are retained as the set of key technical phrases proficient in the talent.

4. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The initial scientific research connection edge between the scientific research project node and the professional and technical talent node is constructed based on the semantic vector space similarity calculation result between the semantic vectors of the key technology phrases of the project and the key technology phrases of the talent's expertise, including: The semantic vectors of the key technology phrases for tackling key problems in the project and the semantic vectors of the key technology phrases for which the talent is good are processed by a vector dot product operation to obtain a phrase-level semantic similarity matrix between the key technology phrases for tackling key problems in the project and the key technology phrases for which the talent is good. The phrase-level semantic similarity matrix is ​​subjected to row-direction max pooling to extract the maximum similarity value between the key technology phrases for each project and the key technology phrases for all talents, which is used as the maximum similarity vector on the project side. The phrase-level semantic similarity matrix is ​​subjected to column-direction max pooling to extract the maximum similarity value between each talent's key technical phrases and all project key technical phrases, which is used as the talent-side maximum similarity vector. The mean value of the maximum similarity vector on the project side is calculated to obtain the average semantic similarity parameter between the scientific research project unit and the professional and technical personnel unit in the project-to-talent direction. The mean value of the maximum similarity vector on the talent side is calculated to obtain the average semantic similarity parameter from talent to project direction between the scientific research project unit and the professional and technical talent unit. Based on preset weighting coefficients, the average semantic similarity parameter from project to talent and the average semantic similarity parameter from talent to project are weighted and fused to generate a comprehensive semantic matching parameter between the scientific research project unit and the professional and technical talent unit. Based on the comprehensive semantic matching degree parameter, an initial scientific research breakthrough association edge carrying the comprehensive semantic matching degree parameter is constructed between the scientific research project node and the professional and technical personnel node.

5. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The step of extracting graph structure features from the scientific research knowledge graph network to generate a set of project node structure embedding features for the scientific research project nodes in the scientific research knowledge graph network and a set of talent node structure embedding features for the professional and technical personnel nodes in the scientific research knowledge graph network includes: The scientific research knowledge graph network is subjected to node neighborhood sampling processing. For each scientific research project node, a set of first-order neighbor nodes that are directly connected to the scientific research project node through the scientific research breakthrough association edge are collected. For each professional and technical talent node in the first-order neighbor node set, recursive sampling is performed, and other scientific research project nodes that are directly connected to each professional and technical talent node through scientific research breakthroughs are collected as the second-order neighbor node set. Construct the local subgraph structure of the scientific research project node in the scientific research knowledge graph network based on the first-order neighbor node set and the second-order neighbor node set; Graph convolution propagation processing is performed on the local subgraph structure of the research project node to aggregate the feature information of neighboring nodes in the local subgraph structure onto the research project node, thereby generating the project node structure embedding feature of the research project node. The neighborhood sampling and local subgraph structure construction processes are repeated for each professional and technical talent node in the scientific research knowledge graph network. Graph convolution propagation processing is then performed on the local subgraph structure of each professional and technical talent node to generate talent node structure embedding features for each professional and technical talent node. The project node structure embedding features of all scientific research project nodes are combined into the project node structure embedding feature set, and the talent node structure embedding features of all professional and technical talent nodes are combined into the talent node structure embedding feature set.

6. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The process involves calling a pre-trained project requirement feature extractor to perform semantic encoding on the text describing the project's research direction for each research project unit, obtaining a text semantic feature vector for that research project unit as the semantic features of the project requirement text, including: The text describing the research project's key research direction is normalized in length, and the text is truncated or padded to a preset fixed text length. The normalized length description text of the project's key points is input into the word embedding layer of the pre-trained project requirement feature extractor, and each lexical unit in the description text of the project's key points is mapped to a sequence of lexical embedding vectors. The word embedding vector sequence is input into the multi-head self-attention layer of the pre-trained project requirement feature extractor to model the semantic dependency relationship between word units in the word embedding vector sequence and generate a context-aware word vector sequence carrying contextual semantic information. The context-aware word vector sequence is input into the pooling layer of the pre-trained project requirement feature extractor, and average pooling is performed on all word vectors in the context-aware word vector sequence to obtain sentence-level semantic vectors of fixed dimensions. The sentence-level semantic vector is input into the fully connected mapping layer of the pre-trained project requirement feature extractor. The sentence-level semantic vector is then subjected to dimensionality transformation, and the text semantic feature vector of the scientific research project unit is output as the project requirement text semantic feature.

7. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The process involves calling a pre-trained talent capability feature extractor to perform semantic encoding on the text describing the professional and technical talent capabilities of the professional and technical talent unit, obtaining the text semantic feature vector of the professional and technical talent unit as the semantic features of the talent capability text, including: The text describing the professional and technical personnel's abilities in the professional and technical personnel unit is divided into paragraphs, and the text describing the professional and technical personnel's abilities is divided into multiple descriptive paragraph units. For each descriptive paragraph unit, text length normalization and word embedding mapping are performed to generate a paragraph word embedding vector sequence for each descriptive paragraph unit; The paragraph word embedding vector sequence of each paragraph description unit is input into the bidirectional long short-term memory network layer of the pre-trained talent ability feature extractor. Forward and reverse sequence encoding processing is performed on each paragraph word embedding vector sequence to obtain the forward hidden state sequence and the reverse hidden state sequence of each paragraph description unit. The hidden state at the last time step of the forward hidden state sequence of each paragraph unit is concatenated with the hidden state at the first time step of the reverse hidden state sequence to obtain the paragraph semantic vector of each paragraph unit. The paragraph semantic vectors of all descriptive paragraph units are weighted and summed, and different weight coefficients are assigned to each descriptive paragraph unit according to its position order in the professional and technical personnel capability description text, thereby generating a document-level semantic vector that integrates multi-paragraph information. The document-level semantic vector is input into the output layer of the pre-trained talent ability feature extractor, and the text semantic feature vector of the professional and technical talent unit is output as the talent ability text semantic feature.

8. The method for matching research directions with professional and technical personnel based on knowledge graphs according to claim 1, characterized in that, The cross-modal feature fusion and matching process is performed on the project node structure embedding feature set, the talent node structure embedding feature set, the project requirement text semantic features, and the talent ability text semantic features to generate a multi-dimensional matching degree measurement parameter set between the scientific research project unit and the professional and technical talent unit, including: Each project node structure embedding feature in the project node structure embedding feature set and each talent node structure embedding feature in the talent node structure embedding feature set are concatenated into a vector to generate a project-talent structure feature pair vector. The semantic features of the project requirements text and the semantic features of the talent capabilities text are concatenated into vectors to generate a project-talent text feature pair vector. The project-talent structure feature pair vector is input into the first matching degree calculation branch, and the project-talent structure feature pair vector is subjected to nonlinear transformation processing through the first multilayer perceptron network to output the structural dimension matching degree parameter. The project-talent text feature pair vector is input into the second matching degree calculation branch, and the project-talent text feature pair vector is subjected to nonlinear transformation processing through the second multilayer perceptron network to output the text dimension matching degree parameter. The project node structure embedding features and the project requirement text semantic features are subjected to feature cross-processing to generate a project-side fused feature vector; The talent node structure embedding features and the talent ability text semantic features are subjected to feature cross-processing to generate a talent-side fusion feature vector; Attention interaction processing is performed on the project-side fusion feature vector and the talent-side fusion feature vector. The attention weight distribution of the project-side fusion feature vector to the talent-side fusion feature vector is calculated. The talent-side fusion feature vector is then weighted and aggregated according to the attention weight distribution to generate interaction attention matching degree parameters. The structural dimension matching parameter, the text dimension matching parameter, and the interaction attention matching parameter are weighted and summed to generate a comprehensive matching score between the research project unit and the professional and technical talent unit. Based on the comprehensive matching score, target professional and technical talent units that match the research project unit are selected from the professional and technical talent information set, and a talent matching recommendation list containing the identifier of the target professional and technical talent unit, the identifier of the research project unit, and the corresponding comprehensive matching score is generated.

9. A knowledge graph-based system for matching research directions with professional and technical personnel, characterized in that, The invention includes a processor and a computer-readable storage medium storing machine-executable instructions, which, when executed by a computer, implement the knowledge graph-based method for matching research directions with professional and technical personnel as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Scientific research talent and research and development task bidirectional matching method and system based on graph representation

    CN118735217A

  • Scientific researcher and scientific research project matching recommendation method and system

    CN120542828A