Knowledge graph-based professional influence evaluation and dynamic deduction method and system, and storage medium

By constructing a knowledge graph-based multi-source heterogeneous data acquisition and fusion system, the problems of one-sided and static evaluation results in existing technologies have been solved. This system enables professionals to perform holistic, objective, and dynamic predictive analysis, and provides fair benchmarking and forward-looking decision support across organizations and regions.

CN120806741AActive Publication Date: 2025-10-17HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN) +1

Patent Information

Application Number
CN202511285632.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-10-17
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing technologies have problems when evaluating professionals, such as information islands, one-sided evaluation dimensions, lack of unified standards and staticness, making it impossible to conduct dynamic predictive analysis.

Method used

A knowledge graph-based multi-source heterogeneous data acquisition and fusion system is constructed. Entity relationships are extracted through pre-trained language models and specific prompting engineering, the absolute score of global influence is calculated, and context-aware analysis and perturbation simulation are performed to achieve dynamic inference.

Benefits of technology

It achieves holistic, objective, and comparable evaluation results, provides fair benchmarking capabilities across organizations and regions, and possesses context-aware and forward-looking decision support capabilities, breaking through the paradigm limitations of static evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806741A_ABST
    Figure CN120806741A_ABST
Patent Text Reader

Abstract

The invention provides a professional influence evaluation and dynamic deduction method and system based on a knowledge graph and a storage medium. The method comprises the steps of 1, constructing a structured knowledge set representing entities and relationships in a professional field; 2, calculating a global influence absolute score of a target entity based on the structured knowledge set in the step 1; step 3, based on the global influence absolute score in the step 2, performing context-aware context influence analysis; and 4, performing prospective influence deduction based on disturbance simulation of the structured knowledge set in the step 1. The method has the beneficial effects that 1, the holography and objectivity of an evaluation basis are realized, and the credibility and fairness of an evaluation result are fundamentally improved; 2, a uniform and data-driven evaluation scale is established, and cross-organization and cross-region fair benchmarking is realized for the first time; and 3, a context-aware multi-level analysis capability is provided, so that the evaluation has unprecedented depth and fineness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer data processing and artificial intelligence, and particularly relates to a professional influence evaluation and dynamic deduction method and system based on a knowledge graph and a storage medium. BACKGROUND

[0002] In the era of knowledge economy, scientific, objective and comprehensive evaluation of professional talents (such as doctors, scientific researchers, engineers and lawyers) has become the cornerstone of talent management, resource allocation and strategic decision-making for various organizations. However, the current technical practice generally faces a core dilemma: the contradiction between the extreme dispersion of information required for evaluation and the serious one-dimensionality of evaluation method dimensions.

[0003] On the one hand, key data such as professional information, academic achievements, social reputation and network position that can comprehensively reflect the comprehensive ability of an individual are widely scattered in a large number of heterogeneous and disconnected data sources such as institutional websites, industry vertical platforms, academic databases, news media and social networks, forming an insurmountable "information island".

[0004] On the other hand, the existing evaluation system mostly relies on simple weighting of a few easily accessible explicit indicators (such as title, length of service and platform internal score), completely ignoring the consideration of the depth of individual skill combination, the pivotal role in the collaboration network and the actual reputation in the industry and other deeper implicit capital. This contradiction leads to the subjectivity of the existing evaluation results, poor comparability, and complete lack of context awareness and multi-level analysis capabilities.

[0005] More importantly, the existing technology is essentially a pure "backward-looking" static summary, and it is completely powerless for strategic problems facing the future, dynamics and prediction, such as "how much quantitative impact will the core talent flow have on the organization?" or "how to form a team with maximum efficiency?". Therefore, there is an urgent need in the field for a new technical solution to fundamentally break through the above technical bottlenecks and upgrade the evaluation of professional personnel from static information display to dynamic and quantifiable decision support.

[0006] The most similar existing implementation scheme to the present application can be mainly divided into two categories: the first category is the online evaluation and ranking system widely used in the business field; the second category is the scientific metrology evaluation system long-term used in the academic field.

[0007] The first existing technical solution: online evaluation and ranking system based on internal platform data; the typical representative of this type of technology is various online medical, legal service or technical Q&A platforms. The technical solution can be decomposed into the following core modules in architecture: 1. Data collection and preprocessing module: Obtain data through two ways, active registration submission (such as title, resume) and passive record of user behavior in the platform (such as consultation volume, score, text evaluation); 2. Index quantification and feature engineering module: Convert the collected information into calculable numerical features. For example, map the text title "chief physician" to a numerical score through a pre-set mapping table, or aggregate and count user behavior data; 3. Weighted scoring model module: Use a linear weighted model set by the platform operator, which can be expressed mathematically as: Where represents the i-th numerical feature, represents the pre-set weight of the feature; 4. Sorting and display module: Sort the professionals in the platform in descending order according to the final calculated Score value and display on the user interface; Second type of existing technical solution: scientific metrology evaluation system based on citation network; this type of technology is mainly applied in the academic field, represented by "H index", "journal impact factor (JIF)", etc. The core logic of its technical solution is: 1. Data collection and network construction: Collect structured data such as papers, authors, journals, and citations from professional academic databases (such as Web of Science, Scopus); 2. Index calculation module: Build a citation network based on the collected data and run specific, fixed mathematical formulas or statistical indicators (such as H index algorithm) on this network to quantify the academic influence of scholars or institutions; The above two types of existing technical solutions, although they meet the basic evaluation needs to some extent, but from a deeper technical perspective, they all have the following fundamental and difficult to overcome defects: 1. One-sidedness of evaluation dimension and limitation of data basis: The data source of the first type of technology is limited to the platform, which is a typical "information island", and cannot reflect the individual's true value outside the platform. The data source of the second type of technology is limited to academic output, which is an "evaluation in the study", and cannot measure the individual's professional practice ability and social network value. Both of them fail to integrate the individual's multi-dimensional ability capital, leading to one-sided and distorted evaluation results; 2. One-sidedness of evaluation dimension and limitation of data basis: The data source of the first type of technology is limited to the platform, which is a typical "information island", and cannot reflect the individual's true value outside the platform. The data source of the second type of technology is limited to academic output, which is an "evaluation in the study", and cannot measure the individual's professional practice ability and social network value. Both of them fail to integrate the individual's multi-dimensional ability capital, leading to one-sided and distorted evaluation results; 3. Lack of unified evaluation scale and comparability: Different platforms have different scoring systems, and the results are completely incomparable across platforms. Even the H-index is difficult to compare directly due to differences in subject areas.

[0008] 4. Static and retrospective nature of the system paradigm (most core defect): Both existing technologies are essentially static data snapshots of individual or organizational past achievements. Their system architecture and algorithm design do not have the technical means to respond to virtual events such as core talent flow, simulate system structure changes, and predict the potential impact of such events on the future in a forward-looking and quantitative manner. This makes such technologies completely useless when faced with future-oriented, dynamic, and predictive strategic problems, limiting their application value to the superficial level of "historical summary". SUMMARY

[0009] To solve the problems in the prior art, the present application provides a professional influence evaluation and dynamic deduction method based on a knowledge graph, comprising: Step one: constructing a structured knowledge set representing entities and relationships in the professional field; Step two: calculating the global influence absolute score of the target entity based on the structured knowledge set of step one; Step three: context-aware situational influence analysis based on the global influence absolute score of step two; Step four: forward-looking influence deduction based on perturbation simulation of the structured knowledge set of step one.

[0010] As a further improvement of the present application, step one comprises: Step S1, multi-source heterogeneous data acquisition: obtaining structured, semi-structured, and unstructured data about professionals and their related activities from multiple predefined online data sources; Step S2, knowledge extraction: using a unified extraction framework based on a pre-trained language model and combined with specific prompt engineering, the framework identifies entities and extracts multi-dimensional relationship triples between them from long text through a three-stage processing pipeline of abstraction-recognition-refinement; Step S3, knowledge fusion and verification: for entities from different data sources pointing to the same real-world object, use a multi-level conflict resolution framework based on signal priority for entity alignment and fact verification; Step S4, knowledge storage: load the cleaned and fused structured knowledge into a data storage system that can handle complex relationships between processed entities.

[0011] The beneficial effects of the present application are: 1. The holographic and objective evaluation basis is realized, and the credibility and fairness of the evaluation results are fundamentally improved; the specific embodiment is: the data basis of the prior art is "isolated" and "one-sided", or limited to the platform or limited to a single academic dimension. The structured knowledge aggregation unit of the present application constructs a holographic knowledge graph containing individual academic, professional, social network and other multi-dimensional ability capital by automatically extracting and fusing information from the whole network multi-source heterogeneous data. The evaluation process is completely driven by data and algorithm, especially the "signal priority" multi-level conflict resolution framework, which ensures the accuracy of the facts and eliminates artificial intervention and subjective bias; this makes the evaluation conclusion of the present application no longer based on one-sided information, but on the accurate portrait based on panoramic data, and its objectivity, comprehensiveness and credibility are completely beyond the reach of the prior art; 2. A unified, data-driven evaluation scale is creatively established, and fair benchmarking across organizations and regions is first realized; the specific embodiment is: the scoring system of the prior art relies on subjective weighting or uses a rigid formula, resulting in non-transparent and non-comparable results. The individual ability quantification unit of the present application creatively calculates an "absolute score of global influence" (which is not limited by any external context) through a pre-trained comprehensive evaluation model (preferably GNN) that can capture complex nonlinear relationships. This score provides a unified "ability measure" for every professional in the world based on their intrinsic comprehensive ability. This makes it possible to make fair and objective horizontal comparisons among professionals in different institutions, different regions, and even different countries, solving the fundamental problem of chaotic evaluation standards and the inability to benchmark in the industry; 3. Provide context-aware multi-level analysis capability, making the evaluation have unprecedented depth and precision; the specific embodiment is: the prior art can only provide a general and single-level ranking. The context influence analysis unit of the present application can provide a multi-level ranking based on the unified "ability measure", and the ranking results are not only accurate and fair, but also have a clear and intuitive interpretation, which can be directly used as a reference for the improvement of individual ability and the optimization of organizational structure. ) that is not limited by any external context. This score provides a unified "ability measure" for every professional in the world based on their intrinsic comprehensive ability. This makes it possible to make fair and objective horizontal comparisons among professionals in different institutions, different regions, and even different countries, solving the fundamental problem of chaotic evaluation standards and the inability to benchmark in the industry; 3. Provide context-aware multi-level analysis capability, making the evaluation have unprecedented depth and precision; the specific embodiment is: the prior art can only provide a general and single-level ranking. The context influence analysis unit of the present application can provide a multi-level ranking based on the unified "ability measure", and the ranking results are not only accurate and fair, but also have a clear and intuitive interpretation, which can be directly used as a reference for the improvement of individual ability and the optimization of organizational structure. The ruler dynamically penetrates the analysis according to the arbitrary range specified by the user (from the microscopic department to the macroscopic country, or even a specific technical field), and uses the "pre-computation and caching" strategy to ensure the instantaneous response of the query; this capability makes the evaluation no longer a one-size-fits-all, but can accurately answer "the person is absolutely core in the hospital, but what is his level in the city?" or "in the 'AI pharmaceutical' sub-track, what is his industry status?" Such complex questions with depth and context provide users with more insightful, multi-dimensional decision-making basis; 4. The paradigm breakthrough from static evaluation to dynamic deduction is realized, and the system is endowed with forward-looking decision support capability (core beneficial effect); The paradigm breakthrough from static evaluation to dynamic deduction is realized, and the system is endowed with forward-looking decision support capability (core beneficial effect); this makes the present application not limited to "evaluating the present situation", but can "predict the future". It changes the evaluation system from a static information display tool to a dynamic "sand table deduction" engine that can support talent introduction risk assessment, team optimization, organizational structure adjustment and other major strategic decisions. This leap from "descriptive analysis" to "predictive and instructive analysis" represents a generational progress in technology application, providing managers with unprecedented data-driven decision-making tools. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a whole function module architecture diagram of the system of the present application; Figure 2 is a structured knowledge aggregation unit workflow schematic diagram of the present application; Figure 3 is a individual ability quantification unit calculation flow schematic diagram of the present application; Figure 4 is a situational influence analysis unit logic schematic diagram of the present application; Figure 5 is a system disturbance deduction unit core flow schematic diagram of the present application; Figure 6 is a simplified medical field knowledge graph Schema example diagram of the present application; DETAILED DESCRIPTION

[0013] In view of the technical problem in the background art that the evaluation system can only perform static and retrospective summary on historical data, and cannot make forward-looking and quantifiable prediction on future events, thereby limiting its application value, the purpose of the present application is to provide a professional influence evaluation and dynamic deduction system and method, which aims to break through the paradigm limitation of the prior art that can only perform static evaluation, and endow the evaluation system with forward-looking and quantifiable dynamic deduction and decision support capability.

[0014] 1. The present application aims to provide a technical solution that can automatically collect, extract and fuse information from multi-source heterogeneous data, and structurally represent it as a unified, machine-readable knowledge network, thereby laying a comprehensive and reliable data foundation for subsequent objective evaluation and dynamic deduction; 2. In view of the technical problem of inaccurate and non-objective evaluation results caused by the simple linear or rigid evaluation model in the prior art, the present application aims to provide a technical solution that can combine high-dimensional and heterogeneous feature vectors into an absolute influence power index that is not limited by external context and can be globally compared based on the aforementioned knowledge network using a comprehensive evaluation model (such as the preferred graph neural network model) that can effectively learn network structure information, thereby establishing a unified and objective evaluation scale; 3. In view of the technical problem of inaccurate and non-objective evaluation results caused by the simple linear or rigid evaluation model in the prior art, the present application aims to provide a technical solution that can combine high-dimensional and heterogeneous feature vectors into an absolute influence power index that is not limited by external context and can be globally compared based on the aforementioned knowledge network using a comprehensive evaluation model (such as the preferred graph neural network model) that can effectively learn network structure information, thereby establishing a unified and objective evaluation scale; 4. In view of the core technical defect that the prior art is essentially a summary of historical data and is completely unable to make forward-looking and quantitative predictions, the present application aims to provide a core technical solution based on structured knowledge set perturbation simulation, which can forwardly and quantitatively deduce the specific impact of a key event on related organizations in multiple macro dimensions by simulating the structural or parameter changes caused by the event on a virtual copy of the knowledge network, thereby transforming the evaluation system from a static information display tool into a dynamic deduction engine that can support major strategic decisions.

[0015] The present application discloses a professional influence evaluation and dynamic deduction method based on a knowledge graph, comprising: Step one: build a structured knowledge set representing entities and relationships in the professional field; this step aims to automatically convert multi-source, heterogeneous raw data into a high-quality structured knowledge model that provides data support for subsequent calculations. In a preferred embodiment, the structured knowledge set is embodied as a domain knowledge graph, and the construction process includes: Step S1, multi-source heterogeneous data collection: through a configured web crawler or API interface, unstructured or semi-structured data about professionals and their related activities is obtained from multiple predefined online data sources (including but not limited to official websites, academic databases, patent databases, news media, etc.).

[0016] Step S2, knowledge extraction: a unified extraction framework based on a pre-trained language model (preferably a large language model LLM) combined with specific prompt engineering is adopted. This framework uses a three-stage processing pipeline of summarization-recognition-refinement to accurately identify entities (such as [professional], [organization], [academic achievement], etc.) that conform to predefined schema from long text and extract multi-dimensional relationship triples (for example, <head entity, relationship type, tail entity>) between them.

[0017] Step S3, knowledge fusion and verification: for entities from different data sources pointing to the same real-world object, a multi-level conflict resolution framework based on "signal priority" is used for entity alignment and fact verification. This framework comprehensively utilizes metadata (such as timestamps, data source authority), internal topological consistency of the graph, and context reasoning based on LLM to ensure the uniqueness, accuracy and completeness of the final knowledge. Step S4, knowledge storage: the structured knowledge after cleaning and fusion is loaded into a data storage system (such as a graph database or a relational database through multi-table association) that can efficiently handle complex association relationships between entities.

[0018] Step two: based on the structured knowledge set in step one, calculate the absolute score of the global influence of the target entity; this step aims to calculate a quantitative ability index for each professional entity in the knowledge graph that is not limited by external context, objective and globally comparable. Its calculation process includes: Step 1, multi-dimensional ability capital variable calculation: based on the structured knowledge set, a pre-defined multi-dimensional ability capital feature vector is calculated for each professional entity through structured query (such as Cypher query) or graph algorithm. This vector contains at least the following three dimensions of variables: • Academic / knowledge capital variable: such as high-impact paper score based on journal impact factor and citation weight, high-level scientific research project dominance score, etc.

[0019] • Professional / skill capital variable: such as comprehensive work experience score, maximum difficulty coefficient of problems handled, and pioneering technical contribution degree, etc.

[0020] • Social / network capital variable: such as influence score in authoritative academic organizations, centrality (preferably betweenness centrality) in cooperation network, etc.

[0021] Step 2, fractional synthesis: input the normalized multi-dimensional capability feature vector into a pre-trained comprehensive evaluation regression model for end-to-end nonlinear mapping, and finally output a scalar value, i.e. the absolute score of global influence ). In a preferred embodiment, the model is a graph neural network (GNN, such as Graph Attention Network, GAT), which takes the feature vector as the initial input of the node, and can effectively capture the network structure information by propagating and aggregating information on the professional cooperation network subgraph, so as to obtain more accurate scores.

[0022] Step three: based on the absolute score of global influence in step two, context-aware situational influence analysis is carried out; this step aims to realize multi-level and penetrating analysis of individual influence based on a unified evaluation scale. Its implementation includes: Step y1, dynamic definition of evaluation range: receive one or more context parameters input by the user, which are used to define a specific evaluation range (for example, organizational range "XX Hospital XX Department", or geographical range "XX City").

[0023] Step y2, evaluation group screening: convert the context parameters into a structured query to the structured knowledge set to obtain all professional entities that meet the range conditions, forming a temporary evaluation group.

[0024] Step y3, relative position calculation: obtain the absolute score of global influence of all members in the evaluation group ), and calculate the ranking (Rank) and percentile (Percentile) of the target professional entity in the group, so as to output its relative influence in the specific context.

[0025] Step four: based on the perturbation simulation of the structured knowledge set in step one, carry out forward-looking influence deduction. This step is the core innovation point of the present application, aiming to realize the leap from static evaluation to dynamic prediction. The deduction process includes: Step s1, definition and baseline calculation of organizational macro indicators: a set of macro variable system capable of measuring the comprehensive capability of the organization (such as department, hospital) is defined in advance (for example, total amount of human capital, proportion of top-notch talents, internal collaboration network density, etc.). Before the deduction starts, first calculate all the macro variable values of the organization to be analyzed in the current state as the baseline (Baseline).

[0026] Step s2, non-persistent perturbation simulation: in response to a simulation event (e.g. "a core staff member flows from organization A to organization B"), a non-persistent, virtual modification is made to the structured knowledge set in a temporary, isolated computing environment (e.g. a memory graph copy). The modification includes at least one of the following: First, topology modification: temporarily add, delete or modify edges representing relationships between entities (e.g. break the-[employed by]-> relationship between the staff member and organization A, and establish the-[employed by]-> relationship between the staff member and organization B).

[0027] Second, attribute / weight modification: temporarily modify numerical attributes or weight parameters associated with related entities or relationships to simulate changes in their state in calculations.

[0028] Third, reachability / flow logic modification: modify at the algorithm level. For example, when calculating the internal network density of organization B after perturbation, the system temporarily includes staff member P in the node set of organization B in the program logic of the graph algorithm, even if the membership edge has not been established at the data level. This way directly simulates the change of membership in the calculation logic.

[0029] Step s3, post-perturbation recalculation and impact quantification analysis: based on the perturbed structured knowledge set, recalculate all macro variables of the organization to obtain new values (New_Value). Finally, by calculating the change rate , generate a clear, quantitative impact analysis report to accurately show the predicted impact of the simulation event on related parties in multiple dimensions.

[0030] Through the coordinated work of the above three steps, the present application completely builds a technical closed loop from data collection, knowledge construction to static evaluation, dynamic deduction, which has a logical and clear technical solution, strong implementability and can produce significant beneficial effects.

[0031] As shown in Figure 1 , as the second embodiment of the present application, this embodiment is illustrated in the medical field, but this should not be understood as limiting the application field of the present application.

[0032] Module one: structured knowledge aggregation unit, as shown in Figure 2 .

[0033] Step one in the method of the present application, i.e. "constructing or obtaining a structured knowledge set representing entities and their relationships in a professional field", this unit is the cornerstone of the entire system, and its core task is to automatically convert multi-source, heterogeneous raw data into a high-quality, high-accuracy, high-completeness structured knowledge model, providing a solid and reliable data foundation for the effective operation of all subsequent evaluation and deduction modules. Its specific implementation scheme includes the following key technical steps and considerations.

[0034] Step S1, multi-source heterogeneous data collection: This step aims to obtain as comprehensive raw data as possible from the Internet and internal databases.

[0035] (1) Definition and classification of data sources The system pre-configures and maintains an extensible data source list, which at least includes: • High-value structured / semi-structured data sources: including but not limited to domestic and foreign academic literature databases (such as PubMed, Web of Science, Google Scholar, China Knowledge Network), patent databases (such as USPTO, EPO, CNIPA), clinical trial registration centers (such as ClinicalTrials.gov), official agency directories (such as the National Health Commission physician registration information query system).

[0036] • High timeliness unstructured data sources: including but not limited to official websites, department homepages, news portals, industry vertical media, professional personal blogs or public social media pages of major hospitals and research institutes.

[0037] (2) Adaptation and execution of collection strategy: different collection strategies are adopted for different types of data sources: In step S1, for API interface: for data sources that provide API (such as PubMed), the system uses the API client written with preset query keywords (such as personnel name, institution name, disease name) and update frequency to perform periodic and incremental API calls to obtain structured data.

[0038] In step S1, for web data: a configurable distributed web crawler cluster based on Scrapy or similar frameworks is used. To achieve efficient and stable collection, the crawler cluster has the following technical features: Dynamic IP proxy pool: integrate multiple commercial or self-built IP proxy services to achieve millisecond-level switching of request IP to cope with IP-based access frequency limits; Dynamic User-Agent rotation: Maintains a list of hundreds of mainstream browser and mobile device User-Agents, randomly selects one for each request to simulate real user behavior. Adaptive request delay and session maintenance: Adjusts the request rate dynamically according to the response status code (such as 200, 429, 503) and response time of the target server, and simulates the login process to maintain the session when necessary.

[0039] 2. Step S2, knowledge extraction: This step uses a unified extraction framework based on a large language model (LLM) combined with specific prompt engineering to achieve high-precision and high-efficiency entity and relationship extraction.

[0040] (1) Long text preprocessing pipeline: To solve the input length limitation of LLM, reduce the calling cost and improve the extraction accuracy, all collected long texts (such as biographies of characters, long news) will go through a three-stage preprocessing pipeline of "summary-identification-refinement" before entering the extraction process: Step a, coarse-grained summary and key information area positioning: First, use an unsupervised extractive summary algorithm (such as TextRank) or a lightweight generative summary model (such as T5-small) to compress long texts into "long summaries" containing core information. In this process, the system will record the position index of each selected sentence in the original text for subsequent backtracking verification.

[0041] Step b, semi-structured knowledge extraction based on LLM: Embed the generated "long summary" text into a pre-designed unified Prompt template for entity and relationship extraction. The Prompt uses "In-Context Learning" and "Chain-of-Thought (CoT)" design patterns to guide the LLM to output all entities and relationships that meet the pre-defined Schema in JSON format at one time.

[0042] Step c, backtracking original text refinement and context verification: For each knowledge point (entity or relationship) extracted by LLM, the system uses the position index recorded in the first step to quickly locate its context fragment in the original text. Then, through a verification Prompt, LLM is required to combine more abundant original text context to perform secondary confirmation, information completion or correction on the extraction results, thereby greatly improving the accuracy of extraction.

[0043] (2) Extensible knowledge graph Schema design: The system pre-defines a hierarchical and extensible domain knowledge graph schema (Schema), which not only specifies the types of entities and relationships, but also defines their attributes, data types, and constraint conditions.

[0044] Entity layer (Nodes): For example, in the medical field, core entity types such as [doctor], [medical institution], [paper], [journal], [disease], [treatment / technology], [research project], and [society / association] are defined, and each entity type is specified with its mandatory or optional attributes (Properties), such as [journal] entities must have "impact factor" (floating-point type) and "partition" (character type) attributes.

[0045] Relationship layer (Edges): Detailed definitions of relationship types between entities, such as [employed by], [graduated from], [published], [co-author is], [good at], [hosted], etc. Some relationships may also have attributes, such as the [published] relationship with "author order" (e.g., 'first author', 'corresponding') and "publication year" attributes.

[0046] Step S3, knowledge fusion and verification: This step aims to solve the information conflict and redundancy problems from different data sources, and ensure the uniqueness, accuracy and integrity of knowledge, including:

[0047] Step c1, Entity Alignment: For entities from different data sources that may refer to the same real-world object (e.g., "XX University Third Hospital" vs. "Beijing Medical Third Hospital"), the system uses a hybrid alignment strategy: Step c10, rule-based fast matching: First, a pre-set synonym dictionary (such as institution aliases, common name writing methods) and string similarity algorithms (such as Jaro-Winkler) are used for a round of fast and high-confidence matching; Step c11, similarity calculation based on representation learning: For ambiguous cases that cannot be solved by rules, the system calculates the cosine similarity of the attribute vectors (based on text attributes) and graph embedding vectors (based on the neighbor structure in the graph, calculated by GraphSAGE and other models) of the entities to be matched; Step c12, pair-wise judgment based on pre-trained language models (LLM): The descriptive information of the two entities is input into a "pair-wise judgment" Prompt, and the LLM makes the final, human expert-like judgment.

[0048] Step c2, the multi-level resolution framework of "signal priority" of fact conflicts: when there are multiple mutually conflicting facts about the same attribute of the same entity (for example, source A says that a certain doctor is a "chief physician", and source B says that he is a "deputy chief physician"), the system starts a multi-level resolution framework, including: Level 1, priority rules based on metadata: first apply hard rules. Prefer information with the latest timestamp; if the timestamps are similar, prefer information from data sources with higher preset authority scores.

[0049] Level 2, topology verification based on internal consistency of the graph: if the rules cannot resolve, check which fact can form a logical closed loop with more existing high-confidence knowledge in the graph. For example, if the author's unit of the doctor's recent published papers is all A hospital, then the confidence of the fact that he "works in A hospital" will be much higher than "works in B hospital".

[0050] Level 3, context-based comprehensive reasoning resolution by LLM: for the most complex conflicts, the system organizes all conflicting facts and their metadata, related context evidence into a structured text, and hands it over to an LLM in the role of a "fact checker" for comprehensive reasoning, and requires it to output the final resolution result and detailed reasoning process.

[0051] Step S4, knowledge storage: this step aims to persistently store the high-quality knowledge after final cleaning and fusion.

[0052] (1) Definition of logical data model: the core of the invention is a logical data model that can represent a set of entities and the N-ary relationships between them.

[0053] (2) Implementation of physical storage: those skilled in the art should understand that the above logical data model can be physically implemented as (but not limited to): Step S40: a property graph database: this is the preferred embodiment of the invention, for example, using Neo4j, JanusGraph, etc. Entities are mapped to nodes, and relationships are mapped to edges with direction and attributes, which can most efficiently support subsequent graph queries and graph algorithms.

[0054] Step S41: an RDF triple store: for example, using Apache Jena.

[0055] Step S42: A set of data tables that are related to each other through foreign keys in a relational database (such as MySQL, PostgreSQL): for example, through a "personnel table", a "paper table" and a "personnel-paper relationship table".

[0056] Step S43: One or more JSON or XML documents: stored in a document database (such as MongoDB), representing the relationship between entities through nesting or ID reference.

[0057] Step S44: A group of object instances in the computer memory are connected to each other through pointers or references, and are used in high-performance real-time computing scenarios.

[0058] Through the above five tightly coupled steps, the structured knowledge aggregation unit of the present invention can build and maintain a high-quality structured knowledge set that can be used for serious evaluation and deduction, and provide a solid and reliable data foundation for the effective operation of all subsequent modules.

[0059] Module 2: Individual capability quantification unit, e.g. Figure 3 shown.

[0060] Step 2 of the method of the present invention is "calculating the absolute global influence score of the target entity based on the structured knowledge set of step 1"; this unit is the core computing engine of the entire system. Its task is to calculate one or more objective and globally comparable quantitative capability indicators for each professional entity in the knowledge graph that are not restricted by any external context (such as the organization or geographical location). In this embodiment, we will calculate a comprehensive "absolute global influence score ( )" as an example. The calculation process of this score is logically rigorous and data-driven, and is the technical cornerstone of all subsequent evaluation and deduction functions. Its specific implementation plan includes the following three key technical steps: Step 1: Refined calculation of the multi-dimensional capability capital variable system: This step aims to transform the discrete, multimodal information stored in the knowledge graph into a computable, high-dimensional feature vector that can fully reflect the individual's comprehensive capabilities. This feature vector is composed of a predefined multidimensional capability capital variable system.

[0061] (1) Definition of variable system: The system includes at least the following three capital dimensions, each of which has multiple specific variables that can be accurately calculated through a structured knowledge set.

[0062] Dimension One: Academic / Knowledge Capital - measures an individual's ability in knowledge creation, academic contribution, and frontier exploration.

[0063] Variable V_Paper_Impact (High Impact Paper Score): Through structured query, all [academic achievement] entities with the target entity as the core author (such as author order 'first author' or 'corresponding author') are traversed, and weighted sum is performed according to the preset impact score (such as impact factor, partition, number of citations, etc.) of the [journal] entity where the achievement is located. Its formula can be expressed as: .

[0064] Wherein, is the author order weight, is the citation number weight.

[0065] Variable V_Project_Lead (High Level Project Lead Score): Through structured query, [scientific research project] entities with the target entity as the core role (such as 'chief scientist' and 'project leader') are filtered out, and weighted sum is performed according to the level (such as 'national level' and 'provincial and ministerial level') and fund size of the project.

[0066] Variable V_Patent_Value (Patent Value Score): The number of [patent] entities of the target entity as the inventor is counted, and comprehensive weighting is performed according to the patent type (such as invention patent and utility model), patent family size, number of citations, and whether technology transfer occurs, etc.

[0067] Dimension Two: Professional / Skill Capital - measures an individual's practical experience, technical level, and ability to solve complex problems in a professional field.

[0068] Variable V_Experience_Comprehensive (Comprehensive Experience Score): A composite variable composed of multiple sub-attributes, calculated by the formula . Wherein, YearsOfPractice and TitleLevelScore are directly obtained from the attributes of [professional] entity, and the weight can be determined by AHP or by a committee of domain experts.

[0069] YearsOfPractice and TitleLevelScore are directly obtained from the attributes of [professional] entity, and the weight Can be calibrated by Analytic Hierarchy Process (AHP) or by a committee of domain experts.

[0070] Variable V_Skill_Pioneering (Technical Pioneering Contribution): Statistics the number of [treatment / technology] entities that the target entity has contributed as a pioneer or first introducer. This is a high-weight sparse variable to identify the technical leaders in the field.

[0071] Dimension Three: Social / Network Capital: Measures the individual's reputation, status, and ability to connect and mobilize resources within the industry.

[0072] Variable V_Network_Appointment_Influence (Academic Appointment Influence Score): Through structured queries, traverse the target entity's [academic appointment] relationships in various [society / association] entities, and add up the weights according to the level of the society (such as international, national, provincial) and the weight of the position (such as chairman, vice-chairman, member).

[0073] Variable V_Network_Centrality (Cooperation Network Centrality): This is a pure graph algorithm-driven variable. The system first constructs a cooperation network subgraph in the structured knowledge set, containing only [professional] entities and their [cooperator is]-> relationships. Then, on this subgraph, calculate the Betweenness Centrality for each node. This index measures to what extent a node is a "bridge" in the shortest path between other nodes in the network, and the higher the score, the stronger the node's role as a hub in knowledge dissemination and resource connection. Those skilled in the art will understand that other graph centrality algorithms such as Eigenvector Centrality or Degree Centrality can also be used as supplements or alternatives.

[0074] Step 2, construction and normalization of multi-variable feature vector, including: Step b1, construction of feature vector: After completing the calculation of all the above variables, the system will generate a unified dimension high-dimensional static capability feature vector for each professional entity in the knowledge graph, composed of the calculation results of these variables .

[0075] Step b2, data normalization: Due to the huge difference in the dimension and numerical range of different variables, the feature vector must be normalized before being sent to the final synthesis model Each dimension of the feature vector is normalized to eliminate the influence of dimensional differences and improve the training efficiency and stability of the model. The Z-Score standardization method is preferably used to convert the values of all variables into a standard normal distribution with a mean of 0 and a standard deviation of 1. Those skilled in the art can also select other normalization methods as needed, such as Min-Max Scaling.

[0076] Global influence absolute score Synthetic model of the global influence absolute score: This step is to synthesize the normalized multi-dimensional feature vector into a single, comparable scalar score . To achieve this purpose, the present application discards the traditional subjective linear weighting model and instead uses a pre-trained comprehensive evaluation regression model that can capture complex nonlinear relationships.

[0077] (1) Model selection and architecture: Preferred embodiment: Graph Neural Network (GNN) model Technical route: The present application preferably uses a Graph Attention Network (GAT). This model takes the feature vector generated in the previous step as the initial embedding representation of each node in the graph. Through multi-layer information propagation on the collaboration network subgraph, each node can selectively and weightedly aggregate the information of its neighbor nodes to update its own representation. This attention mechanism enables the model to automatically learn the deep network structure effect that "who you cooperate with matters more than how many times you cooperate with them".

[0078] Model architecture: The model contains several graph attention layers, each followed by a nonlinear activation function (such as LeakyReLU) and a Dropout layer. Finally, through a Global Pooling Layer and one to two Fully Connected Layers, the final node embedding representation is mapped to a scalar score .

[0079] Alternative embodiment: Other machine learning regression models Those skilled in the art will understand that in the simplified scenario without utilizing network structure information, other traditional machine learning regression models such as Gradient Boosting Decision Tree (XGBoost, LightGBM) or a Multi-Layer Perceptron (MLP) can also be used. These models treat the feature vector X i as a isolated data point, and fit a mapping function from features to scores through supervised learning.

[0080] (2) Model training and deployment include: Training Data Step: Model training isn't performed on every query; it's completed during system deployment or periodic updates. Training data comes from a high-quality dataset manually annotated by multiple domain experts using a back-to-back scoring method and averaging the results. This dataset contains feature vectors and their corresponding "expert consensus scores" for hundreds to thousands of experts.

[0081] Training process: Use Mean Squared Error (MSE) or Smooth L1 Loss as the loss function, and use Adam or AdamW optimizer to train and optimize model parameters until the model performance on the validation set converges.

[0082] Operation and output: In actual operation, the system only needs to convert the normalized feature vector of the person to be evaluated into (and its neighbor information, if GNN is used) is fed into this trained model for an efficient forward propagation, and the final global influence absolute score ( ), whose score range can be normalized to between 0 and 100.

[0083] Through the rigorous design and implementation of the above three steps, the individual ability quantification unit of the present invention can condense the rich, multi-dimensional information in the knowledge graph into an objective, fair and highly discriminatory single score, laying a solid and reliable quantitative foundation for all subsequent multi-level evaluation and dynamic deduction applications.

[0084] Module 3: Situational influence analysis unit, such as Figure 4 shown.

[0085] Step 3 of the method of the present invention, namely, "conducting context-aware situational influence analysis based on the global influence absolute score", is implemented by the situational influence analysis unit in the system. The core task of this unit is to use the "global influence absolute score ( )" as a unified, unchanging evaluation scale, enabling context-aware, multi-level, penetrating analysis of individual influence. To improve system response speed and user experience, this embodiment prefers a "pre-calculation and caching" strategy rather than real-time calculation for each query. Its specific implementation plan includes the following key technical steps.

[0086] 1. Pre-definition and configuration of multi-level evaluation system: This step aims to establish a structured and flexibly configurable evaluation hierarchy as the basis for subsequent batch calculations.

[0087] Step y1, dynamic definition of evaluation range:

[0088] Step y10, definition of evaluation hierarchy: The system pre-defines a set of standard evaluation tiers, which can be tree-like or network-like. In the medical field embodiment of the invention, the hierarchy at least includes: Organizational Tiers: Tier 1 (micro-level): specific department in a hospital (e.g., XX Hospital - Department of Cardiology) Tier 2 (meso-level): single medical institution (e.g., XX Hospital) Geographical Tiers: Tier 3 (city-level): city where the target entity is located (e.g., XX City) Tier 4 (regional / provincial level): province where the target entity is located or defined economic / geographical region (e.g., North China) Tier 5 (national level): nationwide (e.g., China) Domain-specific Tiers: Tier 6 (sub-field): specific professional field defined based on the [professional]-[:expertise]->[disease / technology] relationship (e.g., all professional groups proficient in "coronary intervention").

[0089] Step y11, configuration management: The evaluation hierarchy is stored in a configuration file (such as YAML or JSON file) or database table that can be accessed and modified by system administrators. This design allows flexible addition (such as adding a "global" tier), modification or disabling of specific tiers according to business needs, without modifying the core calculation code of the system, ensuring the scalability and maintainability of the system.

[0090] Step y2, efficient group data extraction based on batch graph query and caching mechanism: The core of this step is how to efficiently obtain the benchmark group data of all personnel in the knowledge graph in all preset tiers at one time.

[0091] Step y20, batch processing flow: This module is implemented as an offline, periodically batch computing task triggered by a scheduled task (e.g. Cron Job). The task will iterate through all entity nodes of type [Professional] in the knowledge graph.

[0092] Step y21, query cache and optimization (key performance guarantee): To avoid huge computing redundancy and database load caused by repeated queries of the same range of group data, this embodiment introduces a query result caching mechanism based on key-value storage (e.g. Redis, Memcached).

[0093] Technical route: Step y210, unique hierarchical instance extraction: when the batch task starts, first iterate through all [Professional] entities to extract all unique hierarchical instance lists to which they belong (e.g. all non-duplicate department names, all non-duplicate city names, etc.).

[0094] Step y211, batch query and cache of group data: the system only performs a structured query on each of these unique hierarchical instances once. The purpose of this query is to obtain the ID of all members under this hierarchical instance and their corresponding global absolute influence score (s_abs) list. After serialization, the query result is stored in the cache with the unique identifier of the hierarchical instance as the key (Key) and the value (Value) as “[{‘id’: ‘p1’, ‘s_abs’: 98.5}, {‘id’: ‘p2’, ‘s_abs’: 95.4}, …]”.

[0095] Step y212, cache life cycle management: the life cycle (TTL, Time-To-Live) of the cache should match the execution period of the batch task to ensure the timeliness of the data. For example, if the batch task is executed at 4:00 AM every day, the cache validity period can be set to 24 hours.

[0096] Step y3, calculation and storage of global relative influence indicators: After obtaining all relevant group data, the system calculates the complete, multi-level relative influence profile for each professional and stores it as a node attribute.

[0097] Step y30, batch computing process: the system iterates through each [Professional] entity p in the knowledge graph: Step y301: create a JSON object for entity p to store its complete influence profile, denoted as InfluenceProfile.

[0098] Step y302: Traverse all predefined evaluation tiers (Tier 1 to Tier 6).

[0099] Step y303: For each tier, first determine the specific instance that the entity p belongs to (e.g., the department A, city B, etc.).

[0100] Step y304: For each tier, first determine the specific instance that the entity p belongs to (e.g., the department A, city B, etc.).

[0101] Step y305: In this list, locate the position of entity p through efficient algorithms such as binary search, thereby calculating its rank (Rank) and percentile (Percentile) within the tier. The percentile can be calculated by the formula to ensure that the percentile of the highest score is 100. The structured result containing information such as scope name, rank, total number of people in the group, percentile, etc. is stored as a sub-object in InfluenceProfile.

[0102] Step y31, data storage and update (pre-computation strategy): Technical route: After calculation, the InfluenceProfile JSON object containing all tier influence indicators will be written as a new attribute directly to the node of the [professional] entity p.

[0103] Storage example (new p.InfluenceProfile attribute added to a doctor node): { "department": { "scope": "XX Hospital - Department of Cardiology", "rank": 1, "total": 25, "percentile": 100.0, "s_abs": 98.5 }, "hospital": { "scope": "XX Hospital", "rank": 5, "total": 850, "percentile": 99.41, "s_abs": 98.5 }, "city": { "scope": "XX city", "rank": 15, "total": 18500, "percentile": 99.92, "s_abs": 98.5 }, "last_updated": "2025-07-18T04:00:00Z" } Advantages and considerations: This "pre-computation and storage" strategy is a typical space-time trade-off solution. It places complex and time-consuming batch computation tasks in the background offline, which is not noticeable to the user. When the upper application (such as the user interface or API) needs to query the complete influence portrait of a professional, only a simple attribute reading operation is required for the node of the person, and all levels of evaluation results can be obtained instantly in milliseconds, greatly improving the response speed of the front-end application and user experience, and significantly reducing the real-time query pressure of the database.

[0104] Through the rigorous design and implementation of the above three steps, the context influence analysis unit of the present application can systematically and efficiently build a complete and multi-level relative influence portrait for each member in the knowledge graph, greatly improving the depth and breadth of evaluation, and providing strong performance guarantee for the instant call of the upper application.

[0105] Module four: system disturbance driving unit, as shown in Figure 5 .

[0106] Step four in the method of the present application, i.e. "based on the perturbation simulation of the structured knowledge set, forward-looking influence deduction", is realized by the system disturbance deduction unit in the system. This step is the most creative part of the present application, and its core task is to realize the paradigm breakthrough from "static evaluation" to "dynamic prediction". It accurately and quantitatively deduces how a simulated event (such as the inter-organizational flow of core professionals) will have a multi-dimensional impact on related parties through controllable and non-destructive perturbation simulation (Perturbation Simulation) on the structured knowledge set. This module is a key bridge connecting data analysis and high-level strategic decision-making, and its specific implementation scheme includes the following key technical steps.

[0107] Step s1, definition and baseline calculation of organizational macro-capability variable system: Before conducting any simulation, a set of computable macro-indicators that can measure the capabilities of organizations (such as hospitals, departments, and research teams) must be established.

[0108] (1) Definition of macro variables: This system predefines a set of configurable macro variables that can comprehensively reflect the overall strength of an organization. These variables can be derived through aggregate queries on structured knowledge sets or through graph algorithms. The core variables include at least: Human capital indicators: V_Org_Talent_Sum: Total amount of talent capital. Calculates the absolute score of global influence of all members in the organization ( Reflects the organization’s “total firepower” or overall strength.

[0109] V_Org_Talent_Mean: The mean value of talent capital. Calculated for all members in the organization The standard deviation of the data reflects the degree of dispersion or balance of talent levels within the organization.

[0110] V_Org_Talent_TopRatio: The ratio of top talents. The ratio of employees with scores in the top 1% (or other configurable threshold) to the total number of employees in the organization. This measures the concentration of core talent within the organization.

[0111] Output and reputation indicators: V_Org_Output_HighImpact: Total high-impact output. Aggregated calculation of the weighted sum of high-impact papers, high-level projects, and high-value patents produced by all members of the organization over the past period, reflecting the organization's academic or technical reputation.

[0112] Network structure indicators (core): V_Org_Network_Density: Internal collaboration network density. Calculate the density of the organization's internal collaboration subgraph, which is the ratio of the actual number of edges to the theoretical maximum possible number of edges. This measures the closeness of collaboration among internal members and reflects team cohesion and knowledge sharing efficiency.

[0113] V_Org_Network_Cohesion: Internal network cohesion. Calculates the number of connected components within the organization's internal cooperative subgraph. If this number is greater than 1, the organization is fragmented into uncooperative "cliques." This metric is key in determining whether an organization is at risk of internal friction.

[0114] V_Org_Network_BridgeCount: Number of bridges to outside organizations. Count how many members in the organization have cooperative relationships in other important organizations outside, these members are the "bridges" for the organization to obtain external information.

[0115] In step s1, the baseline calculation process: When a scenario to be deduced is received, which contains at least one simulated event (such as personnel P flowing from organization A to organization B), the system performs the following operations: Step s10: On the current unmodified, real structured knowledge set, perform a graph query for organization A and organization B respectively to filter out their respective member groups.

[0116] Step s11: Based on the two groups, calculate the current values of all the above macro variables.

[0117] Step s12: Store the calculated macro variable values of organizations A and B as baseline (Baseline) in memory or temporary database for subsequent comparison.

[0118] Step s2: Non-destructive perturbation simulation based on in-memory graph copy.

[0119] To ensure that the deduction process does not affect the online real knowledge graph data and can achieve high-performance real-time calculation, all simulation operations are performed in a temporary, isolated in-memory environment.

[0120] (1) Technical route: In-memory graph copy When a deduction request arrives, the system will not directly manipulate the physically stored graph database. Instead, it will perform an efficient graph query to load the subgraph related to the deduction, which includes the entity to be perturbed, related organization entities, and their respective members and first / second-order neighbor relationships, into the server's memory. This forms a lightweight, temporary graph copy. Those skilled in the art will understand that for large-scale graphs, more advanced graph sampling techniques (such as GraphSAINT) can also be used to obtain a representative and controllable size of the calculation subgraph.

[0121] (2) Implementation of perturbation operations (core protection point): The system performs non-persistent, virtual modifications on the structured knowledge in this in-memory graph copy to simulate events. This modification includes at least one or a combination of the following ways to block all possible circumvention paths: The first, topology modification: directly modify the adjacency relationship of the graph. For example, to simulate the flow of personnel, the system will temporarily delete the edge representing the membership of personnel P to the outflow organization A ((P)-[r: employed]->(A)), and temporarily create a new edge representing the membership of personnel P to the inflow organization B ((P)-[r': employed]->(B)).

[0122] The second, attribute / weight modification: without changing the topology, but by modifying the parameters to simulate state changes. For example, to simulate that a certain person "leaves but still cooperates", the system can temporarily modify the "state" attribute of the membership edge between him and the original organization from "employed" to "consultant", and adjust the weight factor from 1.0 to 0.2 when calculating the macro variables of the organization. To simulate "relationship break", you can temporarily set its weight to 0.

[0123] The third, reachability / flow logic modification: modify at the algorithm level. For example, when calculating the internal network density of organization B after the disturbance, the system temporarily includes personnel P in the node set of organization B in the program logic of the graph algorithm, even if the membership edge has not been established at the data level. This way directly simulates the change of membership in the calculation logic.

[0124] Step s3, post-perturbation index recalculation and impact quantification analysis: After the graph copy is successfully "disturbed", the system will recalculate the macro capabilities of the organization based on this simulated "future" graph state.

[0125] Step s30: Post-Perturbation Recalculation Process: The system recalculates all macro variables defined earlier on the disturbed in-memory graph copy for the member groups of organization A (which no longer contains or has changed the state of personnel P) and organization B (which now contains or has changed the state of personnel P), obtaining a new set of values (New_Value).

[0126] Step s31, impact quantification and analysis: The system compares the recalculated New_Value with the stored Baseline before the simulation begins, and calculates the absolute change and the relative change rate to accurately quantify the changes in each macro variable.

[0127] Step s32, generation and output of structured impact analysis report: The last step is to present all the calculated quantitative results in a clear, intuitive and easy-to-understand way for decision makers.

[0128] Output data structure steps: The system encapsulates all calculated changes into a structured JSON object. This object is clearly divided into two sections: "Impact on Outgoing Organizations" and "Impact on Incoming Organizations." Each section contains the baseline value, new value, absolute change, and relative rate of change for all macro variables.

[0129] Visualization and explanatory output steps: In user-interactive front-end applications, this data can be rendered into intuitive comparative charts. For example, a comparative radar chart can be used to show the comprehensive changes in various organizational capability indicators (such as total talent, network density, and output) before and after the disturbance.

[0130] • Use a waterfall chart to clearly show how the “total amount of talent capital” increases or decreases due to turnover.

[0131] • For changes in key indicators (e.g., network cohesion changes from 1 to 2), the system can automatically generate a natural language explanation: "Warning: The outflow of core member P may cause the internal cooperation network of Organization A to split into two independent sub-groups, posing a risk of team fragmentation." Through the rigorous design and implementation of the above four steps, the system disturbance simulation unit of the present invention can transform the evaluation system from a static information display tool into a data-driven dynamic "sand table simulation" engine that can support major strategic decisions such as talent introduction, risk management, and team building, thereby producing huge beneficial effects that are completely incomparable to existing technologies.

[0132] Those skilled in the art may conceive of technical alternatives to several core modules in the technical solution described in the present invention. However, in-depth technical comparative analysis reveals that while these alternatives may superficially achieve some functionality, they cannot achieve the same level of technical effectiveness, system performance, intelligence, or information utilization as the technical solution selected by the present invention. The technical deficiencies they present also contradict the necessity and creativity of the solution selected by the present invention.

[0133] 1. Analysis of alternatives to the score synthesis method in Module 2 (Individual Ability Quantification Unit) The preferred solution of this invention uses a graph neural network (GNN, such as GAT) model. This solution uses the calculated multidimensional variables as the initial features of the node. Through end-to-end learning on the professional collaboration network subgraph, it automatically aggregates neighbor node information and updates its own representation, ultimately mapping it to a score.

[0134] Alternative: Traditional machine learning regression models such as gradient boosting decision trees (XGBoost, LightGBM) or support vector regression (SVR) can be used. This solution requires all calculated variables as a feature set (X) and relies on a small-scale dataset pre-labeled by domain experts as training labels (Y) to train a regressor for score prediction through supervised learning.

[0135] Comparative analysis and conclusion: Although the alternative solution can also achieve the purpose of synthesizing multiple dimensions into a single score, it has two major defects: 1. Inherent deficiency of information utilization: The alternative solution treats each professional as an isolated data point, completely losing the valuable "relationship" information in the structured knowledge set. It cannot automatically learn and capture deep, non-local network structure effects like "who you work with matters more than how many times you have worked together" and "the value of being a bridge in the knowledge dissemination network" from the data like GNN. Therefore, its accuracy and depth of evaluation results are far inferior to the GNN solution.

[0136] 2. Limitations of model capabilities: Traditional regression models essentially fit a fixed feature space, while GNN learns in a dynamic information propagation space determined by the graph structure. GNN can better handle sparse, high-dimensional, and complex dependent data, and its model expression and generalization capabilities are usually superior to traditional models under the same data volume.

[0137] Therefore, the GGNN solution selected by the present application can produce more accurate, profound, and objective evaluation results by end-to-end fusion learning of node attribute information and graph topology information, and its technical effects are difficult to achieve by traditional machine learning alternative solutions.

[0138] 2. Alternative solution analysis for module four (system disturbance deduction unit) implementation method The present application solution: Adopt the method based on knowledge graph disturbance simulation. This solution simulates events by directly modifying the relationships (or their attributes, weights) between entities on the in-memory copy of the graph structure and recalculates the macroscopic graph indicators of the affected organizations to quantify the impact.

[0139] Alternative: A complex rule-based expert system can be constructed. This solution requires a large number of field experts to be invited in advance, and hundreds of "IF-THEN-ELSE" rules are defined manually through knowledge engineering. For example: "IF the title of the employee = 'Chief Physician' AND the S_abs of the employee > 95 AND the department = 'national key discipline' THEN the academic prestige of the organization = the academic prestige of the organization - 20".

[0140] Comparative analysis and conclusion: The rule-based alternative solution seems feasible in scenarios with simple logic and few influencing factors, but it has three fundamental flaws that make it impossible to truly achieve the purpose of the invention: 1) Unable to capture systemic chain reactions: Expert rules are essentially linear, local, and based on experience. They cannot capture the non-linear, complex network structure chain reactions triggered by personnel changes. For example, the departure of a key "bridge" figure may cause the originally tight internal cooperation network to split into several "small groups", resulting in a sharp decline in the organization's "internal network density" and "network cohesion" indicators. This systemic and emergent change cannot be predicted by any limited and pre-set rules.

[0141] 2) Poor scalability and maintainability: The construction and maintenance of the rule base is extremely costly, requiring a large amount of expert time and effort. When new evaluation indicators or new situations need to be added, the entire rule base may need to be modified and verified on a large scale to avoid conflicts between rules, which is extremely difficult and fragile in practice.

[0142] 3) Lack of dynamic adaptability: Rule-based systems are static, with knowledge fixed in rules. The graph-based simulation method of the invention is data-driven and dynamically adaptive. When the knowledge graph itself is updated due to the addition of new data, the results of the simulation will automatically and correspondingly change without any human intervention.

[0143] In summary, the graph perturbation simulation method of the invention can truly and dynamically reflect the systemic changes triggered by events by directly operating at the data structure level of the system. Its depth, accuracy, and automation far exceed the rigid, rule-based alternative solution.

[0144] Summary: Although there are the above-mentioned alternative technical paths, they either have inherent deficiencies in information utilization or have obvious short boards in model complexity and dynamic adaptability. The technical solution combination set forth in the present application, by organically combining the leading technologies such as graph neural network and graph perturbation simulation, is the best technical choice to achieve the purpose of the present application, and the beneficial effects brought by it are difficult to achieve by other alternative schemes.

[0145] The application further discloses a professional influence evaluation and dynamic deduction system based on a knowledge graph, which comprises a memory, a processor and a computer program stored in the memory.

[0146] The application further discloses a computer readable storage medium, which stores a computer program configured to realize the steps of the method of the application when called by a processor.

[0147] The above is a further detailed description of the application in combination with specific preferred embodiments, and the specific implementation of the application cannot be limited to these descriptions. For ordinary skilled persons in the technical field to which the application belongs, a number of simple deductions or substitutions can be made without departing from the concept of the application, and all of them should be regarded as falling within the protection scope of the application.

Claims

1. A method for evaluating and dynamically deducing the influence of professionals based on knowledge graphs, characterized by: include: Step 1: Construct a structured knowledge set that represents entities and relationships within a professional field; Step 2: Based on the structured knowledge set in step 1, calculate the absolute score of the global influence of the target entity; Step 3: Based on the absolute global influence score from step 2, perform context-aware situational influence analysis; Step 4: Based on the disturbance simulation of the structured knowledge set in step 1, perform forward-looking impact deduction.

2. The professional influence evaluation and dynamic deduction method according to claim 1 is characterized in that: The step one comprises: Step S1, multi-source heterogeneous data collection: obtaining structured data, semi-structured data, and unstructured data about professionals and their related activities from multiple predefined online data sources; Step S2, knowledge extraction: A unified extraction framework based on a pre-trained language model and combined with specific prompt engineering is used. This framework uses a three-stage processing pipeline of summarization-recognition-refinement to identify entities that match predefined patterns from long texts and extract multi-dimensional relationship triplets between them. Step S3: Knowledge fusion and verification: For entities from different data sources that point to the same real-world object, a multi-level conflict resolution framework based on signal priority is used to perform entity alignment and fact verification; Step S4, knowledge storage: the cleaned and integrated structured knowledge is loaded into a data storage system that processes complex relationships between entities.

3. The professional influence evaluation and dynamic deduction method according to claim 2 is characterized in that: In step S1, structured data and semi-structured data include domestic and foreign academic literature databases, patent databases, clinical trial registries, and official institution directories; unstructured data include official websites of hospitals and research institutes, department homepages, science and technology / health sections of news portals, industry vertical media, and personal blogs or public social media pages of professionals; In step S1, the API client performs periodic incremental calls to the data source providing the API using preset keywords and update frequency to obtain structured data; In the step S1, web page data is collected by a distributed web crawler cluster; The crawler cluster has the following technical features: Dynamic IP proxy pool: Integrate multiple commercial or self-built IP proxy services to achieve millisecond-level switching of request IPs to cope with IP-based access frequency restrictions; Dynamic User-Agent rotation: Maintain a list of User-Agents for several major browsers and mobile devices, randomly selecting one for each request to simulate real user behavior; Adaptive request delay and session persistence: Dynamically adjust the request rate based on the target server's response status code and response time, and maintain the session during the simulated login process.

4. The professional influence evaluation and dynamic deduction method according to claim 2 is characterized in that: In step S2, the three-stage processing pipeline includes: Step a: Coarse-grained summarization and key information area location: Using an unsupervised extractive summarization algorithm or a lightweight generative summarization model, the long text is compressed into a long summary containing the core information, and the position index of the selected sentences in the original text is recorded for subsequent backtracking verification; Step b: Extracting semi-structured knowledge based on a pre-trained language model: The long summary text is embedded in a pre-designed unified prompt template for entity and relationship extraction. This prompt template uses contextual learning and thought chain design patterns to guide the pre-trained language model to output all entities and relationships contained in the text that conform to the predefined schema in JSON format at once. Step c: Backtracking the refinement and context verification of the original text: For the knowledge points extracted by the pre-trained language model, use the position index recorded in step a to locate the context fragment in the original text. Through the verification prompt, the pre-trained language model is required to combine the original text context to perform secondary confirmation, information completion, or correction on the extraction results; In step S2, the predefined schema is a predefined domain knowledge graph schema, which is used to specify the types, attributes, data types and constraints of entities and relationships.

5. The professional influence evaluation and dynamic deduction method according to claim 2 is characterized in that: In step S3, the multi-level conflict resolution framework uses internal topological consistency of the metadata graph and contextual reasoning based on a pre-trained language model to ensure the uniqueness, accuracy, and completeness of the final knowledge; The step S3 specifically includes: Step c1, entity alignment: For entities from different data sources that may point to the same real-world object, a hybrid alignment strategy is adopted. The specific steps are as follows: Step c10, rule-based fast matching: high-confidence matching is performed using a preset synonym dictionary and string similarity algorithm; Step c11, similarity calculation based on representation learning: For ambiguous cases that cannot be resolved by rules, the cosine similarity of the attribute vector and graph embedding vector of the entity to be matched is calculated respectively; Step c12, pairwise judgment based on the pre-trained language model: Input the descriptive information of the two entities into the pairwise judgment prompt, and the pre-trained language model makes the final judgment similar to that of a human expert; Step c2, signal priority multi-level adjudication framework for conflicting facts: When there are multiple conflicting facts about the same attribute of the same entity, a multi-level adjudication framework is activated. The specific steps are as follows: Level 1, metadata-based priority rules: Apply hard rules to prioritize information with the latest timestamp. If timestamps are similar, information from a data source with a higher preset authority score is preferred. Level 2: Topological verification based on the internal consistency of the graph: If the rules cannot make a decision, the facts are verified against the existing, high-confidence knowledge in the graph to form a logical closed loop; Level 3: Contextual comprehensive reasoning and adjudication based on pre-trained language models: For complex conflicts, all conflicting facts, their metadata, and relevant contextual evidence are organized into structured text and submitted to the pre-trained language model of the fact-checker role for comprehensive reasoning. The final adjudication result and detailed reasoning process are required to be output.

6. The professional influence evaluation and dynamic deduction method according to claim 2 is characterized in that: In step S4, by defining the logical data model to represent entities and the many-to-many relationships between the logical data model to represent entities, the physical implementation process includes: Step S40: an attribute graph database, where entities are mapped as nodes and relationships are mapped as edges with directions and attributes; Step S41: RDF triple storage; Step S42: multiple data tables in a relational database are associated through foreign keys; Step S43: One or more documents stored in a document database in JSON or XML format, representing the relationship between entities by nesting or ID reference; Step S44: A set of object instances in the computer memory that are interconnected through pointers or references.

7. The professional influence evaluation and dynamic deduction method according to claim 1 is characterized in that: The second step includes: Step 1, multi-dimensional capability capital variable calculation: Based on the structured knowledge set in step 1, a predefined multi-dimensional capability capital feature vector is calculated for the professional entity through structured query or graph algorithm; Step 2, score synthesis: The normalized multi-dimensional capability capital feature vector is input into a pre-trained comprehensive evaluation regression model for end-to-end nonlinear mapping, ultimately outputting a scalar value, namely the absolute global influence score.

8. The professional influence evaluation and dynamic deduction method according to claim 7 is characterized in that: In step 1, the predefined multidimensional capability capital feature vector includes at least three variables of the following dimensions: academic / knowledge capital variable, professional / skill capital variable, and social / network capital variable; In step 2, the method further includes: Step b1, constructing feature vectors: Generate a high-dimensional static capability feature vector of uniform dimension composed of variable calculation results for the professional entity in the knowledge graph X i ; Step b2, data normalization: using a statistically based standardization method, the values ​​of all variables are transformed into a standard normal distribution with a mean of 0 and a standard deviation of 1; Step 2 also includes the training and deployment of a comprehensive evaluation regression model. The specific steps are as follows: Training data: Model training is completed during deployment or periodic updates, not during each query. Training data comes from a manually labeled dataset with back-to-back scoring and averaging by multiple senior experts in the field. This dataset contains feature vectors from multiple professionals and their corresponding expert consensus scores. Training process steps: Use mean square error or smoothed L1 loss as the loss function, and use an adaptive learning rate optimizer to train and optimize model parameters until the performance of the validation set converges; Run and output steps: During actual operation, you only need to feed the normalized feature vector of the person to be evaluated into the trained model for forward propagation to calculate the final absolute global influence score, which is normalized to between 0 and 100.

9. The professional influence evaluation and dynamic deduction method according to claim 1 is characterized in that: The step three includes: Step y1, dynamic definition of evaluation scope: receiving one or more context parameters input by the user to define the evaluation scope; Step y2, evaluation group screening: converting the situational parameters into a structured query on the structured knowledge set to obtain all professional entities that meet the conditions and form a temporary evaluation group; Step y3, relative position calculation: obtain the absolute scores of the global influence of all members in the evaluation group, calculate the ranking and percentile of the target professional entity in the group, and output its relative influence.

10. The professional influence evaluation and dynamic deduction method according to claim 9 is characterized in that: In the step y1, it also includes: Step y10: pre-define a standard, tree-like or network-structured evaluation hierarchy, which at least includes an organizational level, a geographical level, and a professional field level; Step y11: Store the evaluation hierarchy in a configuration file or database table that is accessed and modified by the system administrator, allowing specific levels to be added, modified, or disabled based on business needs without having to modify the core calculation code of the system.

11. The professional influence evaluation and dynamic deduction method according to claim 10 is characterized in that: In the step y2, it further includes: Step y20, batch processing flow: This is performed using an offline, periodic batch computing task triggered by a scheduled task. This task traverses all entity nodes of the professional type in the knowledge graph. Step y21, query caching and optimization: Use a key-value storage-based mechanism to cache query results. The specific steps are as follows: Step y210, unique level instance extraction: when the batch task is started, all professional entities are traversed to extract a list of all unique level instances to which they belong; Step y211, group data batch query and cache: only perform a structured query on each unique level instance, and obtain the IDs of all members under the level and their corresponding global influence absolute scores. After the query results are serialized, they are stored in the cache using the unique identifier of the level instance as the key; Step y212, cache lifecycle management: The cache lifecycle matches the execution cycle of batch tasks to ensure the timeliness of data.

12. The professional influence evaluation and dynamic deduction method according to claim 11 is characterized in that: In the step y3, it further includes: Step y30, batch calculation process: traverse each professional entity p in the knowledge graph, specifically including: Step y301: Create a JSON object for entity p to store its complete influence profile, denoted as InfluenceProfile; Step y302: traverse all predefined evaluation levels; Step y303: For each level, determine the specific instance to which entity p belongs at that level; Step y304: Read the sorted instance directly from the cache Score list; Step y305: From the score list of step y304, locate the position of entity p through efficient algorithms such as binary search, and calculate its ranking and percentile in the hierarchy. The percentile is calculated by the formula , ensuring that the highest percentile score is 100, the structured result containing the range name, rank, total number of people in the group, and percentile information is stored as a sub-object in InfluenceProfile; Step y31, data storage and update: After the calculation is completed, the InfluenceProfile JSON object containing the influence indicators of all levels is written and stored directly on the node of the professional entity p as a new attribute.

13. The professional influence evaluation and dynamic deduction method according to claim 1 is characterized in that: In step 4, the deduction process includes: Step s1, definition of organizational macro indicators and baseline calculation: pre-define a macro variable system to measure the comprehensive capabilities of the organization. Before the simulation begins, calculate the values ​​of all macro variables in the current state of the organization to be analyzed as a baseline. Step s2, non-persistent perturbation simulation: performing non-persistent, virtual modifications to the structured knowledge set in a temporary, isolated computing environment; Step s3, post-disturbance recalculation and impact quantification analysis: Based on the structured knowledge set after the disturbance, recalculate all macro variables of the organization to obtain new values, and calculate the change rate , generate a quantitative impact analysis report to show the predicted impact of simulated events on relevant parties in multiple dimensions.

14. The professional influence evaluation and dynamic deduction method according to claim 13 is characterized in that: In step s1, the macro variables are obtained by aggregate query or graph algorithm calculation of structured knowledge sets, and the core variables include at least talent capital indicators, output and reputation indicators, and network structure indicators. The talent capital indicators include the total talent capital, the average talent capital, and the average talent capital. The output and reputation indicators include the total high-impact output. The network structure indicators include internal collaborative network density, internal network cohesion, and the number of external connection bridges. In step s1, the baseline calculation process is as follows: Step s10: Execute graph queries on the current unmodified, real structured knowledge set for organizations A and B respectively to filter out their respective member groups; Step s11: Based on the two groups in step s10, calculate the current values ​​of all macro variables; Step s12: The calculated macro variable values ​​of organization A and organization B are stored as baselines in the memory or temporary database.

15. The professional influence evaluation and dynamic deduction method according to claim 13 is characterized in that: In step s2, the modification includes at least one of the following methods: The first type is topological structure modification: temporarily adding, deleting, or modifying edges that represent relationships between entities; The second type, attribute / weight modification: temporarily modifying the numerical attributes or weight parameters associated with the relevant entities or relationships to simulate the change of their state in the calculation; The third type is reachability / flow logic modification: At the algorithm level, the computational logic of the graph algorithm is temporarily modified to simulate changes in entity ownership relationships. In the step s3, it further includes: Step s30, post-disturbance recalculation process: on the perturbed memory graph copy, recalculate all previously defined macro variables for the member groups of organization A and organization B respectively to obtain a set of new values; Step s31, impact quantification and analysis: the recalculated new values ​​are compared one by one with the baseline stored before the deduction began, and the change of each macro variable is accurately quantified by calculating the absolute change and relative change rate; Step s32, generation and output of a structured impact analysis report: present all calculated quantitative results in a way that decision makers can understand, including: Output data structure step: Encapsulate all calculated changes into a structured JSON object. This object is clearly divided into two parts: the impact on the outflow organization and the impact on the inflow organization. Each part contains the baseline value, new value, absolute change, and relative change rate of all macro variables. Visualization presentation and explanatory output step: In the front-end application that interacts with users, the data is rendered into comparative charts.

16. A knowledge graph-based professional influence evaluation and dynamic deduction system, characterized by: include: A memory, a processor, and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the method according to any one of claims 1 to 15 when called by the processor.

17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of any one of claims 1 to 15 when called by a processor.

Citation Information

Patent Citations

  • Measuring method of social network communication influence and measure system thereof

    CN104008182A

  • Method for assessing and sorting citation network academic influences based on credibility

    CN107391659A

  • Talent information processing method and system based on occupational attainment and big data analysis

    CN119784346A

  • Knowledge graph-based influence prediction method and system

    CN120494185A

  • GNN-based talent evaluation model and method and talent selection method

    CN120509771A

Cited By

  • Large-scale high-speed text training comparison data set production device

    CN121278393A