Method for evaluating achievement transformation potential based on dynamic knowledge graph
By generating structures such as session packages and evidence packages in the dynamic knowledge graph, and using methods such as session hashing and slot templates, the problem of inconsistent collection standards across data sources is solved, and the data probe and slot template are firmly bound together, ensuring the traceability of the evaluation model and the stability of the results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN NATURAL SELECTION INT INTELLECTUAL PROPERTY CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies lack unified constraints in dynamic knowledge graphs, resulting in inconsistent collection criteria across data sources, difficulty in forming stable criteria for session packets, difficulty in fully preserving temporal payloads of evidence packets, difficulty in consistently merging candidate packet relationship records, and unclear traceability entry points between historical transformation training parameters inferred by the evaluation model and output results, making it difficult to achieve stable implementation of potential scores, risk indicators, and path evidence interpretation.
By generating a continuous and consistent process of session packages, evidence packages, candidate packages, admission packages, version packages, subgraph packages, chain feature packages, and vector packages, and using methods such as session hash registration, slot templates, event timestamps, database entry timestamps, and delay indicators, the data probes are permanently bound to the slot templates, ensuring a consistent standard and traceability for data collection and processing.
It achieves a unified standard for cross-data source collection and processing, ensuring that data probe configuration and slot templates are consistent throughout the probe collection and subsequent processing chain. It provides an associative session-level index entry, solves the consistency problem of data organization and field meaning across steps, and improves the traceability of the evaluation model and the stability of the results.
Smart Images

Figure CN122045259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dynamic knowledge graphs and data probe acquisition and processing, and particularly to a method for evaluating the potential of achievement transformation based on a dynamic knowledge graph. Background Art
[0002] In the field of dynamic knowledge graphs and data probe acquisition and processing, existing solutions for evaluating the potential of achievement transformation around an achievement identification set and a data source list usually rely on the concatenation of multiple data source access, entity recognition, entity alignment, relationship extraction, and evaluation model inference. However, there is a lack of unified constraints in aspects such as cross-data source acquisition caliber, consistency of time-series fields, and traceability of graph evolution. There are limitations such as difficulty in forming a stable caliber for session packets, difficulty in completely retaining the time-series payload of evidence packets, and difficulty in consistently merging candidate packet relationship records. Most existing methods directly use the records obtained by probe acquisition for relationship extraction or scoring calculation, and there is a lack of a systematic constraint link for event timestamps, warehousing timestamps, and delay metrics. In scenarios where multi-data source asynchronous warehousing and multi-batch updates occur concurrently, problems such as difficulty in locating multi-version conflicts of the same relationship, lack of a unified determination basis for the credibility score and conflict mark of relationship records within candidate packets are likely to occur, making it difficult to stably achieve output potential scores, risk indicators, confidence levels, and path evidence explanations. For the joint processing between session packets, evidence packets, candidate packets, and admission packets, existing technologies generally lack an integrated management link for graph snapshots, snapshot hashes, and differential packets in the connection of gating resolution, graph update, and version packet generation. It is difficult to form a continuous and consistent process from evidence packets to candidate packets to admission packets and further generate version packets, sub-graph packets, chain feature packets, and vector packets in application scenarios constrained by relationship types, time windows, and credibility thresholds, resulting in unclear correspondence between the change sources of graph snapshots and differential packets, unstable time-series association between window indices and chain feature packets, and unclear traceability entry between historical transformation training parameters relied on by evaluation model inference and output results, thus causing the objective impacts of difficult review and difficult closed-loop management of the evaluation link in actual engineering operations. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a method for evaluating the potential of achievement transformation based on a dynamic knowledge graph, including:
[0004] S100. Obtain an achievement identification set and a data source list, configure data probes according to the data source list and generate slot templates, register a session hash, and generate a session packet;
[0005] S200. Perform probe acquisition on the session packet to generate an evidence packet; the evidence packet includes an event timestamp, a warehousing timestamp, and a delay metric;
[0006] S300: Perform entity recognition, entity alignment, and relation extraction on the evidence package to generate candidate packages, and bind slot identifiers, probe identifiers, credibility scores, and conflict markers to the relations;
[0007] S400. Perform gated resolution on the candidate packets to generate admission packets and generate a verification set;
[0008] S500: Perform graph update on the admission package to generate a version package; the version package includes a graph snapshot, a snapshot hash, and a differential package;
[0009] S600. Perform subgraph extraction on the version package, locate the result entity in the graph snapshot, extract and evaluate subgraphs according to relation type, time window and confidence threshold, and generate subgraph package;
[0010] S700. Perform link construction on the sub-graph package to generate a chain feature package; the chain feature package includes traction chain features, risk chain features, and differential packages;
[0011] S800: Perform time-series calculations on the chain feature package to generate a window index and synthesize a vector package;
[0012] S900, Perform evaluation model inference on the vector package; the evaluation model includes historical transformation training parameters and outputs potential score, risk index, confidence level, and path evidence interpretation.
[0013] Furthermore, the data probe includes a sampling granularity field, a field mapping table, an entity dictionary version, a data source access identifier, an extraction rule identifier, and a quality verification rule identifier; the quality verification rule identifier is associated with a missing test flag, an outlier threshold, and a duplicate record discrimination rule.
[0014] Furthermore, the slot template includes a slot identifier, an entity type set, a relation type set, a set of required fields, a set of evidence fields, a credibility threshold, and a conflict type code.
[0015] Furthermore, the delay index is obtained by subtracting the event timestamp from the entry timestamp; the gating resolution includes generating delay tags according to the delay index threshold and writing the relationships with delay tags into the review set.
[0016] Furthermore, the entity alignment uses name similarity, attribute similarity, and semantic vector similarity to synthesize an alignment score, constructs an entity alignment mapping table, and writes it into the candidate package.
[0017] Furthermore, the conflict marker includes a conflict cluster identifier, a conflict type code, a conflict density value, and a primary slot identifier.
[0018] Furthermore, the differential package includes a set of newly added nodes, a set of deleted nodes, a set of newly added relationships, a set of deleted relationships, a set of weight changes, and a change time window index; the change time window index is associated with the event timestamp.
[0019] Furthermore, the traction chain features include chain length, path edge weight product, number of industry nodes covered by the path, number of investment and financing event nodes, and path credibility composite value; the risk chain features include negative event node count, ownership conflict node count, similar achievement competition node count, and risk propagation depth.
[0020] Furthermore, the window index is synthesized from the position of the mutation point, the trend slope, the volatility, and the heat decay coefficient of the trend sequence; the trend sequence is updated driven by the change time window index in the difference package.
[0021] Further, after S900, drift detection is performed on the potential score, the risk index, and the path evidence interpretation; the drift detection includes comparing the historical window index interval and generating a disposal transaction package; the disposal transaction package includes a disposal type code, a target node identifier, a target relationship identifier, a session hash, a snapshot hash, an evaluation model version number, and a disposal timestamp.
[0022] The key innovations of this invention include:
[0023] (1) Under the constraints of the result identifier set and the data source list, the session package is generated by the session hash registration, and the data probe and the slot template are fixed in the session package, so that the generation links of the probe collection, the evidence package, the candidate package and the admission package have a consistent session-level input caliber.
[0024] (2) In the process of entity identification, entity alignment and relation extraction to generate the candidate package, the relationship is bound to the slot identifier, probe identifier, credibility score and conflict marker, and the candidate package is subjected to admission determination in the gating resolution, the admission package is generated and the review set is generated simultaneously.
[0025] The following are its main beneficial effects:
[0026] (1) In view of the problem that the existing solution lacks a unified standard when organizing collection and processing across the data source list, which makes it difficult for the data probe configuration and the slot template to be integrated into the probe collection and subsequent processing link, the present invention generates the session package by registering the session hash and solidifies the data probe and the slot template, so that the session package can be called as the same session-level configuration payload in the running link of the probe collection to generate the evidence package, the evidence package to generate the candidate package, and the candidate package to generate the admission package. This makes the data organization and field meaning of the cross steps consistent at the session dimension, and provides a session-level index entry that can be associated with the review set and the version package.
[0027] (2) In view of the problems that existing solutions directly use the collected records for relation extraction or scoring processing and lack the link constraint of the event timestamp, the entry timestamp and the delay index, the difficulty in consistently merging the relation records in the candidate package in the multi-source asynchronous entry scenario, and the lack of a unified judgment basis for the credibility score and the conflict mark, the present invention binds the relationship to the slot identifier and the probe identifier when generating the candidate package and generates the credibility score and the conflict mark. At the same time, in the gating resolution, the delay index is combined to perform admission judgment on the candidate package, output the admission package and write the resolution decision and the record to be reviewed into the review set, so that the relation records in the admission package have a verifiable gating source and conflict handling path, and provide a consistent admission path for the input boundary of the subsequent graph update.
[0028] (3) In view of the problems of existing solutions lacking integrated management of the graph snapshot, the snapshot hash and the differential package during the graph update process, resulting in unclear traceability entry between the graph change source and the evaluation link, and unstable temporal correlation between the window index and the chain feature package, this invention triggers the graph update with the admission package and generates the version package containing the graph snapshot, the snapshot hash and the differential package. Then, based on the version package, the subgraph extraction, the link construction and the temporal calculation are performed to form the vector package and input into the evaluation model for inference. This enables the potential score, the risk index, the confidence level and the path evidence interpretation to form a versioned link that can be associated with the graph snapshot and its differential package. Thus, under the constraints of the relationship type, the time window and the confidence threshold, a continuous and traceable process is established from the admission package to the version package, the subgraph package, the chain feature package and finally to the vector package and the inference output of the evaluation model. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating a method for evaluating the commercialization potential of research results based on dynamic knowledge graphs, as provided in an embodiment of this application. Detailed Implementation
[0030] Reference Figure 1 This is a flowchart illustrating a method for evaluating the commercialization potential of research results based on dynamic knowledge graphs, provided by an embodiment of the present invention. The flowchart may include at least steps S100-S900:
[0031] S100: Obtain the result identifier set and data source list, configure data probes according to the data source list and generate slot templates, register session hashes, and generate session packets;
[0032] S200. Perform probe collection on the session packet to generate an evidence package; the evidence package includes an event timestamp, an entry timestamp, and a delay index.
[0033] S300: Perform entity recognition, entity alignment, and relation extraction on the evidence package to generate candidate packages, and bind slot identifiers, probe identifiers, credibility scores, and conflict markers to the relations;
[0034] S400. Perform gated resolution on the candidate packets to generate admission packets and generate a verification set;
[0035] S500: Perform graph update on the admission package to generate a version package; the version package includes a graph snapshot, a snapshot hash, and a differential package;
[0036] S600. Perform subgraph extraction on the version package, locate the result entity in the graph snapshot, extract and evaluate subgraphs according to relation type, time window and confidence threshold, and generate subgraph package;
[0037] S700. Perform link construction on the sub-graph package to generate a chain feature package; the chain feature package includes traction chain features, risk chain features, and differential packages;
[0038] S800: Perform time-series calculations on the chain feature package to generate a window index and synthesize a vector package;
[0039] S900, Perform evaluation model inference on the vector package; the evaluation model includes historical transformation training parameters and outputs potential score, risk index, confidence level, and path evidence interpretation.
[0040] To address the problems in existing technologies, such as inconsistent identification standards for achievements, scattered data sources with varying access methods, leading to difficulties in aligning the reference identifiers and field meanings of the same achievement across different systems, and the inability to establish a consistent input standard for subsequent data collection, further resulting in a lack of structured evidence and auditing relevance, this invention completes the structured access of the achievement identifier set and data source list through step S100, and solidifies the data probes and slot templates into a traceable configuration payload of the session packet, providing a unified input standard for subsequent probe collection and evidence structuring from the source. Specifically, this includes:
[0041] S100: Obtain the result identifier set and data source list, configure data probes according to the data source list and generate slot templates, register session hashes, and generate session packets;
[0042] In the specific implementation of this invention, the achievement identifier set is read from the achievement registration database by the achievement management terminal and aggregated by the session orchestration unit. The achievement identifier set is used to indicate the scope of achievement entities to be evaluated. Each achievement identifier in the achievement identifier set includes achievement number, achievement name, affiliated institution identifier, achievement type identifier, and registration timestamp fields. The achievement type identifier is used to distinguish the registration categories of patent achievements, paper achievements, software copyright achievements, and trade secret achievements. The data source list is loaded by the data source management terminal during system initialization or when the evaluation task is triggered. The data source list is used to indicate the external or internal data access objects participating in this evaluation session. The data source list includes data source access identifier, data source type identifier, access endpoint description, update frequency description, and authorization credential indicator fields. The data source type identifier is thus limited to patent data sources, paper data sources, industry dynamic data sources, investment and financing event data sources, and enterprise product data sources required for knowledge graph construction, and a corresponding relationship is established with the achievement type identifiers in the achievement identifier set. Specifically, after receiving the evaluation task trigger signal, the session orchestration unit takes the result identifier set and the data source list as input, and first performs a consistency verification process. The consistency verification process includes performing a uniqueness check on the result number field, a reachability check on the data source access identifier, and a validity check on the authorization credential indication field. The verification failure record is written into the session log structure to form a traceable record. If the consistency verification passes, the process link of configuring data probes according to the data source list is entered.
[0043] The data probes are generated and loaded by the probe configuration unit. The data probes are used to collect and extract structured data from specified data sources. The minimum set of data probes consists of a sampling granularity field, a field mapping table, an entity dictionary version, a data source access identifier, an extraction rule identifier, and a quality verification rule identifier. The sampling granularity field is used to constrain the discretization method of the collection time window. The field mapping table is used to map data source fields to unified fields. The entity dictionary version is used to constrain the vocabulary scope of entity recognition and entity alignment. The data source access identifier is used to bind the access endpoint description in the data source list. The extraction rule identifier is used to bind the rule set of structured extraction. The quality verification rule identifier is used to bind the missing test flag, outlier threshold, and duplicate record discrimination rules. Specifically, the probe configuration unit parses the data source list item by item, selects the corresponding sampling granularity field based on the data source type identifier, and the generation of the sampling granularity field is completed by the time window management unit. The time window management unit reads the update frequency description and generates an evaluation time window index, which is then written into the sampling granularity field. Subsequently, the probe configuration unit generates a unified field list based on the field mapping table. The field mapping table is loaded and formed by the field governance unit. The field governance unit pre-sets a unified field specification for each data source type and establishes a mapping relationship from the source field to the unified field. The unified field includes at least the record primary key, event timestamp, database entry timestamp, source identifier, text payload field, and structured attribute field. The event timestamp and database entry timestamp are used for the evidence package formation in subsequent S200, and the source identifier is used for the relationship binding probe identifier in subsequent S300. Furthermore, the entity dictionary version is loaded by the terminology management unit. When the evaluation task is triggered, the terminology management unit reads the domain ontology and the synonym merging table, generates an entity dictionary version number, and writes it into the data probe. This is used to constrain the spelling consistency and name disambiguation in the subsequent entity recognition and entity alignment processes. The extraction rule identifier is generated by the rule management unit. The rule management unit selects an extraction template and registers the template version number according to the data source type identifier. The extraction template includes page field parsing rules, text cleaning rules, entity candidate generation rules, and relation candidate generation rules. The quality verification rule identifier is generated by the quality management unit. The quality management unit binds the missing test marking rule, outlier threshold rule, and duplicate record discrimination rule to the quality verification rule identifier and writes it into the data probe. The missing test marking rule is used to write a missing test mark to null value fields mapped by the field mapping table. The outlier threshold rule is used to write an outlier mark to numerical or time anomalies in structured attribute fields. The duplicate record discrimination rule is used to write a duplicate mark to duplicate sampling of the same record primary key and the same event timestamp combination.Understandably, the aforementioned minimum set of data probes constitutes the core parameter set of S100, which supports the probe acquisition and structured generation of evidence packages in S200. On this basis, the probe configuration unit also loads preferred extended fields, which include conflict priority, source credibility baseline, incremental cursor field and acquisition failure retry policy field. The acquisition failure retry policy field is used to constrain the number of retries, backoff interval and circuit breaker conditions of automated operation, and writes retry and circuit breaker events into the session log structure.
[0044] After the data probe configuration is completed, the slot template is generated by the slot modeling unit. The slot template is used to constrain the evidence structure and credibility threshold of the relationship in the subsequent candidate package. The minimum set of the slot template consists of slot identifier, entity type set, relationship type set, required field set, evidence field set, credibility threshold and conflict type code. Specifically, the slot modeling unit takes the result type identifier of the result identifier set and the data source type identifier of the data source list as input, first generates a slot identifier namespace, and the slot identifier namespace is generated and registered according to the concatenation rule of "result type identifier - data source type identifier - relation family identifier"; then, the slot modeling unit solidifies the entity type set according to the entity dictionary version. The entity type set includes type identifiers of result entities, patent entities, paper entities, enterprise entities, product entities, technical concept entities, and investment and financing event entities. The relation type set includes type identifiers of citation relationships, family relationships, cooperation relationships, investment relationships, application relationships, substitution relationships, and competition relationships. The required field set is provided by the field governance unit and generated according to the field mapping table. The required field set includes record primary key, event timestamp, entry timestamp, source identifier, and relation type identifier fields. The evidence field set includes text payload field, structured attribute field, extraction rule identifier field, and quality mark field. Furthermore, the credibility threshold is generated by the threshold configuration unit, which assigns a credibility threshold value based on the data source type identifier and the quality verification rule identifier and writes it into the slot template. The conflict type code is generated by the conflict encoding unit, which defines a conflict type code set based on the relation type set and registers the code table version number. The conflict type code set contains encoded entries for field conflicts, time conflicts, subject conflicts, and source conflicts, which are used to form conflict markers in subsequent S300 and referenced when performing gating resolution in S400. Through the above processing, the slot template and the data probe form a one-to-one or many-to-one correspondence. The correspondence is generated by the slot mapping unit and written into the session packet metadata structure. When generating the correspondence, the slot mapping unit performs a coverage check on the entity type set and the relation type set. Mapping entries that fail the coverage check are written into the session log structure and marked with an exception type code.
[0045] After the data probe and slot template are generated, the session hash is generated and registered by the hash registration unit. The session hash is used to identify the key configurations of this evaluation session for consistency. Specifically, the hash registration unit performs serialization and normalization processing on the result identifier set, the data source list, the data probe, and the slot template. The serialization and normalization processing includes field order fixing, null value field removal rule fixing, and encoding format fixing. Based on the normalization result, a hash digest is generated, and the hash digest is registered as the session hash. The hash registration unit writes the session hash, session number, generation timestamp, entity dictionary version, extraction rule identifier version number, quality verification rule identifier version number, and conflict type code table version number into the session registration table structure. The session registration table structure is stored in the session storage unit for subsequent steps and for audit traceability. Further, the session orchestration unit performs atomic write processing on the session registration table structure. Atomic write processing includes locking the session number before writing and releasing the lock after writing. If locking fails, a session conflict alarm is generated and written to the session log structure. The alarm triggers the collection task suspension policy and waits for the next trigger signal to restart the generation process.
[0046] After completing session hash registration, the process of generating session packets is completed by the session encapsulation unit. The session packet is the output product of S100 and includes an encapsulation structure containing a session number, session hash, result identifier set, data source list, data probe, slot template, and slot mapping relationship. Specifically, the session encapsulation unit writes the session number and session hash into the session header field, the result identifier set into the result input field, the data source list into the data source input field, the data probe into the probe orchestration field, the slot template into the slot template field, and the slot mapping relationship into the mapping field. The session encapsulation unit performs integrity checks on the session packet. The integrity checks include checking the existence of the session hash field, checking the consistency of the access identifier between the probe orchestration field and the data source input field, and checking the consistency of the slot identifier between the slot template field and the mapping field. If the checks fail, the reason for the failure and the associated field name are written into the session log structure and an exception type code is marked. Finally, the session packet is written to the session storage unit as the output of S100, and is scheduled to be read when the trigger condition of the evaluation task scheduling unit is met, serving as the input of S200 to "perform probe acquisition on the session packet". The probe orchestration field in the session packet is the input position for S200 to call the data source access identifier and sampling granularity field. The slot template field in the session packet is the reference basis for S200 to write the evidence field set and quality mark field when generating the evidence packet. The session hash in the session packet is carried and transmitted as the session association field of S200 to generate the evidence packet.
[0047] Summary of the technical effects of this step: This step completes the structured access of the result identifier set and data source list, and solidifies the data probes and slot templates into a traceable configuration payload of the session packet, forming a unified input standard for subsequent probe collection and evidence structuring. This step forms the session packet under the constraints of session hash registration and version number binding, supporting cross-step configuration consistency and audit correlation. This step writes probe orchestration, slot templates, and mapping relationships into the session packet and specifies subsequent input positions, establishing a hierarchical input link from S100 to S200.
[0048] To address the problem that existing technologies often lack session-level configuration constraints and unified field mapping in cross-data source acquisition, making it difficult to express event occurrence and data entry delays with consistent time-series fields, and further hindering subsequent entity identification and relation extraction from distinguishing between "event time" and "data arrival time," thus introducing time-series noise, and lacking reliable delay criteria for conflict determination and gating resolution, this invention completes cross-data source probe acquisition and unified field mapping under session packet constraints in step S200, and encapsulates the acquisition records into an evidence package with time-series fields, clearly defining and solidifying event timestamps, data entry timestamps, and delay indicators, providing traceable input for subsequent relation binding, conflict marking, and gating resolution. Specifically, this includes:
[0049] S200. Perform probe collection on the session packet to generate an evidence package; the evidence package includes an event timestamp, an entry timestamp, and a delay index.
[0050] In the specific implementation of the present invention, the input source of S200 is the session packet generated and written to the session storage unit by S100. The session packet carries a session number, session hash, result input field, data source input field, probe orchestration field, slot template field and mapping field. The probe orchestration field is bound with sampling granularity field, field mapping table, entity dictionary version, data source access identifier, extraction rule identifier and quality verification rule identifier. Specifically, after receiving the evaluation task trigger signal, the evaluation task scheduling unit reads the session packet from the session storage unit and performs a readiness check. The readiness check includes verifying the existence of the session hash field, probing the reachability of the access endpoint description in the data source input field, verifying the validity period of the authorization credential indication field, and validating the evaluation time window index in the sampling granularity field. If the reachability probe fails or the credential verification fails, the evaluation task scheduling unit writes the failure reason, the associated data source access identifier and session number to the session log structure, and triggers the backoff retry process corresponding to the acquisition failure retry policy field. When the number of backoff retry attempts reaches the upper limit of the acquisition failure retry policy field, the evaluation task scheduling unit writes a circuit breaker flag to the session log structure and marks the acquisition task of the session packet as suspended. The suspended state will re-enter the readiness check link when the next round of evaluation task trigger signal arrives.
[0051] After the session packet passes the readiness check, probe acquisition is undertaken by the acquisition execution unit. This acquisition execution unit, along with the data source access unit, field governance unit, quality management unit, and evidence encapsulation unit, forms a serial processing link. Specifically, the data source access unit is responsible for establishing a connection session based on the data source access identifier and access endpoint description; the field governance unit is responsible for converting source fields into unified fields by calling the field mapping table; the quality management unit is responsible for generating missing test markers, anomaly markers, and duplicate markers by calling the quality verification rule identifier; and the evidence encapsulation unit is responsible for encapsulating the processed records into evidence packets. At the beginning of each acquisition round, the data source access unit parses the evaluation time window index from the sampling granularity field and converts it into an acquisition time window description. This acquisition time window description includes a start boundary, an end boundary, and a sliding step size field, with the sliding step size field corresponding to the update frequency description. Understandably, for different data source type identifiers, the data source access unit adopts corresponding access strategies to adapt to the data source publishing method. The minimum set of access strategies includes a batch retrieval strategy based on pagination cursors, an incremental retrieval strategy based on change flags, and a subscription retrieval strategy based on event streams. The incremental retrieval strategy uses the incremental cursor field in the preferred extended fields to record the cursor position of the last successful acquisition, and reads the cursor position as the starting offset at the beginning of the current acquisition round. The cursor position is associated with the session number and written into the session log structure. The subscription retrieval strategy registers the subscription topic and writes it into the subscription identifier field after the connection session is established. The subscription identifier field is associated with the session hash and stored for source tracing of subsequent evidence packets.
[0052] The acquisition execution unit performs parsing and merging processing on the raw records returned by the connection session. The parsing and merging process first parses the text payload field and structured attribute field according to the extraction rules, and then performs merging determination on duplicate returned records within the same acquisition round based on the record primary key. Specifically, the field governance unit inputs the parsed source fields into the field mapping table and generates a unified field list, which is then written into the intermediate evidence record structure. The intermediate evidence record structure includes at least the record primary key, source identifier, text payload field, structured attribute field, event timestamp, and database entry timestamp fields. The event timestamp is obtained by mapping the event occurrence time recorded by the source system, and the database entry timestamp is generated and written by the acquisition execution unit when the record is written to the evidence storage unit. Furthermore, the quality management unit performs quality verification processing on the intermediate evidence record structure. This quality verification processing includes missing test marking, anomaly marking, and duplicate record detection. The missing test marking process, based on the missing test marking rules in the quality verification rule identifier, writes missing test marks to required fields in the unified field list that contain empty values and records the missing test field names. These required fields correspond to the set of required fields in the slot template fields. The anomaly marking process, based on the anomaly threshold rules, writes anomaly marks to time anomalies, numerical anomalies, and format anomalies in the structured attribute fields, and writes the anomaly type code and associated field name into the session log structure. The duplicate record detection process, based on the duplicate record detection rules, writes duplicate marks to records with the same primary key and the same event timestamp combination, and writes the duplicate marks and source identifier together into the intermediate evidence record structure. Understandably, the quality verification process belongs to the core processing set of S200, and its output provides traceable input for the quality mark field of the subsequent evidence package; on this basis, preferably, the quality management unit also loads the source credibility baseline and writes the source credibility baseline into the source confidence field of the intermediate evidence record structure, which is used as an input caliber when S300 scores the credibility of the relationship binding.
[0053] After completing parsing, mapping, and quality verification, the evidence encapsulation unit encapsulates the intermediate evidence record structure into the evidence package. The evidence package is the output of S200 and carries key fields required for connection with subsequent steps. Specifically, the evidence encapsulation unit writes the session number and session hash as evidence package header fields, and writes the source identifier, data source access identifier, and probe identifier into the evidence source field. The probe identifier is generated by the acquisition execution unit based on the probe orchestration field and establishes a mapping relationship with the slot identifier in the slot template field. The evidence encapsulation unit writes the record primary key, text payload field, structured attribute field, and quality mark field into the evidence content field, and writes the event timestamp and entry timestamp into the evidence time sequence field. Furthermore, the evidence encapsulation unit performs a consistency check on the event timestamp and the entry timestamp. This consistency check includes event timestamp missing verification, entry timestamp missing verification, and time sequence verification. When an event timestamp is detected as missing, the evidence encapsulation unit writes a missing flag to the quality flag field and writes the record to the session log structure. This missing flag is retained and transmitted during entity identification and relation extraction in S300. When a time sequence anomaly is detected, the evidence encapsulation unit writes a time anomaly flag to the quality flag field and writes the anomaly type code to the session log structure. The delay index is generated after the evidence encapsulation unit completes the entry action. The delay index is obtained by subtracting the event timestamp from the entry timestamp and is written to the delay index field of the evidence package. The delay index field and the evidence time sequence field together constitute the input basis for the generation of delay flags in the subsequent S400 gating resolution. Understandably, the event timestamp, the database entry timestamp, and the delay index belong to the minimum set of fields in the evidence package. The remaining fields are necessary carrier fields encapsulated for connection with subsequent S300 and S400. Among them, the probe identifier is used as the input position for S300 to bind the relationship probe identifier, and the quality mark field is used as the reference input for S300 to generate credibility score and conflict mark.
[0054] In the engineering implementation, the data source list includes patent data sources, paper data sources, industry dynamic data sources, and investment and financing event data sources. The collection and execution unit starts multiple connection sessions in parallel under the same session number. Among them, the patent data source adopts a batch pull strategy based on pagination cursor to cover the updates of public text within the evaluation time window index; the paper data source adopts an incremental pull strategy based on change markers and records the pull breakpoint with the incremental cursor field; the industry dynamic data source adopts a subscription pull strategy to access the event stream and records the subscription session with the subscription identifier field; and the investment and financing event data source adopts a hybrid strategy of batch pull and incremental pull to take into account both historical completion and real-time updates. The collection and execution unit uniformly enters the field governance unit to map the original records returned by parallel collection into the field mapping table, and performs quality verification processing on the combination of missing event timestamps, abnormal structured attribute field formats, and duplicate records in the quality management unit. Finally, the evidence encapsulation unit generates a unified structured evidence package and writes it into the evidence storage unit. When writing, the evidence storage unit partitions and stores the evidence package according to the session number and evaluation time window index, and writes the successful entry mark into the session log structure. When a write failure occurs, the evidence storage unit writes the failure reason and the primary key of the associated record into the session log structure and triggers the supplementary acquisition process corresponding to the acquisition failure retry strategy field. After the retry is successful, the supplementary acquisition process adds the entry timestamp and recalculates the delay index field.
[0055] Finally, the evidence package output by S200 is written to the evidence storage unit with the session number as the association key, and is read by the evaluation task scheduling unit in the next processing stage as input for S300 to "perform entity recognition, entity alignment, and relation extraction on the evidence package". The text payload field and structured attribute field in the evidence package are the input positions for entity recognition in S300, the probe identifier in the evidence package is the input position for relation binding probe identifier in S300, the quality mark field and source confidence field in the evidence package are the reference inputs for S300 to generate credibility score and conflict mark, and the event timestamp, storage timestamp and delay index in the evidence package are the input basis for S400 to generate delay mark and form review set through gate resolution.
[0056] In summary, this step achieves the following technical results: Under session packet constraints, it completes cross-data source probe collection and unified field mapping, and encapsulates the collected records into an evidence package with time-series fields. This step solidifies the event timestamp, database entry timestamp, and latency metrics into the time-series payload of the evidence package, and writes probe identifiers and quality marker fields into the evidence package to support subsequent relationship binding and gating resolution. This step forms the evidence package and specifies its input positions in S300 and S400, establishing a hierarchical input link from S200 to subsequent steps.
[0057] To address the problem that existing technologies often process entity recognition, entity alignment, and relation extraction separately using a loose pipeline, lacking a mechanism to bind relation records with acquisition probes and slot templates under the same evidence structure and session constraints, resulting in untraceable relation sources, inconsistent credibility standards, and difficulties in expressing conflicts with structured tags, thus leading to a lack of a unified input standard for subsequent gating resolution, this invention completes the concatenated processing of entity recognition, entity alignment, and relation extraction under the evidence package structure and session hash constraints in step S300, encapsulates relation records into candidate packages, and binds slot identifiers and probe identifiers to relations while generating credibility scores and conflict tags, specifically including:
[0058] S300: Perform entity recognition, entity alignment, and relation extraction on the evidence package to generate candidate packages, and bind slot identifiers, probe identifiers, credibility scores, and conflict markers to the relations;
[0059] In a specific implementation of the present invention, the input source of S300 is the evidence package written to the evidence storage unit by S200. The evidence package carries a session number, a session hash, an evidence source field, an evidence content field, and an evidence timing field. The evidence source field includes a source identifier, a data source access identifier, and a probe identifier. The evidence content field includes a record primary key, a text payload field, a structured attribute field, and a quality marker field. The evidence timing field includes an event timestamp, an entry timestamp, and a delay index. Specifically, after the evaluation task scheduling unit is triggered, the candidate generation unit reads the evidence package from the evidence storage unit according to the session number and evaluation time window index, and performs an input integrity check. The input integrity check includes the existence verification of the text payload field, the parsing consistency verification of the structured attribute field, the review of the missing and abnormal markers of the quality marker field, and the association consistency verification of the probe identifier and the session hash. In the event of a missing field or parsing failure, the candidate generation unit writes the failure reason, the primary key of the associated record, the source identifier and the session number into the session log structure, and writes the evidence package into the verification cache structure. The verification cache structure maintains the same caliber as the subsequent S400 generation of the verification set processing link.
[0060] After the evidence package passes the input integrity check, entity recognition is performed by the entity recognition unit, which is a combination of a natural language processing unit and a dictionary constraint unit. The natural language processing unit performs sentence segmentation, word segmentation, part-of-speech tagging, and candidate segment detection on the text payload field, while the dictionary constraint unit performs entity dictionary version constraints and type constraints on the candidate segment detection output. The entity dictionary version has been generated as a core field of the data probe in S100 and transmitted to S200 through the session packet. S200 then inherits it in the evidence packet header field using a session hash association method. Therefore, at the beginning of S300, the entity recognition unit reads the entity dictionary version that matches the session hash from the session storage unit and loads the domain ontology and the synonym merging table accordingly. Specifically, the natural language processing unit performs segmented parsing on the title, summary, and body segments in the text payload field, generating a segment index for each segment and retaining the character position offset. Subsequently, candidate segment detection generates a set of entity candidate segments under the constraints of the segment index. The set of entity candidate segments includes the segment text, segment start and end positions, segment context window, and source identifier fields, where the segment context window is formed by concatenating the windows of adjacent sentences. Further, the dictionary constraint unit performs dictionary matching on the set of entity candidate segments. Dictionary matching includes exact matching, synonym matching, and abbreviation matching. When an abbreviation appears for the first time, a Chinese-English equivalent entry is written and bound to the entity dictionary version. When an English abbreviation appears in the set of entity candidate segments, the dictionary constraint unit reads the abbreviation entry from the synonym merge table and writes it into a comparison field of the full Chinese name and the full English name. This comparison field is written into the candidate package along with the entity recognition result for subsequent auditing. Understandably, the minimum input set for entity recognition is the text payload field and the entity dictionary version, and the minimum output set for entity recognition is the entity candidate set. The entity candidate set includes the entity name, entity type identifier, source identifier, evidence fragment index, fragment start and end positions, and entity confidence field, where the entity confidence field is a composite value of dictionary matching confidence and context consistency confidence. On this basis, preferably, the entity recognition unit also reads the structured name field from the structured attribute field, performs mutual verification on the entity candidate fragment set, and writes a mutual verification conflict flag and retains the conflict evidence fragment index when the mutual verification fails.
[0061] After generating the entity candidate set, entity alignment is handled by an entity alignment unit. This unit is a combination of a name similarity calculation unit, an attribute similarity calculation unit, a semantic vector similarity calculation unit, and an alignment decision unit. Name similarity, attribute similarity, and semantic vector similarity are defined alignment components. Specifically, the name similarity calculation unit performs normalization on the entity names in the entity candidate set. Normalization includes case normalization, full-width / half-width normalization, symbol cleaning, and stop word removal. After normalization, it performs string similarity calculation on the candidate entity pairs to generate name similarity. The attribute similarity calculation unit reads attribute items such as institution identifier, author identifier, publication date, classification number, and investment round identifier from the structured attribute fields and performs attribute consistency comparison on the candidate entity pairs to generate attribute similarity. The attribute consistency comparison includes complete consistency judgment and interval consistency judgment. The interval consistency judgment is used to check the compatibility between date-type attributes and interval-type attributes. The semantic vector similarity calculation unit generates semantic vector representations for the context windows related to entities in the text payload field and performs vector similarity calculation on the candidate entity pairs to generate semantic vector similarity. Furthermore, the alignment decision unit inputs name similarity, attribute similarity, and semantic vector similarity into the alignment score synthesizer to generate alignment scores, and generates an entity alignment mapping table based on the alignment scores and entity type identifier constraints. The entity alignment mapping table includes source entity identifier, target entity identifier, alignment score, alignment basis fragment index, and alignment decision type code. The alignment basis fragment index points to the evidence fragment index in the evidence package. Understandably, the minimum input set for entity alignment is the entity candidate set and the structured attribute fields, and the minimum output set for entity alignment is the entity alignment mapping table. Based on this, preferably, the entity alignment unit performs name disambiguation processing for cases with multiple entities having the same name. Name disambiguation processing generates a disambiguation key based on a combination key of the organization identifier and the publication date and writes it into the alignment decision type code. The disambiguation key is associated with the session number and written into the session log structure.
[0062] After entity recognition and alignment are completed, relation extraction is performed by a relation extraction unit, which is a combination of a pattern extraction subunit and a learning extraction subunit. The pattern extraction subunit performs rule extraction based on the extraction rule identifier, while the learning extraction subunit performs semantic extraction based on the training parameters. The outputs of both are fused and deduplicated in the relation fusion subunit. Specifically, the pattern extraction subunit reads the extraction template version number corresponding to the extraction rule identifier from the evidence source field of the evidence package, and performs relation candidate generation in parallel on the text payload field and the structured attribute field according to the extraction template. The reference field, cooperation field, investment field, and application field in the structured attribute field are directly mapped to relation candidates, and the trigger words in the text payload field are matched with the syntactic dependency structure to generate relation candidates. The learning extraction subunit performs semantic classification on the entity context window aligned by the entity alignment mapping table. The semantic classification outputs the relation type identifier and relation confidence field, and writes the relation evidence fragment index into the relation candidate. Furthermore, the relation fusion subunit performs primary key alignment and deduplication on the relation candidates output by the pattern extraction subunit and the learning extraction subunit. Primary key alignment generates relation primary keys based on a composite key of "subject entity identifier—object entity identifier—relationship type identifier—event timestamp." Deduplication retains candidates with higher relation confidence fields under the same relation primary key and writes the remaining candidates into a conflict candidate cache structure, which is referenced in subsequent conflict marker generation. Understandably, the minimum input set for relation extraction consists of an entity alignment mapping table, text payload fields, structured attribute fields, and extraction rule identifiers. The minimum output set for relation extraction is the relation candidate set, which includes subject entity identifiers, object entity identifiers, relation type identifiers, relation primary keys, relation confidence fields, event timestamps, source identifiers, and relation evidence fragment indexes.
[0063] After generating the candidate set of relationships, a candidate package is generated by the candidate encapsulation unit. The candidate package is the output of S300, and within the candidate package, the relationship is bound with slot identifiers, probe identifiers, credibility scores, and conflict flags. Specifically, the slot binding unit reads the slot template field and mapping field that match the session hash from the session storage unit, and inputs the relationship type identifier from the candidate set into the slot mapping relationship to obtain the slot identifier and write it into the slot identifier field of the relationship record. The slot identifier field is consistent with the slot identifier namespace generated by S100. The probe binding unit reads the probe identifier of the evidence package and performs a consistency check with the slot identifier field. If the consistency check passes, the probe identifier is written into the probe identifier field of the relationship record. If the consistency check fails, the reason for the inconsistency and the primary key of the associated record are written into the session log structure, and the relationship record is written into the cache structure to be reviewed. Furthermore, the credibility score is generated by a credibility scoring unit. This unit uses the credibility thresholds in the relation confidence field, quality label field, source confidence field, and slot template field as inputs to generate the credibility score. The generation process includes quality downweighting, source baseline fusion, and threshold comparison. The quality downweighting applies a downweighting coefficient to the relation confidence field based on missing, anomaly, and duplicate labels. The source baseline fusion combines the source confidence field with the downweighted relation confidence field. The threshold comparison compares the combined value with the credibility threshold and writes it into the credibility score field. The credibility score field is retained and passed as one of the core inputs for subsequent S400 gating resolution. Meanwhile, conflict markers are generated by a conflict marker unit. This unit takes the conflict type code, conflict candidate cache structure, and relation primary key as input to generate a conflict cluster identifier, a conflict type code, a conflict density value, and a primary slot identifier. The conflict cluster identifier is used to aggregate multiple relation candidates under the same subject entity identifier and the same object entity identifier. The conflict density value is calculated by combining the number of candidates within the same conflict cluster with the number of identifiers from different sources. The primary slot identifier is written by the slot identifier field. The conflict marker field is written to the relation record and output with the candidate package, providing an input for the subsequent S400 to perform gating resolution to generate an admission package and a review set.
[0064] In an engineering embodiment, for a patent achievement registered in the achievement management terminal, the evidence storage unit outputs a public text evidence package from the patent data source, a paper abstract evidence package from the paper data source, and an announcement evidence package from the investment and financing event data source under the same session number. The candidate generation unit uniformly performs entity recognition on the above evidence packages and generates an entity candidate set under the constraint of entity dictionary version. The entity alignment unit performs name similarity, attribute similarity, and semantic vector similarity to synthesize alignment scores for multiple types of entity candidates of "achievement entity - patent entity - enterprise entity - investment and financing event entity" and constructs an entity alignment mapping table. The relationship extraction unit extracts investment relationships from the announcement evidence package, cooperation relationships from the paper abstract evidence package, and citation and application relationships from the patent public text evidence package under the constraint of extraction rule identifier. Subsequently, the candidate encapsulation unit binds slot identifiers to each relationship record according to the mapping field and inherits the probe identifiers in the evidence package. The credibility scoring unit generates a credibility score by combining the quality mark field and the source confidence field. The conflict marking unit generates a conflict cluster identifier and conflict density value according to the conflict type code and writes it into the main slot identifier. Finally, a candidate package is generated and written into the candidate storage unit.
[0065] The candidate package output by S300 includes an entity alignment mapping table and a set of relationship records formed by entity recognition, entity alignment and relationship extraction. The set of relationship records carries a slot identifier field, a probe identifier field, a credibility score field and a conflict flag field. The candidate package is written to the candidate storage unit according to the session number and is read by the evaluation task scheduling unit in the next processing stage as input for S400 to "perform gating resolution on the candidate package". The credibility score field, conflict flag field, event timestamp, entry timestamp and delay index together constitute the input basis for S400 to generate the admission package and the review set.
[0066] This step's technical effects can be summarized as follows: Under the constraints of the evidence packet structure and session hashing, this step completes the concatenated processing of entity recognition, entity alignment, and relation extraction, and encapsulates the relation records into candidate packets. This step binds slot identifiers and probe identifiers to relation records and generates credibility scores and conflict markers, forming a unified input standard for subsequent gating resolution. This step outputs candidate packets and indicates their input positions in S400, establishing a hierarchical input link from S300 to S400.
[0067] To address the problems in existing technologies where relation extraction results are often directly written into the graph without a linked gating mechanism based on latency indicators, slot template reliability thresholds, and conflict type codes, leading to high-latency, low-reliability, or conflicting relations being mixed into the graph, uncontrollable version evolution, and a lack of verification entry points, this invention, through step S400, performs gating judgment and conflict resolution on candidate packets under the constraints of slot templates and latency indicators. Passed relation records are encapsulated into admission packets, and blocked or manually verified records are written into a verification set. Specifically, this includes:
[0068] S400. Perform gated resolution on the candidate packets to generate admission packets and generate a verification set;
[0069] In a specific implementation of this invention, the input source for S400 is the candidate packet written to the candidate storage unit by S300. The candidate packet encapsulates an entity alignment mapping table and a set of relationship records under session number and session hash constraints. Each relationship record carries a slot identifier field, a probe identifier field, a credibility score field, and a conflict flag field. Simultaneously, it retains a traceable index of the event timestamp, entry timestamp, and delay index on the evidence packet side. Specifically, the gating resolution unit consists of a gating rule loading subunit, an admission determination subunit, a conflict resolution subunit, a delay flag subunit, a review generation subunit, and an admission encapsulation subunit. The evaluation task scheduling unit starts S400 when it detects a candidate packet in the candidate storage unit that matches the session number and meets the scheduling trigger conditions. These scheduling trigger conditions include a candidate packet arrival flag, a session log structure readability flag, and a slot template matching the session hash being in an active version flag. After the gating resolution unit starts, it first performs input verification processing, which includes verifying the existence of the slot identifier field and the slot template, verifying the access consistency of the probe identifier field and the data probe, verifying the numerical field of the credibility score field, and verifying the legality of the conflict type code in the conflict marker field. In the event of a verification failure, the gating resolution unit writes the failure reason, the association primary key, the session number and the session hash into the session log structure, and writes the corresponding relationship record into the review buffer. The review buffer is one of the sources of the subsequent review set.
[0070] After input validation passes, the gating rule loading subunit reads the slot template matching the session hash from the session storage unit, and extracts the credibility threshold, conflict type code, and required field set from the slot template to form a gating parameter group. Simultaneously, the gating rule loading subunit reads the evidence package index associated with the candidate package from the evidence storage unit, and uses the event timestamp, entry timestamp, and delay index in the evidence package as the timing constraint input for gating resolution. Understandably, the minimum input set for this step consists of the candidate package, the slot template, and the delay index that is traceably associated with the candidate package. The delay index is obtained by subtracting the event timestamp from the entry timestamp and is already embedded within the evidence package. Therefore, at the beginning of S400, the gating resolution unit reads the delay index and compares it with the delay threshold in the gating parameter group to obtain a delay marker. Specifically, the delay tagging subunit generates a delay tagging field for each relation record. The delay tagging field contains a delay level code and a delay source identifier. The delay level code is obtained by segmented mapping between delay indicators and delay thresholds. The delay source identifier is taken from the probe identifier field and associated with the source-side index of the data source access identifier. The delay tagging field is written back to the relation record and enters the subsequent conflict resolution and admission determination link. At the same time, the relation record with the delay tagging field is written into the candidate queue of the review set. At the end of this step, the candidate queue of the review set is merged with other review sources to generate the final review set.
[0071] The admission determination subunit performs gating and filtering on the set of relation records. The core of the gating and filtering process is the comparison of credibility thresholds and the comparison of required field sets. Specifically, the admission determination subunit locates the corresponding entry in the slot template according to the slot identifier field, reads the credibility threshold and required field set bound to the entry, and performs threshold comparison on the credibility score field in the relation record. For relation records that do not meet the credibility threshold, a credibility failure mark is generated and written into the candidate queue of the review set. For relation records that meet the credibility threshold, the required field set comparison is performed. The required field set comparison includes the existence verification of the subject entity identifier, object entity identifier, relation type identifier, relation evidence fragment index and event timestamp, and the combination consistency verification of the probe identifier field and the slot identifier field. If the verification fails, a field missing mark is generated and written into the candidate queue of the review set. Furthermore, the admission determination subunit generates candidate admission tags for relation records that have passed the credibility threshold comparison and the mandatory field set comparison. The candidate admission tag fields include admission reason code, gating version tag and session hash. The gating version tag is extracted from the effective version of the slot template and bound to the session number and written into the relation record. The candidate admission tag fields are one of the input conditions for the subsequent conflict resolution subunit.
[0072] The conflict resolution subunit performs conflict cluster-level resolution processing on relation records with candidate admission flag fields. The conflict clusters are defined by conflict cluster identifiers in the conflict flag field, which includes the conflict cluster identifier, conflict type code, conflict density value, and primary slot identifier. Specifically, the conflict resolution subunit aggregates relation records according to the conflict cluster identifier to form an intra-cluster candidate set, and selects the corresponding resolution strategy entry from the intra-cluster candidate set according to the conflict type code. The resolution strategy entry is part of the gating parameter set loaded by the gating rule loading subunit, and includes primary slot priority rules, source priority rules, credibility score sorting rules, and timestamp consistency rules. Furthermore, the primary slot priority rule assigns priority tags to relationship records where the primary slot identifier and slot identifier fields are consistent; the source priority rule assigns source weight tags to the data source access identifier corresponding to the probe identifier field; the credibility score sorting rule sorts according to the credibility score field and generates a sorting index; and the timestamp consistency rule verifies the consistency between the event timestamp and the entry timestamp and generates a time-series conflict tag for relationship records with reverse order. After the above tags are generated, the conflict resolution subunit performs a selection process on the candidate set within the cluster. The selection process combines the priority tags, source weight tags, sorting index, and time-series conflict tags into a resolution decision record. The resolution decision record is written into the relationship record and forms a resolution result tag field. Understandably, the minimum input set for conflict resolution consists of the conflict cluster identifier, conflict type code, and credibility score field. The minimum output set for conflict resolution is the resolution result tag field and the set of eliminated relationship records. The set of eliminated relationship records is written into the candidate queue of the review set, carrying the resolution decision record, for easy subsequent review and tracking. Furthermore, if there are candidates with the same credibility score field and the same source weight label in the candidate set within the cluster, the conflict resolution subunit performs secondary discrimination on the candidates. The secondary discrimination is based on the quality label field in the evidence package backtracked by the relation evidence fragment index, and generates a quality discrimination label according to the combination of anomaly label and duplicate label. The quality discrimination label is written into the resolution decision record.
[0073] The review generation subunit performs review set construction processing on the candidate queue of the review set. This process includes review trigger condition determination, review entry encapsulation, and review routing information generation. Specifically, the review trigger condition determination merges and deduplicates the sources of trustworthiness failure markers, missing field markers, timing conflict markers, delay markers, and eliminated relation record sets within the candidate queue. For each review entry, a review type code and a review priority code are generated. The review priority code is synthesized from the conflict density value, delay level code, and deviation magnitude of the trustworthiness score field. The review entry encapsulation process encapsulates the session number, session hash, relation primary key, slot identifier field, probe identifier field, trustworthiness score field, conflict marker field, and resolution decision record into a review entry structure. This structure is then written into the review set, which is written to the review storage unit using the session number as the partition key. Furthermore, the route information for review is generated based on the mapping relationship between the probe identifier field and the data source access identifier. The generation timestamp of the review set and the version mark of the review set are recorded in the session log structure. The version mark of the review set and the gated version mark are kept in the same namespace.
[0074] The admission encapsulation subunit performs admission package encapsulation processing on relation records that have passed gating and conflict resolution and have a "reserved" status in the resolution result marker field, thus obtaining the admission package. Specifically, the admission encapsulation subunit uses the session number and session hash as header fields, writes the associated subset of the set of relation records in the reserved status and the entity alignment mapping table into the admission package body field, and generates an admission index in the admission package body field. The admission index includes a slot identifier field index, a relation type identifier index, and a time window index, wherein the time window index is mapped from the event timestamp and maintains alignment with the change time window index in the differential package of subsequent S500 graph updates. At the same time, the admission encapsulation subunit retains the delay marker field in the admission package body field and writes the gating version marker and session hash into the admission package header field, forming an auditable admission package structure. Understandably, the minimum output field set of the admission package includes a session hash, a gated version flag, a set of relationship records, and an admission index. After the admission package is written to the admission storage unit, it is read by the evaluation task scheduling unit and used as input for S500 to "perform a graph update on the admission package". The time window index in the admission index is entered into S500 to drive the generation of the change time window index of the differential package. The slot identifier field and probe identifier field in the admission package body field are entered into S500 for the source tracing and registration of nodes and relationships. At the same time, the review set is pushed to the review processing channel by the scheduling unit and associated with the session log structure. When the status of the review entry changes, the review processing channel writes back the review result flag and realizes cross-round tracking through the session hash in subsequent iteration sessions.
[0075] In the engineering implementation, for multi-source evidence input of the same result entity, the candidate storage unit outputs a candidate package containing investment relationships, cooperation relationships, and citation relationships. The gating resolution unit reads the slot template matching the session hash and loads the credibility threshold and conflict type code. For the relationship records in the candidate package, a delay mark field is first generated, and then the threshold entries are located according to the slot identifier field to perform credibility threshold comparison. When a relationship record from a certain investment and financing event generates a delay level code due to storage delay and triggers the delay mark field, the relationship record is written into the candidate queue of the review set. When a citation relationship has two candidates from different probe identifiers in the conflict cluster and the conflict density value is high, the conflict resolution subunit sorts according to the source weight mark and credibility score field and generates a resolution decision record, retains one, and writes the set of eliminated relationship records into the review set. The remaining relationship records that meet the threshold and are retained are encapsulated into an admission package by the admission encapsulation subunit and an admission index is generated. The time window index and slot identifier field index in the admission index are used in S500 for differential package generation and graph source registration.
[0076] Finally, the admission packet output by S400 is a structured input carrier after gating resolution. The admission packet contains a session hash, gating version tag, relationship record set and admission index, and is written into the admission storage unit for S500 to call. At the same time, the review set output by S400 contains a review entry structure and carries a review type code, review priority code and resolution decision record. It is written into the review storage unit and associated with the session log structure to form a cross-step audit link.
[0077] This step's technical effects can be summarized as follows: Under the constraints of slot templates and latency metrics, this step performs gating and conflict resolution on candidate packets, and encapsulates the passed relationship records into admission packets. This step writes latency markers and conflict resolution decision records into a review set, and then archives the review set along with the session hash. This step outputs admission packets and specifies their input positions in the S500, forming a hierarchical input chain from candidate packets to admission packets.
[0078] To address the problem that existing graph updates often lack a version closed loop of "differential package generation—differential application—snapshot hash index," leading to difficulties in tracing the update process, accurately locating the introduction time window and version source of a certain relationship, and consequently affecting the auditability of subsequent subgraph extraction, path construction, and trend calculation, this invention completes differential package generation and differential application under the constraints of the admission package and baseline snapshot pointer in step S500, and generates an updated graph snapshot. Simultaneously, hash calculations are performed on the graph snapshot and differential package to form a snapshot hash index, which is then packaged into a version package. Specifically, this includes:
[0079] S500: Perform graph update on the admission package to generate a version package; the version package includes a graph snapshot, a snapshot hash, and a differential package;
[0080] In a specific implementation of the present invention, the input source of S500 is the admission packet written to the admission storage unit by S400. The admission packet encapsulates a set of relation records after gating and resolution under the constraints of session number and session hash, and retains slot identifier field, probe identifier field, credibility score field and delay flag field in the relation records. At the same time, it carries an admission index, which includes slot identifier field index, relation type identifier index and time window index. The time window index is obtained by mapping the event timestamp and maintains an alignment caliber with the change time window index in the subsequent differential packet. Specifically, the graph update is completed collaboratively by the graph access unit, version management unit, change calculation unit, hash calculation unit, rollback control unit, and write commit unit. The graph access unit is used to access the existing graph versions in the graph storage unit. The version management unit is used to maintain the graph version chain and snapshot hash index. The change calculation unit is used to parse the admission package and generate a differential package. The hash calculation unit is used to generate a snapshot hash for the graph snapshot and the differential package. The rollback control unit is used to handle commit failures and consistency verification failures. The write commit unit is used to write the version package to the version storage unit and register a traceable index.
[0081] In the initial stage of S500, the evaluation task scheduling unit initiates the graph update process when it detects an admission packet matching the session number in the admission storage unit and meets the update triggering conditions. These update triggering conditions include an admission packet arrival flag, a session log structure being in a writable state flag, and a graph access unit completing a locking flag. Specifically, the graph access unit reads the currently effective graph version and generates a baseline snapshot pointer, which is associated with the snapshot hash index of the currently effective graph snapshot. Simultaneously, the version management unit reads the session hash and establishes an update transaction record. This update transaction record includes the session number, session hash, baseline snapshot pointer, and transaction timestamp, and is written to the session log structure as a reference for subsequent rollback control units. Understandably, the minimum input set for this step consists of the admission packet, the baseline snapshot pointer, and the update transaction record. The admission packet provides a set of relationship records and an admission index, the baseline snapshot pointer provides the location entry point for the current graph snapshot, and the update transaction record provides a session-level audit entry point.
[0082] The change calculation unit performs structured parsing processing on the admission package. This structured parsing processing includes admission index expansion, relation record normalization, entity location, and change candidate generation. Specifically, the change calculation unit groups relation records based on the slot identifier field index in the admission index, further buckets the relation records within each group based on the relation type identifier index, and forms a time window queue consistent with the event timestamp based on the time window index. Within each time window queue, the change calculation unit performs normalization processing on the relation records. This normalization processing includes unifying the entity identifier caliber, performing dictionary mapping on the relation type identifier, expanding and registering the association between the probe identifier field and the data source access identifier, retaining the credibility score field and the delay flag field, and generating source metadata. The source metadata is written into the change candidate record and associated with the slot identifier field. Furthermore, the change calculation unit calls the baseline snapshot pointer to access the graph storage unit and locates the subject entity identifier and object entity identifier associated with the relationship record in the graph snapshot. The location process is completed in the entity index subunit. The entity index subunit maintains a mapping cache from entity identifier to node address. When the mapping cache is not hit, the graph access unit is triggered to perform disk index retrieval and backfill the cache. If the node address has been obtained, the change calculation unit generates a change candidate set. The change candidate set provides a candidate opcode for each relationship record. The candidate opcode includes one of three types: addition, deletion, and weight change. The candidate opcode, along with the relationship primary key, node address, relationship type identifier, slot identifier field, probe identifier field, credibility score field, and delay flag field, is written into the change candidate record.
[0083] After the change candidates are generated, the change calculation unit performs differential package construction processing. This process takes the change candidate records as input and outputs the differential package, which includes a node addition set, a node deletion set, a relationship addition set, a relationship deletion set, a weight change set, and a change time window index. Specifically, the generation of the node addition set is driven by the associated subset of the entity alignment mapping table. The change calculation unit reads the entity alignment mapping table entries associated with the relationship record set from the admission package body field, identifies entity identifiers not present in the graph snapshot, and generates node addition records. Each node addition record includes an entity identifier, an entity type identifier, source metadata, and a write time window index. In this embodiment, the node deletion set is constrained by the rollback control unit and the lifecycle policy. The lifecycle policy generates candidate deletion records using failure markers in the session log structure as input, and these records are only included in the node deletion set after consistency verification. The relationship addition set is formed by extracting candidate records for addition with candidate opcodes. Each new relationship record includes a subject entity identifier, object entity identifier, relationship type identifier, slot identifier field, probe identifier field, credibility score field, delay flag field, and write time window index. The relationship deletion set is formed by candidate records for deletion with candidate opcodes, retaining the original relationship primary key and deletion reason code in the deleted records. The weight change set is formed by candidate records for weight changes with candidate opcodes, containing the original weight value index, new weight value index, change source metadata, and write time window index. Further, the change time window index is obtained by merging and deduplicating the time window index queue, and a time window identifier consistent with the event timestamp caliber is written into the differential packet. Simultaneously, an association flag with the session hash is registered in the change time window index for subsequent updates of the trend sequence in S800.
[0084] After the differential package is constructed, the write-commit unit performs pre-commit verification on the differential package. The pre-commit verification includes relation consistency verification, referential integrity verification, and conflict review association verification. Specifically, relation consistency verification verifies the existence of subject entity identifiers and object entity identifiers in the node addition set and graph snapshot in the relation addition set and relation deletion set, and verifies the legality of relation type identifiers in the relation type dictionary; referential integrity verification verifies the alignment between the evidence fragment index retained in the relation record and the index table of the session log structure; conflict review association verification checks the association between the delayed flag field and the review entry structure of the review set. If a delayed flag field exists and a review entry structure with the same relation primary key already exists in the review set, a review association flag is written for the new record of that relation in the differential package, and the write priority code is reduced. Understandably, the above verifications are optional extended functions of this step. The minimum set of verifications necessary to implement the core improvement of this step are relation consistency verification and reference integrity verification. These two are used to constrain the differential package to have executable write conditions when it is implemented in the project. Conflict review association verification is a preferred function, used to retain a traceable entry point for the review link in the write link.
[0085] After the pre-commit verification passes, the graph access unit performs differential application processing, applying the differential packet to the graph snapshot pointed to by the baseline snapshot pointer to generate an updated graph snapshot. Specifically, the differential application processing is performed in the graph update buffer, which is an isolated write area outside the graph storage unit, used to complete the merged writing of node addition, relation addition, relation deletion, and weight change before submission. In the node addition stage, the graph access unit assigns a node address to the new node record and writes it to the entity index subunit, while registering the source set and session hash of the probe identifier field in the node metadata. In the relation writing stage, the graph access unit assigns an edge address to the new relation record according to the relation type and writes it to the edge index subunit, while registering the slot identifier field, credibility score field, and delay flag field in the edge metadata. In the relation deletion stage, the graph access unit performs mark deletion on the relation deletion record and writes the deletion reason code to the edge metadata. In the weight change stage, the graph access unit updates the edge weight field of the weight change record and registers the change source metadata. Furthermore, after the differential application processing is completed, the graph access unit generates an updated graph snapshot pointer and returns it to the version management unit. The version management unit binds the updated graph snapshot pointer to the update transaction record to form a version candidate record.
[0086] After the version candidate record is generated, the hash calculation unit performs hash calculation processing on the updated graph snapshot and differential packet to obtain the snapshot hash. Specifically, the snapshot hash is a hash digest generated by the node segment, edge segment, and index segment pointed to by the snapshot pointer of the graph snapshot according to a predetermined serialization caliber. The hash digest and the session hash are jointly written into the snapshot hash index of the version management unit. At the same time, the hash calculation unit generates a differential hash for the differential packet and writes it into the differential packet metadata field. The differential hash and the snapshot hash form a consistency check item, which is used by the rollback control unit to locate the rollback boundary when a commit failure occurs. The hash appearing for the first time in this embodiment refers to the digest value generated by the secure hash function (Secure Hash Algorithm). Both the snapshot hash and the differential hash are generated according to the output caliber of the secure hash algorithm and stored in the metadata field of the version packet in hexadecimal string form. Understandably, the snapshot hash is one of the smallest output fields in this step, used for the subsequent S600 to locate the graph snapshot of the subgraph extraction, and also serves as a key index field of the session-level audit link.
[0087] The write-commit unit performs version package encapsulation and commit processing to form the version package and write it to the version storage unit. Specifically, the version package consists of a version package header field and a version package body field. The version package header field includes the session number, session hash, snapshot hash, baseline snapshot hash index, transaction timestamp, and version sequence number. The version package body field includes the graph snapshot, differential package, and differential hash, where the graph snapshot is the serialized image of the updated graph snapshot in the version storage unit, and the differential package is the differential package generated in this step. During commit, the write-commit unit performs atomic commit processing on the version package. Atomic commit processing includes writing the pre-write log, writing the version package to disk, updating the snapshot hash index, and writing the transaction commit flag. If a failure occurs during the version package to disk stage, the rollback control unit reads the pre-write log and undoes the changes to the graph update buffer, while writing the failure flag, session number, and session hash to the session log structure. If a failure occurs during the snapshot hash index update stage, the rollback control unit triggers the index repair process and writes the repair record to the session log structure. Furthermore, after a successful commit, the write commit unit sets the updated graph snapshot pointer as the new effective pointer and registers the mapping relationship between the version number and the snapshot hash in the version management unit, thus forming a version chain.
[0088] In an engineering embodiment, for a scenario where a certain result entity adds new cooperative relationships and investment and financing relationships within a session cycle, the admission package is read by the evaluation task scheduling unit and triggers a graph update. The change calculation unit groups the results entity according to the slot identifier field index and locates the corresponding node address of the result entity in the graph snapshot, generating a differential package containing the set of new relationships and the set of weight changes, and writing the change time window index obtained by mapping the event timestamp into the differential package. The graph access unit applies the differential package to the graph update buffer to generate an updated graph snapshot. The hash calculation unit generates a snapshot hash and writes it into the snapshot hash index. The write submission unit encapsulates the version package and submits it to the version storage unit. The graph snapshot and differential package in the version package body field are used as input by the subsequent S600 to locate the result entity in the graph snapshot and extract the evaluation subgraph.
[0089] The version package output by S500 includes a graph snapshot, a snapshot hash, and a differential package. The graph snapshot and snapshot hash are used for graph snapshot location and version selection in the "subgraph extraction" stage of S600. The change time window index in the differential package is used to trigger the update of the trend sequence of concern in subsequent links. At the same time, the session hash and snapshot hash in the version package header field are written into the session log structure and associated with the update transaction record, forming a cross-step version traceability entry.
[0090] In summary, this step achieves the following technical results: Under the constraints of the admission packet and baseline snapshot pointer, this step generates and applies the differential packet, and produces an updated graph snapshot. This step performs hash calculations on the graph snapshot and the differential packet to form a snapshot hash index, supporting traceable recording of the version chain. This step outputs the version packet and indicates its input position in the S600, forming a hierarchical input chain from the admission packet to the version packet.
[0091] To address the problem that existing technologies often calculate results directly on the entire graph or perform only local statistics based on static relationships, lacking a mechanism for extracting evaluation subgraphs centered on the results entity and linked by relationship type and time window under version package constraints, leading to evaluation range drift, uncontrollable subgraph boundaries, and difficulty in preserving traceable associations between slots and probes, thus affecting subsequent link construction and evidence interpretation, this invention completes graph snapshot loading and results entity positioning under version package constraints in step S600, and forms a linked extraction link based on relationship type, time window, and confidence threshold, assembling the evaluation subgraph and encapsulating it into a subgraph package, specifically including:
[0092] S600. Perform subgraph extraction on the version package, locate the result entity in the graph snapshot, extract and evaluate subgraphs according to relation type, time window and confidence threshold, and generate subgraph package;
[0093] In a specific implementation of the present invention, the input source of S600 is the version package written to the version storage unit by S500. The version package includes a graph snapshot, a snapshot hash, and a differential package, and associates the session number and session hash in the header field of the version package. The snapshot hash is used as the retrieval key of the graph snapshot to locate the corresponding graph snapshot. At the same time, the differential package contains a change time window index and associates the event timestamp caliber, which is used to drive the time window selection and extraction boundary determination in this step. Specifically, the S600 is completed collaboratively by a version reading unit, a result location unit, an extraction rule unit, a credibility filtering unit, a time window pruning unit, a subgraph assembly unit, and a subgraph verification unit. The version reading unit is used to parse the version package and load the graph snapshot and the difference package. The result location unit is used to locate the result entity in the graph snapshot. The extraction rule unit is used to construct the extraction strategy based on the relationship type. The credibility filtering unit is used to perform edge filtering based on the credibility threshold. The time window pruning unit is used to perform windowed pruning on nodes and relationships based on the time window. The subgraph assembly unit is used to generate the evaluation subgraph and encapsulate it into a subgraph package. The subgraph verification unit is used to perform consistency verification on the subgraph package and record the structure of abnormal entries.
[0094] In the initial stage of S600, the version reading unit takes the version packet as input, performs consistency verification on the header field of the version packet, and loads the graph snapshot and the differential packet. The consistency verification includes at least a match check between the snapshot hash and the graph snapshot, and a match check between the session hash and the header field of the version packet. When the verification fails, the version reading unit writes a failure flag to the session log structure and adds the version packet to the system-side review queue of the review set, without proceeding to subsequent extraction links. Further, the version reading unit parses the change time window index in the differential packet to obtain a candidate time window set and writes it to the window selection cache of the time window pruning unit. The window selection cache includes a time window identifier, window start and end boundaries, and a window source flag, where the window source flag distinguishes windows from those from the change time window index from those from the evaluation task configuration. Understandably, the minimum input set for subgraph extraction in this step is the graph snapshot, snapshot hash, change time window index, and result identifier. The graph snapshot and snapshot hash provide a consistent graph basis, the change time window index provides the window clipping boundary, and the result identifier is used to locate the result entity in the graph snapshot. Other features, such as extraction rule extension and subgraph verification enhancement, are preferred extension functions.
[0095] The result localization unit performs result entity localization in the graph snapshot. The result entity is a graph node corresponding to the result identifier. The result identifier can be passed in by the upper-level evaluation task under the session number constraint and is traceable in the session log structure. Specifically, the result localization unit reads the result identifier and calls the entity index subunit to obtain the node address. The entity index subunit maintains the mapping relationship from the entity identifier to the node address. When the mapping relationship does not exist, the result localization unit triggers the candidate localization process. The candidate localization process takes the alignment entries in the entity alignment mapping table as input, performs threshold judgment on the alignment score synthesized from name similarity, attribute similarity, and semantic vector similarity, and selects a candidate entity identifier and retryes the entity index retrieval. When the candidate localization still fails, the result localization unit writes the result identifier into the missing entity record of the subgraph verification unit and associates the missing entity record with the session hash and snapshot hash. Then, it outputs an empty subgraph packet and ends the processing of this result instance in this step. Further, after successful localization, the result localization unit generates a result localization record. The result localization record includes the result identifier, the result entity node address, the entity type identifier, and the localization timestamp. The result entity node address is used as the input node for the subsequent extraction rule unit.
[0096] After the result entity is located, the extraction rule unit generates an extraction strategy based on the slot template and the evaluation task configuration. The slot template is generated in stage S100 and stored in the session log structure using a session hash index. In this step, the extraction rule unit reads the slot template that matches the session hash and parses the set of relation types and the set of entity types in the slot template as prior constraints for relation type filtering. Specifically, the extraction strategy includes at least a relation type whitelist and a relation expansion depth. The relation type whitelist is obtained by the intersection of the relation type set in the slot template and the relation type set configured in the evaluation task. The relation expansion depth is used to limit the hop count boundary for expansion from the result entity. In this embodiment, the relation expansion depth can be set to one hop or multiple hops, and different expansion depths can be used for different relation types. Furthermore, the extraction rule unit generates extraction rule entries for each relation type. These entries include a relation type identifier, a direction constraint code, a time window reference flag, and a credibility threshold reference flag. The direction constraint code indicates whether the expansion uses the result entity as the subject or object entity; the time window reference flag indicates which window set in the window selection cache to use; and the credibility threshold reference flag indicates the credibility threshold configured in the credibility filtering unit to be invoked. Understandably, the relation type whitelist and credibility threshold are key parameters for the core improvement in this step, representing the minimum set of essential parameters for evaluating subgraph extraction. The direction constraint code, relation expansion depth, and time window reference flag are preferred parameters used to adapt to the relation semantics and event lag characteristics of different data sources during engineering implementation.
[0097] After the extraction strategy is generated, the credibility filtering unit and the time window pruning unit perform joint extraction on the graph snapshot. Specifically, the credibility filtering unit retrieves the candidate edge set associated with the address of the result entity node in the edge index sub-unit of the graph snapshot. Each edge in the candidate edge set retains a slot identifier field, a probe identifier field, a credibility score field, and a delay marker field. The credibility filtering unit reads the credibility threshold reference marker in the extraction rule entry and obtains the credibility threshold. It performs threshold determination on the credibility score field. Candidate edges below the credibility threshold are written into the low credibility rejection record and associated with the statistical field of the conflict type code. At the same time, for candidate edges with a delay marker field, the credibility filtering unit writes the delay marker field into the delay relationship queue and retains its probe identifier field for subsequent verification set entry association. Furthermore, the time window pruning unit takes the candidate time window set in the window selection cache as input and performs time window pruning on the candidate edges that have passed the confidence screening. The time window pruning is based on the event timestamp in the edge metadata and refers to the start and end boundaries of the window of the changed time window index. Candidate edges located outside the window are written into the window removal record and retain their relation type identifier and slot identifier fields. At the same time, the time window pruning unit performs node pruning on the nodes connected to the candidate edges. Node pruning is based on the first appearance timestamp and the most recent update timestamp of the node. If the node does not overlap with the window, the node is marked as an outside node and is not included in the evaluation subgraph during subgraph assembly. Understandably, this step uses the confidence threshold and time window as dual gating conditions, which together constitute the essential constraint set for extracting the evaluation subgraph in this step. The confidence threshold constraint comes from the retention of the confidence score field in the candidate package during graph update, and the time window constraint comes from the parsing result of the changed time window index in the difference package. Both are consistent with the cross-step approach of this invention.
[0098] The subgraph assembly unit writes nodes and relationships with dual gating into the evaluation subgraph and encapsulates them into the subgraph package. Specifically, the subgraph assembly unit constructs the subgraph skeleton with the address of the result entity node as the central node, and adds candidate edges to the subgraph edge set according to the whitelist of relationship types determined by the extraction rules. At the same time, it adds the nodes associated with the subgraph edge set to the subgraph node set. During the addition process, the subgraph assembly unit retains the relationship type identifier, slot identifier field, probe identifier field, credibility score field, and delay flag field for each subgraph edge, and writes the writing time window index of the weight change record or relationship addition record associated with the edge in the differential package into the change time window index field of the subgraph edge, thereby establishing a traceable association with the differential package within the subgraph. Furthermore, the subgraph assembly unit performs deduplication on the subgraph edge set. Deduplication is performed by aggregating edges based on the relation primary key and conflict cluster identifier. When multiple candidate edges exist for the same relation primary key, edges with higher confidence scores and no conflict markers are prioritized for inclusion in the evaluation subgraph. For edges with conflict markers, the subgraph assembly unit writes them into the conflict bypass segment of the subgraph package and retains the conflict cluster identifier and conflict type code for bypass reference during subsequent S700 link construction phases. Understandably, the subgraph package is the output of this step. The subgraph package contains at least a subgraph node set, a subgraph edge set, a result location record, and a window selection cache summary. The subgraph node set and subgraph edge set constitute the evaluation subgraph, the result location record identifies the central result entity of the evaluation subgraph, and the window selection cache summary identifies the time window boundary criteria used for this extraction.
[0099] After the subgraph package is generated, the subgraph verification unit performs consistency checks and records abnormal entries. Specifically, the consistency check includes at least the reference consistency check between the subgraph node set and the subgraph edge set, the dictionary validity check of the slot identifier field and the relation type identifier, and the traceability check of the probe identifier field in the subgraph edge set. The traceability check is implemented through the probe registration table in the session log structure. When it is found that the probe identifier field cannot be traced back, the subgraph verification unit writes the edge into a non-traceable record and writes a verification mark field into the subgraph package. The verification mark field is used for selective skipping or downgrading of the edge in the subsequent S700 link construction stage. Further, the subgraph verification unit performs correlation verification between the relation primary key in the delayed relation queue and the verification entry structure of the verification set. If the verification entry structure is missing, a verification missing record is written and the record is associated with the session hash and snapshot hash as a session-level traceability entry. This correlation verification is a preferred extended function and is used in this embodiment to extend the delayed mark field to subsequent links.
[0100] In an engineering implementation, for a scenario where a certain achievement entity experiences an increase in the number of investment and financing event nodes and the simultaneous occurrence of negative event nodes within a few recent time windows, the version reading unit loads the graph snapshot based on the snapshot hash and parses the differential packet to obtain the change time window index. The achievement positioning unit locates the achievement entity node address in the graph snapshot and generates an achievement positioning record. The extraction rule unit generates a whitelist of relationship types and extraction rule entries based on the relationship type set in the slot template. The credibility filtering unit filters out low-credibility candidate edges based on the credibility threshold. The time window trimming unit trims candidate edges based on the change time window index and outputs the nodes and relationships within the window. The subgraph assembly unit encapsulates the nodes and relationships within the window into a subgraph package and retains the slot identifier field, probe identifier field, credibility score field, and delay marker field. The subgraph verification unit completes the consistency verification and writes the verification marker field into the subgraph package. The final output subgraph package serves as the input for S700's "perform link construction on the subgraph package". The relation type identifier and change time window index fields in the subgraph edge set are used to select candidate paths for traction chains and risk chains during link construction. The probe identifier field and credibility score field in the subgraph edge set are used for path credibility synthesis and conflict bypass processing during link construction.
[0101] Summary of the technical effects of this step: This step completes the loading of the graph snapshot and the location of the resulting entities under the constraints of the version package, and forms a linked extraction link based on relation type, time window, and confidence threshold. This step assembles the extracted nodes and relations into an evaluation subgraph and encapsulates it into a subgraph package, while retaining the traceable association of the slot identifier field, probe identifier field, and change time window index field. This step uses the subgraph package as input to S700 and writes the verification mark field and bypass segment into the subgraph package, forming a usable data structure connection across steps.
[0102] To address the problem that existing technologies for evaluating technology transfer often rely on node-level scoring or simple neighborhood statistics, lacking path enumeration and path-level secondary constraints with the central achievement entity as the entry point, and failing to distinguish between semantic link construction methods for "traction chains" and "risk chains," this invention addresses the issue of insufficient evidence to simultaneously express both the link evidence driving transfer and the risk evidence inhibiting transfer within the same structure. This results in a lack of reusable link feature inputs for subsequent time-series trend calculations and model interpretation. Specifically, step S700 completes path enumeration and path-level secondary constraints under subgraph package constraints, and semantically generates link records according to the relationship type of traction chains and risk chains, forming a chain feature package for use in step S800.
[0103] S700. Perform link construction on the sub-graph package to generate a chain feature package; the chain feature package includes traction chain features, risk chain features, and differential packages;
[0104] In a specific implementation of the present invention, the input source of S700 is the subgraph package output by S600. The subgraph package includes at least a subgraph node set, a subgraph edge set, a result location record, and a window selection cache summary. The subgraph edge set retains a relation type identifier, a slot identifier field, a probe identifier field, a credibility score field, a delay marker field, and a change time window index field. The result location record is used to determine the central result entity of the link construction, the window selection cache summary is used to determine the windowing filtering caliber during link construction, and the probe identifier field and the credibility score field are used for path credibility synthesis and bypass processing during link construction. Specifically, S700 is completed collaboratively by a link construction unit, a path enumeration subunit, a path constraint subunit, a traction chain generation subunit, a risk chain generation subunit, a path credibility synthesis subunit, a differential mapping subunit, and a chain feature encapsulation unit. Among them, the link construction unit is responsible for scheduling and resource orchestration; the path enumeration subunit is responsible for generating reachable paths on the subgraph node set and subgraph edge set; the path constraint subunit is responsible for applying secondary constraints on relation type, time window, and credibility threshold; the traction chain generation subunit is responsible for extracting traction chains from candidate paths; the risk chain generation subunit is responsible for extracting risk chains from candidate paths; the path credibility synthesis subunit is responsible for performing credibility score synthesis on candidate paths; the differential mapping subunit is responsible for aligning the differential package with the path edges and generating path-level change labels; and the chain feature encapsulation unit is responsible for generating the chain feature package and outputting it to S800 as input.
[0105] In the initial stage of S700, the link construction unit takes the subgraph package as input, first reads the result location record to obtain the node address of the central result entity, and reads the relation type identifier and slot identifier fields of the subgraph edge set to establish a relation availability index. Simultaneously, the link construction unit reads the window selection cache summary and generates a link window configuration. The link window configuration includes the window start and end boundaries, window source marker, and window priority marker, where the window priority marker is used to determine the main window for path enumeration when multiple windows exist. Further, the link construction unit performs preprocessing on the subgraph edges carrying the verification marker field. Preprocessing includes writing the edges corresponding to non-backtrackable records into a weighted queue and writing the edges with delay marker fields into a delay bypass queue. Both the weighted queue and the delay bypass queue are associated with probe identifier fields for reference by the path credibility synthesis subunit during the synthesis process. Understandably, the minimum input set for link construction in this step consists of the node address of the central result entity, the relation type identifier in the subgraph edge set, the credibility score field, and the change time window index field. The node address of the central result entity is used to determine the link start point, the relation type identifier is used to distinguish between traction chains and risk chain candidate edges, the credibility score field is used for path credibility synthesis, and the change time window index field is used for path temporal alignment and differential mapping. The slot identifier field, probe identifier field, and delay mark field are key fields of this invention for enhancing traceability and bypass processing, and are preferably retained and integrated into subsequent steps in engineering implementation.
[0106] The path enumeration subunit generates candidate paths on the subgraph node set and subgraph edge set. Specifically, the path enumeration subunit uses the node address of the central result entity as the starting point, interprets the edge direction according to the direction constraint code of the relation type identifier, and limits the set of available edges under the link window configuration constraints. The edge direction interpretation can be implemented using a combined mapping table of relation type identifier and slot identifier fields. The combined mapping table is parsed by the slot template in the S100 stage and is traceable under the session hash index. Further, the path enumeration subunit sets an upper limit for enumeration depth and a upper limit for branches. The upper limit for enumeration depth is used to limit the number of hops that can be extended outward from the central result entity, and the upper limit for branches is used to limit the number of candidate edges that can be extended from each node. When the number of candidate edges exceeds the upper limit for branches, the path enumeration subunit prioritizes the edges with higher confidence scores and no conflict markers to enter the candidate extension set, and writes the remaining edges into the enumeration truncation record. The enumeration truncation record retains the conflict cluster identifier and conflict type code for subsequent traceability of the review set. Understandably, the upper limit of enumeration depth and the upper limit of branches are preferred parameters used to control the computational overhead and path explosion risk of link construction. Their implementation can be dynamically generated by the link construction unit according to the resource budget parameter group in the session packet, and the generation process and triggering conditions are recorded in the session log structure through the session hash, so that the link construction has traceable automated orchestration paths at the session level.
[0107] After candidate paths are generated, the path constraint subunit performs secondary screening and structured annotation on the candidate paths. Specifically, the path constraint subunit performs a relation type consistency check on each candidate path. The relation type consistency check determines whether the path belongs to the traction chain candidate set or the risk chain candidate set based on the set constraints of the relation type identifier. The traction chain candidate set is a combination of relation types corresponding to the transformation traction semantics, and the risk chain candidate set is a combination of relation types corresponding to the risk propagation semantics. This correspondence is further refined by the extraction rule unit of this invention based on the relation type whitelist in stage S600, and loaded into the link rule table by the link construction unit in this step. Further, the path constraint subunit performs a time window consistency check on the candidate paths. The time window consistency check determines the overlap between the change time window index field of the path edge and the start and end boundaries of the window configured in the link window. If there is an edge outside the window within the path, the path is marked as a cross-window path and written into the cross-window path queue. For cross-window paths, the path constraint subunit can choose to retain only the sub-paths within the window and record the pruning point position. The pruning point position is used as the input field for the subsequent traction chain generation subunit and risk chain generation subunit. Furthermore, the path constraint subunit performs a credibility threshold check on the candidate paths. This check takes the credibility score field of each edge within the path as input and combines it with the historical stability records of the data source access identifier and quality verification rule identifier corresponding to the probe identifier field to generate a path-level credibility threshold. If the composite credibility value of a candidate path is lower than the path-level credibility threshold, the path is added to the low-credibility path queue and associated with the probe identifier field for locating the data source. Understandably, the path constraint subunit elevates relation type, time window, and credibility threshold from edge-level constraints to path-level constraints and writes path marker fields and pruning point location fields into the candidate path structure. These fields are subsequently used by the chain feature encapsulation unit to form an interpretable entry point for the chain feature package.
[0108] After path selection, the traction chain generation subunit and the risk chain generation subunit respectively construct links for the candidate paths. Specifically, the traction chain generation subunit selects paths from the traction chain candidate set. Path selection is performed according to a combination sorting rule of the number of industry nodes covered by the path, the number of investment and financing event nodes, and the product of path edge weights. The number of industry nodes covered by the path is obtained by matching the entity type set of the path nodes, the number of investment and financing event nodes is obtained by aggregating the relationship type identifiers of the path nodes, and the product of path edge weights is obtained by aggregating the multiplication of the edge weight fields within the path. The edge weight fields can be written by the graph snapshot during graph updates and bound to the snapshot hash version. Further, the traction chain generation subunit generates traction chain records for the selected paths. The traction chain record includes the chain length, the product of path edge weights, the number of industry nodes covered by the path, the number of investment and financing event nodes, and the synthesized path credibility value. The traction chain record retains the slot identifier field and probe identifier field of the key edges within the path. The selection of key edges can be achieved based on the key relationship priority table of the relationship type identifier, which is traceable under the session hash index. Simultaneously, the risk chain generation subunit selects paths from the candidate risk chain set. Path selection is performed according to a combined sorting rule of negative event node count, ownership conflict node count, similar outcome competition node count, and risk propagation depth. The negative event node count, ownership conflict node count, and similar outcome competition node count are obtained by aggregating the relationship type identifiers of path nodes or path edges, while the risk propagation depth is synthesized from the path length and the proportion of risk relationship types. Further, the risk chain generation subunit generates risk chain records for the selected paths. These records include the negative event node count, ownership conflict node count, similar outcome competition node count, and risk propagation depth, and also retain the slot identifier field and probe identifier field for key edges within the path. For risk paths containing delayed bypass queue edges, the risk chain generation subunit writes a delayed bypass marker field into the risk chain record and associates this marker field with a delayed marker field for subsequent review set tracing.
[0109] During the generation of the traction chain and risk chain, the path credibility synthesis subunit performs credibility score synthesis on the path, generates the path credibility synthesis value, and writes it into the traction chain record and risk chain record. Specifically, the path credibility synthesis subunit takes the credibility score field of each edge within the path as the base input and performs correction processing in combination with the conflict marker field and the delay marker field. The conflict marker field includes a conflict cluster identifier, a conflict type code, a conflict density value, and a main slot identifier. The correction processing includes only counting the main edge corresponding to the main slot identifier for edges belonging to the same conflict cluster identifier, and including the credibility score field of the remaining edges in the conflict penalty item. For edges with a delay marker field, the path credibility synthesis subunit generates a delay penalty item based on the delay interval corresponding to the delay index of the entry timestamp and the event timestamp, and includes the delay penalty item and the conflict penalty item in the synthesis process of the path credibility synthesis value. Furthermore, the path credibility synthesis subunit introduces a quality history summary from the probe identification field. This quality history summary is statistically obtained from the session-level historical execution by the missing test flags, outlier thresholds, and duplicate record discrimination rules associated with the quality verification rule identifiers. The path credibility synthesis subunit incorporates the quality history summary as a weight adjustment term into the path credibility synthesis process, thereby forming a traceable reflection of data source quality fluctuations in the chain feature package. Understandably, the credibility score field, conflict flag field, and delay flag field together constitute the core parameter set for path credibility synthesis in this step, representing the minimum set of fields indispensable for realizing the improved link construction of this invention. The quality history summary is a preferred extended field used to enhance the stability of path credibility synthesis during engineering operation.
[0110] The differential mapping subunit aligns the differential package with the traction chain record and risk chain record, generates a link-level differential mapping result, and writes it into the chain feature package. Specifically, the differential mapping subunit uses the relation primary key and change time window index field of each edge within the path as indexes to retrieve corresponding entries for relation addition sets, relation deletion sets, and weight change sets in the differential package, and writes the retrieved entries into the link differential segment. The link differential segment retains the change type marker, change time window index, and association relationship type identifier. For edges in the path that are not hit in the differential package but exist in the graph snapshot, the differential mapping subunit marks them as stable edges and writes them into the stable edge count, which is used as input for constructing the trend of interest sequence in the subsequent S800 time series calculation stage. Further, for paths in the cross-window path queue, the differential mapping subunit prioritizes writing the differential mapping results of sub-paths within the window into the link differential segment, and writes the pruning point position field into the link differential segment, so that subsequent steps can reproduce the link boundary after window pruning within the same link record. Understandably, the differential mapping subunit maps the differential packet from a version-level structure to a path-level structure, thereby forming a link change label within the chain feature packet that is consistent with the snapshot hash version. This link change label is subsequently used as one of the direct inputs for the S800 to generate the window index.
[0111] In the output phase of this step, the chain feature encapsulation unit encapsulates the traction chain record, risk chain record, and link differential segments into the chain feature package and outputs it. Specifically, the chain feature package includes traction chain features, risk chain features, and a differential package. The traction chain features are obtained by summarizing the traction chain records and include at least the chain length, path edge weight product, number of industry nodes covered by the path, number of investment and financing event nodes, and path credibility composite value. The risk chain features are obtained by summarizing the risk chain records and include at least the number of negative event nodes, number of ownership conflict nodes, number of similar result competition nodes, and risk propagation depth. The differential package is carried in the chain feature package in the form of link differential segments and associated with the change time window index. Further, the chain feature encapsulation unit writes the session hash and snapshot hash into the header field of the chain feature package, and writes the result identifier and result entity node address of the central result entity, so that the chain feature package can be consistent with the window selection cache summary and graph snapshot version when used as input in S800. Understandably, the traction chain features and risk chain features in the chain feature package are used by S800 to "perform time series calculations on the chain feature package" to generate window indices and synthesize vector packages. The link difference segments in the chain feature package are used by S800 to construct the trend sequence of interest and drive the calculation of mutation point locations and trend slopes. At the same time, the probe identification field and slot identification field retained in the chain feature package are used for retrospective reference in the subsequent path evidence interpretation during the S900 output stage.
[0112] In summary, this step achieves the following technical results: Under the constraints of the subgraph package, this step completes path enumeration and path-level secondary constraints with the central result entity as the entry point, and generates link records according to the semantic relationship type between the traction chain and the risk chain. This step introduces the credibility score field, conflict marker field, and delay marker field into the path credibility synthesis, and maps the differential package to link differential segments and writes them into the chain feature package. The chain feature package output by this step forms an input connection with S800 under the constraints of session hash and snapshot hash, and retains key fields for subsequent time series calculations and evidence backtracking.
[0113] To address the problem that existing technologies often use single time series statistics to replace the popularity or attention trends of achievements, lacking a method for constructing attention trend sequences driven by map version difference, resulting in a disconnect between trend calculation and map changes, and making it difficult to simultaneously obtain trend slope, volatility, and mutation point locations under the same window management constraint and establish a correspondence with link evidence, this invention generates attention trend sequences by driving the difference package in the chain feature package according to the change time window index in step S800, forming window-level features and synthesizing window indices under window management constraints, and simultaneously encapsulating the window index, traction chain features, and risk chain features into a vector package for inference in step S900, specifically including:
[0114] S800: Perform time-series calculations on the chain feature package to generate a window index and synthesize a vector package;
[0115] In the specific implementation of this invention, the input source of S800 is the chain feature packet output by S700. The chain feature packet carries session hash and snapshot hash in the packet header field, and contains traction chain features, risk chain features and differential packets in the packet body. The traction chain features include at least chain length, path edge weight product, number of industry nodes covered by the path, number of investment and financing event nodes, and path credibility composite value. The risk chain features include at least negative event node count, ownership conflict node count, similar result competition node count, and risk propagation depth. The differential packets are carried in the form of link differential segments and are at least associated with change time window index and relationship type identifier. S800 is jointly completed by a time series calculation unit, a window management subunit, a sequence construction subunit, an anomaly and gap handling subunit, a trend extraction subunit, a mutation detection subunit, an index synthesis subunit, and a vector encapsulation unit. Among them, the time series calculation unit is responsible for executing scheduling and reading session hash and snapshot hash to establish the calculation context of this step; the window management subunit is responsible for generating or loading time windows and maintaining window sliding rules; the sequence construction subunit is responsible for converting discrete change events driven by the change time window index in the difference packet into the trend sequence of interest; the anomaly and gap handling subunit is responsible for handling sequence holes caused by missing test markers and delay markers; the trend extraction subunit is responsible for outputting the trend slope and volatility; the mutation detection subunit is responsible for outputting the mutation point location; the index synthesis subunit is responsible for synthesizing the window index according to the heat decay coefficient; and the vector encapsulation unit is responsible for generating vector packets and outputting them to S900 as input.
[0116] In the initial stage of S800, the time-series computation unit reads the session hash and snapshot hash of the chain feature packet, and combines them with the registered sampling granularity field and the windowing scope of the window selection cache digest in the session packet to generate a time-series computation configuration. This configuration includes at least window granularity, window step size, window start and end boundaries, upper limit of the number of windows, and window priority flag. Window granularity and window step size together define the sliding rules of the window management subunit. The window start and end boundaries limit the time index range of the sequence construction subunit. The upper limit of the number of windows limits the computational overhead of concurrent windows. The window priority flag is used to select the main window when multiple source windows coexist. Understandably, the minimum input set for this step's time-series computation is the traction chain feature, the risk chain feature, and the change time window index in the difference packet. The traction chain feature and risk chain feature provide the basic quantities for statistical summarization within the window, and the change time window index provides the discrete trigger points of the sequence time axis. The sampling granularity field and the window selection cache digest are preferred extended fields used for automated loading and session-level consistency recording of window granularity and sliding rules during project operation. Furthermore, the time-series calculation unit writes the time-series calculation configuration into the running log structure of this step, and records the configuration loading trigger conditions under the session hash index. The trigger conditions include at least snapshot hash change trigger and window number limit trigger. The former corresponds to the recalculation caused by version change after the graph update, and the latter corresponds to window splitting or merging caused by excessively dense link differential segments.
[0117] The window management subunit establishes a set of time windows based on the time series computation configuration and provides it to the sequence construction subunit for invocation. Specifically, the window management subunit organizes the set of time windows into a window index table, which includes window number, window start and end boundaries, window step offset, and window priority flag. The window number is associated with the central result identifier in the chain feature package, allowing the same result to use the window numbering system across different snapshot hash versions. Furthermore, when the window management subunit detects a change in snapshot hash, it performs a remapping of historical window numbers according to the window start and end boundary alignment strategy and records the comparison of window numbers before and after remapping in the runtime log structure. This enables the subsequent S900 evaluation model inference to identify the version caliber of the window index within the session hash range. For changed time window indices outside the window boundaries, the window management subunit writes them into an out-of-bounds record and associates them with a relationship type identifier. The out-of-bounds record is used by the anomaly and gap handling subunit to subsequently generate gap flags.
[0118] After the window index table is established, the sequence construction subunit performs time-series assembly on the differential package, generates the trend of interest sequence, and writes it into the intermediate sequence package. Specifically, the sequence construction subunit takes the link differential segment as input, locates the corresponding window number on the window index table according to the change time window index, and aggregates the change events according to the window number. The aggregation dimension includes at least a change type marker and a relationship type identifier. The change type marker is used to distinguish between relationship addition, relationship deletion, and weight change, and the relationship type identifier is used to distinguish between changes related to the traction chain and changes related to the risk chain. Further, the sequence construction subunit maps the aggregation results within each window to sequence points and writes a time window index field, a traction change count, a risk change count, and a stable edge count for each sequence point. The traction change count is obtained by aggregating changes related to the traction chain, the risk change count is obtained by aggregating changes related to the risk chain, and the stable edge count is obtained by counting the entries marked as stable edges in the link differential segment. The traction change count, risk change count, and stable edge count together constitute the minimum set of core fields of the trend of interest sequence. Understandably, the trend sequence is used as a direct input for window index synthesis in this step, and also as the time positioning basis for path evidence interpretation in S900. Therefore, when generating each sequence point, the sequence construction sub-unit further retains the key relationship type identifier summary and the corresponding window number associated with that sequence point, so as to make a reverse lookup in subsequent cross-step connections.
[0119] During the generation of the trend sequence, the anomaly and gap handling subunit processes the impact of missing data and delays in the sequence. Specifically, the anomaly and gap handling subunit reads the missing data markers and duplicate record discrimination rules associated with the quality verification rule identifiers, and combines them with the delay bypass marker field and delay marker field retained in the chain feature package to perform hole detection on the trend sequence. Hole detection includes at least window continuity detection and event density detection. When window continuity detection finds a break between adjacent window numbers, the anomaly and gap handling subunit writes a gap marker in the intermediate sequence package and records the start and end window numbers of the gap. When event density detection finds that the change event density within a window exceeds the density threshold, the anomaly and gap handling subunit marks the window as a congested window and records the congestion reason code. The congestion reason code is associated with a conflict cluster identifier summary or a delay bypass marker field summary, used to indicate that the congestion comes from a conflict set or a delayed entry set. Furthermore, for windows with gap markers, the anomaly and gap handling subunit can execute an interpolation strategy. The interpolation strategy includes two types: the hole-preserving strategy and the adjacent window smoothing strategy. The hole-preserving strategy directly passes the gap marker to the exponential synthesis subunit, while the adjacent window smoothing strategy generates a smooth value based on the traction change count and risk change count of adjacent windows and writes the interpolation marker into the intermediate sequence package. The hole-preserving strategy belongs to the minimum implementation set, while the adjacent window smoothing strategy belongs to the preferred extended function. Both are controlled by the anomaly handling mode field in the time series calculation configuration and the triggering conditions are recorded in the runtime log structure.
[0120] After the intermediate sequence package contains the trend sequence of interest, the trend extraction subunit calculates the trend slope and volatility of the trend sequence of interest and writes it into the window feature package. Specifically, the trend extraction subunit performs sequence differencing on the traction change count and risk change count for each window number, and calculates the local rate of change within the sliding range across the window to obtain the trend slope. Simultaneously, the trend extraction subunit performs discrete volatility statistics on the discrete changes of the traction change count and risk change count within the same sliding range to obtain the volatility, and writes the trend slope and volatility into the window feature package along with the window number and time window index fields. Further, the trend extraction subunit writes the stable edge count as a stationarity reference field into the window feature package, and when there is a gap marker or congestion window marker, writes the trend slope and volatility into a quality marker field. The quality marker field is used by the index synthesis subunit to perform weight adjustment. Understandably, the trend slope and volatility are the minimum set of core constituent fields of the window index, and together with the location of the mutation point and the heat decay coefficient, they constitute the synthetic input of the window index; the stable edge count and quality mark fields are preferred extended fields, used to enhance the interpretability and auditability of window features when implemented in engineering.
[0121] The mutation detection subunit performs mutation point location detection on the window feature package and writes it into the mutation record. Specifically, the mutation detection subunit takes the trend slope, volatility, traction change count, and risk change count in the window feature package as input, and generates a mutation candidate set using a sliding window-based mutation discrimination rule. The mutation discrimination rule includes at least a threshold trigger condition and a persistence trigger condition. The threshold trigger condition is used to identify cases where the mutation intensity within a single window exceeds a threshold, and the persistence trigger condition is used to identify cases where multiple consecutive windows satisfy the same direction of change. When a mutation candidate set is matched, the mutation detection subunit outputs the mutation point location and represents it as a combined location value of the window number and the time window index field. The combined location value is written into the mutation record and associated with a trigger type flag. Further, the mutation detection subunit adopts a bypass strategy for windows with notch marks, excluding them from the mutation candidate set generation and writing the bypass reason into the mutation record. The bypass reason includes at least a notch mark and an interpolation mark. Understandably, the mutation point location is a key field in the window index that indicates local structural changes in the trend sequence of concern. It belongs to one of the minimum output sets of this step and forms a cross-step connection with the subsequent path evidence interpretation of S900, enabling the evaluation model to infer that the mutation point location can be used to label the time source of potential scores and risk indicators.
[0122] After trend extraction and mutation detection are completed, the index synthesis subunit performs window index synthesis on the window feature package and mutation records, and generates window segment inputs for the vector package. Specifically, the index synthesis subunit reads the heat decay coefficient and interprets it as a decay weight on the contribution of historical windows. The heat decay coefficient can be determined by the resource budget parameter group and the upper limit of the number of windows in the session package, and its loading source is recorded in the running log structure. Further, the index synthesis subunit uses the in-window statistics, trend slope, volatility, and mutation point location of the trend sequence as synthesis inputs, and outputs the window index according to the window number. The window index at least includes the window number, time window index field, window index value, mutation point location summary, and quality label field in its field structure. The quality label field comes from the window feature package and is used to indicate whether the window index is affected by gap label, congestion window label, or imputation label. Understandably, the window number, time window index field, and window index value together constitute the minimum set of core fields of the window index. The mutation point location summary and quality label field are preferred extended fields used for subsequent evaluation of the interpretation of model inference and audit support for drift detection.
[0123] In the output phase of this step, the vector encapsulation unit vectorizes the window index and the traction chain features and risk chain features in the chain feature package, generating a vector package and outputting it to the S900 as input. Specifically, the vector encapsulation unit writes the session hash and snapshot hash into the vector package header field, and writes the central result identifier and window number range; the vector encapsulation unit writes the window index into the time segment of the vector package in window number order, and writes the traction chain features and risk chain features into the static segment of the vector package, where the time segment is used for the S900's evaluation model to infer the evolution information of the capture trend sequence of interest, and the static segment is used for the S900's evaluation model to infer the reference link structure and link statistics. Further, the vector encapsulation unit writes a source mapping field into the vector package, which is associated with the window number and the corresponding differential packet link differential segment summary, so that the S900 can perform reverse lookup within the session hash and snapshot hash range when outputting path evidence interpretation.
[0124] This step's technical effects can be summarized as follows: This step generates a trend sequence of interest by driving the difference packets in the chain feature package with the change time window index, and forms window-level trend slopes, volatility, and abrupt change point locations under window management constraints. This step synthesizes window-level features based on the heat decay coefficient and outputs a window index, while simultaneously encapsulating the window index, traction chain features, and risk chain features into a vector package. The vector package output by this step carries session hashes and snapshot hashes and serves as the inference input for the S900 evaluation model, achieving cross-step consistency and traceability entry.
[0125] To address the problems in existing technology where technology transfer potential assessment models often only output a single score and lack interpretable correlation with the path evidence in the graph, and where model training parameters and inference versions lack session-level registration, leading to untraceable inference conclusions, difficulty in correcting confidence levels, and a lack of stable input objects for subsequent drift detection, this invention, through step S900, loads and registers the assessment model and historical transfer training parameters after receiving the vector packet, and performs joint inference based on the window index time series segment and chain feature static segment to output potential scores and risk indicators. Simultaneously, it generates path evidence interpretations aligned with the time range field and encapsulates the assessment result package, specifically including:
[0126] S900, Perform evaluation model inference on the vector package; the evaluation model includes historical conversion training parameters and outputs potential score, risk index, confidence level, and path evidence interpretation;
[0127] In the specific implementation of this invention, the input source of S900 is the vector packet output by S800. The vector packet carries session hash and snapshot hash in the packet header field, and organizes the time-series segments of window index and the static segments of traction chain features and risk chain features in the packet body. The time-series segment includes at least the window number, time window index field, window index value, and associated mutation point location summary and quality label field. The static segment includes at least the chain length, path edge weight product, number of industry nodes covered by the path, number of investment and financing event nodes, path credibility composite value, number of negative event nodes, number of ownership conflict nodes, number of similar achievement competition nodes, and risk propagation depth. The S900 is jointly completed by an evaluation model service unit, a feature loading subunit, a model version management subunit, an inference execution subunit, an output verification subunit, an interpretation generation subunit, and a result encapsulation unit. The evaluation model service unit loads and invokes the evaluation model; the feature loading subunit unpacks the vector package and constructs the model input tensor structure; the model version management subunit loads historical transformation training parameters and registers the evaluation model version number; the inference execution subunit performs model inference and generates potential scores and risk indicators; the output verification subunit performs boundary constraint processing on confidence scores and abnormal outputs; the interpretation generation subunit generates path evidence interpretations; and the result encapsulation unit outputs an evaluation result package containing the potential score, the risk indicator, the confidence score, and the path evidence interpretation, which serves as input for subsequent drift detection.
[0128] In the initial stage of S900, the evaluation model service unit obtains an evaluation model loading request from the model version management subunit. This loading request includes at least a session hash, a snapshot hash, an evaluation model version number index, a historical transformation training parameter index, and a resource budget constraint field. The historical transformation training parameters are defined as a set of parameters trained from historical transformation samples. This parameter set includes at least feature normalization parameters, model weight parameters, threshold caliber parameters, and calibration parameters. The feature normalization parameters map window exponent values and chain feature scales to a unified dimension. The model weight parameters express the feature combination relationships within the model. The threshold caliber parameters constrain the segmented mapping caliber of risk indicators. The calibration parameters perform posterior correction on the confidence level. Further, when loading historical transformation training parameters, the model version management subunit performs a parameter integrity check and generates a parameter loading record. This record includes at least the evaluation model version number, parameter timestamp, parameter source identifier, and hash digest. The hash digest is bound to the snapshot hash and written into the session-level audit record, enabling subsequent evaluation result packages to be traced back to the parameter caliber within the session hash range. Understandably, the historical transformation training parameters are one of the smallest sets of inputs that can be used in this step, and together with the vector package, they constitute the necessary inputs for the inference execution subunit; the resource budget constraint field is a preferred extended field, used to schedule and limit the number of concurrent inferences and the batch window size during project execution.
[0129] The feature loading subunit performs unpacking, alignment, and input construction processing on the vector packet to obtain the model input packet, which is then used by the inference execution subunit. Specifically, the feature loading subunit first reads the session hash and snapshot hash of the vector packet and writes them as the context identifier for this inference into the header field of the model input packet. The feature loading subunit further reads the temporal and static segments of the vector packet, aligns the temporal segments according to window numbering order to form a window exponential sequence structure, and maps the static segments to a chain feature vector structure. Further, the feature loading subunit performs normalization processing on the window exponential value and the chain feature vector structure based on the feature normalization parameters and writes a normalization caliber marker field into the model input packet. Simultaneously, the feature loading subunit reads the quality marker field and mutation point location summary, converts them into a quality mask structure and a mutation hint structure, and writes them into the model input packet. The quality mask structure is used within the inference execution subunit to adjust the weights of windows affected by gap markers, congestion window markers, or interpolation markers. The mutation hint structure is used within the inference execution subunit to assign differentiated attention weights to mutation neighborhood windows. Understandably, the window exponential sequence structure and the chain feature vector structure are the minimum feature set for model inference, while the quality mask structure and mutation hint structure are preferred extended features used to enhance the adaptability to changes in data quality and temporal structure without changing the field system of the vector package. When generating the above structures, the feature loading subunit also writes the window number range and the time window index field range into the time range field of the model input package, so that subsequent path evidence interpretation can use the same time range field to complete cross-step connection.
[0130] After the model input package is constructed, the inference execution subunit performs evaluation model inference on the model input package, outputting potential scores, risk indicators, and confidence levels. Specifically, the inference execution subunit uses the window exponential sequence structure and the chain feature vector structure as joint inputs, and reads the model weight parameters to generate an inference computation graph structure. The inference computation graph structure logically includes a temporal encoding submodule, a static fusion submodule, and an output header submodule. The temporal encoding submodule is used to extract temporal representations from the window exponential sequence structure, the static fusion submodule is used to fuse the chain feature vector structure with the temporal representations, and the output header submodule is used to output the raw output values of the potential score and risk indicator respectively, and pass the raw output values to the confidence calculation branch. Furthermore, the inference execution subunit invokes the quality mask structure during the time-series encoding stage to perform weighted suppression processing on window index values marked as low quality, and performs neighborhood aggregation enhancement processing on the window representation within the mutation neighborhood indicated by the mutation hint structure. The usage conditions of the aforementioned quality mask structure and mutation hint structure are controlled by the threshold caliber parameters loaded by the model version management subunit. The threshold caliber parameters include at least a quality weight threshold and a mutation neighborhood width threshold. The former is used to define the suppression boundary of low-quality windows, and the latter is used to define the number of windows that expand to both sides of the mutation point location summary. Understandably, the potential score is defined as the inference output of the evaluation model on the comprehensive transformation trend and traction chain structure signal of the outcome entity within a given time range field, and the risk index is defined as the inference output of the evaluation model on the risk chain structure signal and the fluctuation structure of the window index sequence. Both are output by the inference execution subunit and written into the intermediate inference record, which also includes a session hash, a snapshot hash, and the evaluation model version number for subsequent encapsulation and auditing.
[0131] After the inference output is generated, the output verification subunit performs boundary constraints and anomaly output processing on the potential score, the risk index, and the confidence level. Specifically, the output verification subunit reads the calibration parameters and performs posterior correction on the confidence level. The correction process includes mapping the original confidence level value to a calibration mapping table loaded according to the historical sample distribution, and writing the mapped confidence level into the result field. Simultaneously, the output verification subunit performs range pruning and missing data backfilling on the potential score and the risk index. Range pruning is based on the upper and lower bounds in the threshold caliber parameters. Missing data backfilling is triggered when a gap marker is detected in the model input package and the gap coverage exceeds the gap threshold. After triggering, the potential score and the risk index are marked as gap backfilling status and a backfilling reason code is written into the verification record. Further, when the output verification subunit detects a non-numerical anomaly or a sequence length inconsistency anomaly in the inference output, it generates an abnormal inference event and associates it with the session hash, snapshot hash, window number range, and evaluation model version number, writing it into an abnormal event package. The abnormal event package is used for reference in the system operation and maintenance link or subsequent drift detection process. Understandably, range pruning and confidence correction belong to the minimum implementation set of this step, while missing data backfilling and abnormal inference events are preferred extended functions used to improve the continuous operation capability of the inference link and retain audit entry points in engineering deployment.
[0132] The explanation generation subunit generates path evidence explanations and writes them into the explanation package based on the products of the inference execution subunit and the output verification subunit. Specifically, the explanation generation subunit uses the source mapping field in the vector package as the explanation index entry, and associates the source mapping field with the window number and the difference segment summary of the difference package link. The explanation generation subunit further reads the time range field and the mutation point location summary in the model input package, constructs the explanation time anchor, and locates the contribution candidate window from the source mapping field in the neighborhood of the explanation time anchor. Further, the explanation generation subunit performs relation type identifier summary backtracking within the contribution candidate window. The backtracking process aligns and matches the relation type identifier in the difference segment summary of the difference package link with the traction chain feature and risk chain feature to generate an explanation fragment set. The explanation fragment set includes at least the window number, time window index field, relation type identifier summary, change type marker summary, and chain feature reference summary. The chain feature reference summary is defined as a reference identifier for static segment fields such as chain length, path edge weight product, investment and financing event node count, negative event node count, and ownership conflict node count, which is used to establish a correspondence between the explanation fragment and the static chain feature. Understandably, the minimum set of fields for path evidence interpretation includes at least a window number, a time window index field, and a relation type identifier digest, enabling the interpretation to point to the temporal position of the window index sequence structure and the type position of the link difference within the session hash range. The interpretation fragment set and chain feature reference digest are preferred extensions used to enhance the readability and audit traceability of the interpretation in engineering implementations. After generating the interpretation package, the interpretation generation subunit binds the interpretation package with the session hash, snapshot hash, and evaluation model version number, and writes the interpretation package hash digest for reference by subsequent processing transaction packages.
[0133] In the output phase of this step, the result encapsulation unit encapsulates the potential score, risk indicator, confidence level, and path evidence interpretation into an evaluation result package and outputs it. Specifically, the result encapsulation unit writes the session hash, snapshot hash, and evaluation model version number into the evaluation result package header field, and also writes the time range field and window number range. The result encapsulation unit writes the potential score field, risk indicator field, confidence level field, and path evidence interpretation field into the evaluation result package body, where the path evidence interpretation field carries an interpretation package or interpretation package hash digest and is associated with an interpretation time anchor. Further, the result encapsulation unit writes the evaluation result package into the session-level result storage and registers the output timestamp and output status code under the session hash index. The output status code includes at least a normal output status and a gap-filling status. Simultaneously, the result encapsulation unit passes the reference handle of the evaluation result package to the drift detection process executed after S900, enabling drift detection to use the potential score, risk indicator, and path evidence interpretation as input objects and generate a disposal transaction package by comparing with the historical window index range.
[0134] In summary, this step achieves the following technical results: After receiving the vector packet, it loads and registers the evaluation model and historical transformation training parameters, and performs joint inference based on the window exponential time series segment and chain feature static segment to output potential scores and risk indicators. This step corrects the confidence level in the output verification link and retains entry points for abnormal inference records. Simultaneously, it generates path evidence explanations aligned with the time range field based on the source mapping field. The evaluation result packet output by this step carries the session hash, snapshot hash, and evaluation model version number, and serves as input for the subsequent drift detection and handling transaction packet generation.
Claims
1. A method for evaluating the commercialization potential of research results based on dynamic knowledge graphs, characterized in that, include: S100: Obtain the result identifier set and data source list, configure data probes according to the data source list and generate slot templates, register session hashes, and generate session packets; S200: Perform probe acquisition on the session packet to generate an evidence packet; The evidence package includes event timestamps, data entry timestamps, and delay metrics; S300: Perform entity recognition, entity alignment, and relation extraction on the evidence package to generate candidate packages, and bind slot identifiers, probe identifiers, credibility scores, and conflict markers to the relations; S400. Perform gated resolution on the candidate packets to generate admission packets and generate a verification set; S500: Perform graph update on the admission package to generate a version package; the version package includes a graph snapshot, a snapshot hash, and a differential package; S600. Perform subgraph extraction on the version package, locate the result entity in the graph snapshot, extract and evaluate subgraphs according to relation type, time window and confidence threshold, and generate subgraph package; S700: Perform link construction on the sub-graph package to generate a chain feature package; The chain feature package includes traction chain features, risk chain features, and differential packages; S800: Perform time-series calculations on the chain feature package to generate a window index and synthesize a vector package; S900, Perform evaluation model inference on the vector package; the evaluation model includes historical transformation training parameters and outputs potential score, risk index, confidence level, and path evidence interpretation.
2. The method according to claim 1, characterized in that, The data probe includes a sampling granularity field, a field mapping table, an entity dictionary version, a data source access identifier, an extraction rule identifier, and a quality verification rule identifier; the quality verification rule identifier is associated with a missing test flag, an outlier threshold, and a duplicate record discrimination rule.
3. The method according to claim 1, characterized in that, The slot template includes slot identifier, entity type set, relation type set, required field set, evidence field set, credibility threshold, and conflict type code.
4. The method according to claim 1, characterized in that, The delay index is obtained by subtracting the event timestamp from the entry timestamp; the gating resolution includes generating delay tags according to the delay index threshold and writing the relationships with delay tags into the review set.
5. The method according to claim 1, characterized in that, The entity alignment uses name similarity, attribute similarity, and semantic vector similarity to synthesize an alignment score, constructs an entity alignment mapping table, and writes it into the candidate package.
6. The method according to claim 1, characterized in that, The conflict marker includes a conflict cluster identifier, a conflict type code, a conflict density value, and a primary slot identifier.
7. The method according to claim 1, characterized in that, The differential package includes a set of newly added nodes, a set of deleted nodes, a set of newly added relationships, a set of deleted relationships, a set of weight changes, and an index of change time windows; The change time window index is associated with the event timestamp.
8. The method according to claim 1, characterized in that, The traction chain features include chain length, path edge weight product, number of industry nodes covered by the path, number of investment and financing event nodes, and path credibility composite value; the risk chain features include negative event node count, ownership conflict node count, similar achievement competition node count, and risk propagation depth.
9. The method according to claim 1, characterized in that, The window index is synthesized from the position of the mutation point, the slope of the trend, the volatility, and the heat decay coefficient of the trend sequence; the trend sequence is updated driven by the change time window index in the difference package.
10. The method according to claim 1, characterized in that, After S900, drift detection is performed on the potential score, the risk index, and the path evidence interpretation; the drift detection includes comparing the historical window index interval and generating a disposal transaction package; the disposal transaction package includes a disposal type code, target node identifier, target relationship identifier, session hash, snapshot hash, evaluation model version number, and disposal timestamp.