Free implementation analysis knowledge base construction method based on intelligent agent and workflow
By constructing a molecular-patent knowledge graph and using Markush template distance assessment, the problems of omissions and over-conservatism in the free-to-implement analysis of innovative drug development are solved. This enables the unified expression of chemical structure and legal information and automated risk assessment, improving the accuracy and traceability of the analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU HUIYIDAO TECH CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are insufficient to fully identify potential coverage areas, leading to omissions or excessive conservatism in the free-performance analysis during innovative drug development. Furthermore, they lack unified modeling capabilities for chemical structures, claim texts, targets, and historical legal outcomes, making it difficult to support the continuous analysis needs from molecular design to project initiation decisions.
A knowledge base construction method for free implementation analysis based on agents and workflows is proposed to achieve risk assessment of free implementation of candidate molecules by constructing a molecular-patent knowledge graph and combining Markush template distance and causal calibration.
It achieves the associated expression of chemical structure, patent text and target information in the same graphical model, quantitatively characterizes the relationship between molecules and patent coverage, automates risk screening and recommendation, and ensures that the assessment results are consistent with actual legal risks and are traceable.
Smart Images

Figure CN122024912A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cheminformatics technology, and in particular to a method for constructing a free-implementation analytical knowledge base based on intelligent agents and workflows. Background Technology
[0002] In innovative drug development, Freedom to Execute (FTO) analysis has become a crucial step in determining whether candidate molecules can enter clinical trials and be marketed. With the rapid increase in the number of global patents, a large number of structural patents employing Markush claims are intertwined across multiple countries and legal systems. Relying solely on manual searches and experience-based judgment is insufficient to comprehensively identify potential coverage, easily leading to omissions or overly conservative approaches, thus impacting project decisions. Most existing information systems only support searches by literature or structural fragments, lacking the ability to uniformly model chemical structures, claim texts, targets, and historical legal outcomes. They also lack measurable and comparable methods for characterizing freedom to execute risks, making it difficult to support the continuous analytical needs from molecule design to project initiation decisions. Summary of the Invention
[0003] To address the numerous problems existing in the prior art, this invention provides a method for constructing a knowledge base for free implementation analysis based on intelligent agents and workflows. This invention is based on a molecular-patent knowledge graph, combined with Markush template distance and causal calibration, to achieve risk assessment of free implementation of candidate molecules.
[0004] This specification provides one or more embodiments of a method for constructing a free-implementation analytical knowledge base based on intelligent agents and workflows, including the following steps: A free-implementation knowledge graph is constructed, which establishes the relationships between molecular nodes, claim nodes, and target nodes based on chemical structure data and patent text data. Markush templates are generated based on the structured claims in the free implementation knowledge graph. The free implementation distance vector of each molecule relative to each Markush template is calculated, and the Markush template and the free implementation distance vector are written into the free implementation knowledge graph. A target analysis workflow is constructed based on the free implementation knowledge graph and the free implementation distance vector. The candidate molecule generation agent and the free implementation evaluation agent are called to obtain the free implementation risk assessment results and recommendation results of the candidate molecules. A causal analysis model is established based on workflow logs, patent examination results, and infringement judgment results. The aggregation method of the free implementation distance vector and the free implementation risk assessment rules are then calibrated according to the causal analysis model.
[0005] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By constructing a free implementation knowledge graph based on molecular nodes, claim nodes, and target nodes, the associated expression of chemical structure, patent text, and target information is realized in the same graph model, solving the problem of the dispersion and difficulty in linkage analysis of structural data and legal information in the prior art.
[0006] By extracting Markush templates from claims with structural formulas and calculating multi-component free implementation distance vectors, a quantifiable characterization of the relationship between molecules and patent coverage is achieved, overcoming the shortcomings of relying mainly on empirical judgments and making it difficult to compare the size of the free implementation space of different candidate molecules.
[0007] By building a target analysis workflow on top of a knowledge graph and introducing candidate molecule generation agents and free-implementation evaluation agents, automated risk screening and recommendation from target to candidate molecules is achieved, reducing the workload of repeated manual novelty searches and comparisons.
[0008] By utilizing workflow logs, patent examination results, and infringement judgments to establish a causal analysis model and regularly calibrating distance aggregation methods and evaluation rules, adaptive updates of freely implemented evaluation rules are achieved, ensuring that evaluation results remain consistent with actual legal risks and are traceable. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of the execution flow of the method of the present invention; Figure 2 This is a schematic diagram of the HER2 free implementation analysis process in a specific embodiment of the present invention; Figure 3 This is a comparison chart of the hit rate of free implementation risk level before and after calibration in a specific embodiment of the present invention. Detailed Implementation
[0010] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure.
[0011] An intelligent agent is a computational entity capable of autonomously invoking tools, scheduling data, and executing steps based on a task description under preset objective constraints. Its typical characteristics include task decomposition, state maintenance, result feedback, and anomaly handling capabilities. It can collaborate with external databases, retrieval services, inference models, and workflow orchestrators through standardized interfaces. The free-operation analysis knowledge base is a structured knowledge carrier for patent free-operation analysis scenarios. It unifies the organization of chemical structure information, patent claim text, Markush structure representation, target and activity-related information, and risk assessment records generated by the analysis process, ensuring repeatability and auditability for retrieval, comparison, traceability, and updates. Based on these capabilities, this invention proposes a method for constructing a free-operation analysis knowledge base based on intelligent agents and workflows. Through knowledge graph organization, Markush template generation and distance measurement, workflow-driven intelligent agent collaboration, and causal calibration mechanisms, it achieves a computable expression and sustainable optimization of the free-operation risk of candidate molecules.
[0012] like Figure 1 As shown, a method for constructing a free-implementation analytical knowledge base based on agents and workflows includes the following steps: A free-implementation knowledge graph is constructed, which establishes the relationships between molecular nodes, claim nodes, and target nodes based on chemical structure data and patent text data. The construction of a free-implementation knowledge graph includes: establishing molecular nodes based on chemical structure data, where molecular nodes record molecular structural formulas and molecular identifiers; establishing claim nodes based on patent text data, where claim nodes record claim texts and publication numbers; and establishing associations between molecular nodes, claim nodes, and target nodes through a graph database.
[0013] In one embodiment, the present invention first constructs a free implementation knowledge graph to centrally represent the correspondence between chemical structure information and patent text information. The chemical structure data can come from existing chemical structure databases or an enterprise's internal R&D system, and includes at least the molecular structural formula, molecular name, and molecular identifier. The patent text data can come from a patent search system and includes at least the patent number, claim text, and compound-related fields.
[0014] When constructing a free-implementation knowledge graph, the system uses a graph database as the storage engine. Each chemical structure data entry is created as a molecular node, which stores the machine-readable representation of the molecular structure and the molecular identifier. For example, the structure is converted into a linear representation and saved as an attribute. Each patent claim is created as a claim node, which stores the full text of the claim, the publication number, and the basic information of the patent to which it belongs. For target names or target identifiers appearing in the claim text, target nodes are created, which record the target name, target type, and source database.
[0015] Subsequently, the system establishes associations between nodes through rule matching and text parsing. First, based on the comparison of molecular identifiers, molecular names, and structural formulas, a "disclosure" relationship edge is established between the molecular node and the compound explicitly disclosed in the claim text. For molecules that demonstrate a binding or regulatory relationship with a specific target through experimental data or external activity databases, the system establishes an "action" relationship edge between the molecular node and the target node, recording the source of the activity data and the verification method in the relationship edge attributes. Furthermore, if the claim text points to a target through a functional description, the system establishes a "target-related" relationship edge between the claim node and the target node, thus forming a "claim—target—molecule" path.
[0016] During data import, the system performs standardization operations on chemical structure data and patent text data, including removing duplicate molecules, standardizing molecular identifier formats, cleaning irrelevant fields from patent texts, and normalizing target names using synonyms to ensure consistent referencing across different data sources within the graph structure. Once constructed, the free-implementation knowledge graph provides query interfaces based on node and relationship types, allowing subsequent modules to quickly locate relevant claim nodes and target nodes based on molecular nodes. This provides a unified data foundation for subsequent Markush template generation, free-implementation distance calculation, and target analysis workflows.
[0017] Markush templates are generated based on the structured claims in the free implementation knowledge graph. The free implementation distance vector of each molecule relative to each Markush template is calculated, and the Markush template and the free implementation distance vector are written into the free implementation knowledge graph. In this embodiment, after the free implementation knowledge graph is constructed, the system generates Markush templates for the claims with structure in the free implementation knowledge graph, calculates the free implementation distance vector of each molecule relative to each Markush template based on the Markush template, and then writes the Markush template and the free implementation distance vector into the free implementation knowledge graph as the basic data for subsequent free implementation evaluation.
[0018] First, the claim nodes in the free implementation knowledge graph are screened to select claims that explicitly reference the structural formula in the claim text and have established structural formula image associations in the graph. For each selected claim, the system reads the corresponding structural formula image and claim text, and uses a structural formula recognition tool to convert the structural formula image into molecular skeleton and substituent position markers. Simultaneously, the system performs syntactic parsing on phrases such as "substituent is," "substituent is selected from," and "group is selected from," extracting the allowed set of substituents, prohibited substituent types, and quantity restrictions for each substituent position. The system encapsulates the molecular skeleton and substituent positions and their restrictions together into a Markush template object and creates a corresponding Markush template node in the graph database, recording the template identifier, molecular skeleton representation, and constraint tables for each substituent position.
[0019] After generating the Markush template, the system calculates the freedom implementation distance vector based on the relationship between molecular nodes in the freedom implementation knowledge graph and Markush template nodes. To do this, the system first performs symbolic matching, mapping each molecular structure to the backbone of the Markush template, attempting to determine whether each substituent in the molecule falls into the allowed substituent set defined in the template. When a substituent cannot be matched, the system records the mismatch position and obtains the symbolic constraint distance based on the number and severity of mismatch positions. The larger the symbolic constraint distance, the greater the degree to which the molecule deviates from the textual constraints of the Markush template, and the lower the risk of freedom implementation.
[0020] Secondly, the system constructs molecular graph representations for molecular structures and Markush templates. Using existing molecular fingerprinting algorithms or graph neural network embedding models, each molecule and each Markush template is converted into a fixed-length vector representation. Then, the distance between the molecular vector and the Markush template vector is calculated as the embedding spatial distance, used to measure the degree of similarity between the molecule and the structure covered by the Markush template in terms of overall structure and physicochemical properties. This distance is used to supplement similarity information that is difficult to cover in symbol matching; for example, in cases where the types of substituents are different but the overall shape and electronic properties are similar, the embedding spatial distance can reflect potential coverage risks.
[0021] Furthermore, the system estimates the minimum structural modifications required for a molecule to satisfy the Markush template constraints through structural editing operations. Specifically, allowed editing operations are defined as adding substituents, deleting substituents, changing substituent types, and adjusting some chemical bond types. For each molecule, the system searches for an editing path from the current structure to a structure satisfying the Markush template constraints, calculates the required number of editing steps, or performs a weighted sum based on different editing operations, and uses this result as the structural editing distance. The smaller the structural editing distance, the more easily the molecule can fall within the scope of the claims with minimal structural modifications, but the higher the risk of permissive freedom.
[0022] The symbolic constraint distance, embedding space distance, and structural editing distance mentioned above together constitute the free implementation distance vector. In the graph database, the system establishes relational edges between molecular nodes and Markush template nodes, storing the free implementation distance vector as an attribute of these relational edges. Thus, in subsequent target analysis workflows, the candidate molecule generation agent and the free implementation evaluation agent can directly obtain the free implementation distance vector between each candidate molecule and the relevant Markush template by querying the relational edge attributes, eliminating the need to repeatedly parse the claim text and structural formula, significantly reducing redundant calculations, and improving the consistency and traceability of free implementation analysis.
[0023] Generating Markush templates based on claims with structured representations in a free-implementation knowledge graph includes: identifying structural reference markers in the claim text, extracting structural images corresponding to the structural reference markers, converting the structural images into molecular skeleton diagrams and substituent position sets using a structural recognition algorithm, and storing the molecular skeleton diagrams and substituent position sets as Markush templates.
[0024] After constructing the free-implementation knowledge graph, the system performs structured parsing and template generation on the claims containing structured elements, unifying textual descriptions and graphical structures into Markush templates. Specifically, the system first performs format and semantic analysis on the claim text to identify markers representing structural references, such as fixed-format numbers like "Formula I," "Compound 1," and "General Formula A." The parsing module uses regular expression matching and syntactic rules to determine the position of each structural reference marker in the claim and, based on this, locates the associated structural image file path or identifier within the free-implementation knowledge graph.
[0025] After obtaining the structural image, the system invokes a structural recognition algorithm to binarize the image, extract lines, and detect nodes, converting atomic symbols, bond lines, and substituent markings into a machine-readable molecular skeleton diagram. The molecular skeleton diagram is represented using a graph data structure, where nodes represent atoms in the skeleton, edges represent chemical bonds, and substituent markings such as "R," "R1," and "R2" on atoms are identified as substituent positions. During the recognition process, the system employs template matching for common aromatic rings, heterocycles, and condensed ring structures to improve the stability of skeleton recognition. It also performs simple corrections for any broken lines or overlapping markings in the structural image to ensure that the generated molecular skeleton diagram is chemically topologically connected and logical.
[0026] Simultaneously, the system parses the substituent statements in the corresponding claim text. For statements such as "R1 is", "R2 is selected from", and "R1 and R2 are independently represented", the system uses word segmentation and dependency parsing techniques to extract the list of allowed substituents, excluded group types, quantities, and value ranges for each substituent position. For example, when the text contains "R1 is hydrogen, halogen, or C1-6 alkyl", the system parses the statement into a set of candidate substituents and their types for the R1 position, and records them in a structured form; when the text contains "R1 and R2 can be linked to form a ring", the system adds a ring-forming relationship marker between the two corresponding substituent positions. For further limitations on the same substituent position in different dependent claims, the system integrates multiple statements through citation relationships to form a consistent set of substituent constraints.
[0027] After completing the dual analysis of the molecular skeleton diagram and substituent positions, the system performs alignment and verification steps, matching the substituent numbers appearing in the text with the substituent markers on the molecular skeleton diagram one-to-one. It checks for mismatched numbers or redundant markers, and corrects inconsistencies using manual rules or marks them as anomalies for manual review. Upon successful alignment, the system constructs a Markush template data object. The template includes at least the molecular skeleton diagram, a set of substituent positions, a set of allowed substituents for each substituent position, prohibited substituent types, and the relationships between substituents. This template object is persisted as a Markush template node in the free implementation knowledge graph. Simultaneously, an association is established between the Markush template node and the corresponding claim node, allowing subsequent modules to quickly access the corresponding Markush template definition from the claim node when handling free implementation risks. Through this process, complex textual claims and structural images are uniformly abstracted into formalized Markush templates, facilitating automated measurement and comparison of the coverage relationships between different molecules and the template in subsequent steps.
[0028] The free implementation distance vector includes symbolic constraint distance, embedding space distance, and structural editing distance. The symbolic constraint distance is calculated based on the substituent constraints in the Markush template and the matching results with the molecular structure. The embedding space distance is calculated based on the distance between the molecular vector generated by the molecular graph representation and the corresponding molecular vector in the Markush template. The structural editing distance is calculated based on the minimum number of editing steps required to perform structural editing operations on the molecular graph to meet the Markush template constraints.
[0029] In one embodiment, after generating the Markush template and establishing associations between molecular nodes and Markush template nodes, this invention calculates a free implementation distance vector for each pair of molecules and the Markush template. This free implementation distance vector consists of three parts: symbolic constraint distance, embedding space distance, and structural editing distance. It is used to measure the risk of a molecule falling within the patent protection scope defined by the Markush template from different perspectives. By combining these three distances into a single vector and storing it in the relationship edge attributes between molecular nodes and Markush template nodes in the free implementation knowledge graph, a directly accessible quantitative indicator can be provided for subsequent free implementation evaluation.
[0030] The symbolic constraint distance characterizes the coverage relationship from the perspective of substituent restrictions in the claim text. Based on the substituent positions and allowed substituent sets recorded in the aforementioned Markush template, the system traverses the target molecule structure, matching each substituent position in the molecule with its corresponding position in the template. For positions restricted by the template to "selected from a certain set of groups," the system checks whether the actual substituent in the molecule belongs to that set; for positions restricted by the template to "hydrogen or unsubstituted," the system checks whether there are substituents in the molecule that violate this restriction. The system counts the number of mismatches and the severity of the mismatch types among all substituent positions. For example, completely prohibited groups and groups that only exceed the carbon number range can be assigned different weights. The system then normalizes the distance based on the total number of substituent positions to obtain the symbolic constraint distance. The smaller this distance, the closer the molecule is to the template structure at the textual restriction level, and the higher the potential coverage risk.
[0031] Embedding spatial distance characterizes the covering relationship from the perspective of overall structural similarity. The system first generates a molecular graph representation for each molecule and each Markush template. This molecular graph representation can be achieved using molecular fingerprints or vector representations obtained from graph neural networks. Its core is encoding information such as atom type, bond type, and ring structure into fixed-length numerical vectors. The molecular vector corresponding to the Markush template can be obtained by encoding the template skeleton and typical substituent combinations. For any pair of molecular vectors and Markush template vectors, the system calculates the numerical distance between them based on a preset distance metric (e.g., Euclidean distance or cosine distance), which serves as the embedding spatial distance. Beyond symbol matching, the embedding spatial distance reflects the degree of similarity between the molecule and the template covering structure in terms of overall characteristics such as three-dimensional shape and electron distribution, mitigating the risk of imprecise similarity representation that is difficult to achieve with purely textual constraints.
[0032] Structural edit distance characterizes coverage relationships from the perspective of structural accessibility. The system predefines a set of allowed structural edit operations, including at least adding, deleting, or replacing substituents at specific sites, and changing some chemical bond types. For a target molecule, the system searches the set of structural edit operations, provided that chemical rationality is satisfied, to find an edit path that transforms the current molecular structure into a representative structure falling within the Markush template's defined range, and counts the required number of edit steps. When different types of edit operations have varying importance, weights can be assigned to different edit operations to calculate the structural edit distance in a weighted manner. A smaller structural edit distance indicates that the molecule can fall within the template's coverage range with fewer structural modifications, and the corresponding risk of freedom of implementation is more prominent; a larger structural edit distance indicates that a significant structural change is required, and the risk of patent coverage is relatively lower.
[0033] After calculating the symbolic constraint distance, embedding space distance, and structural edit distance, the system combines these three into a free implementation distance vector and writes it into the relationship edge attribute field between molecular nodes and Markush template nodes in the free implementation knowledge graph. The relationship edges explicitly mark the attribute names used to store the symbolic constraint distance, embedding space distance, and structural edit distance, ensuring that subsequent modules can read and use them through a unified interface. In practical applications, the free implementation evaluation agent can directly perform weighted aggregation based on the free implementation distance vector to obtain a free implementation risk score for ranking and filtering, without needing to re-parse the claims and structural formulas. This improves the efficiency and consistency of free implementation analysis and provides traceable basic data for subsequent causal analysis and parameter calibration.
[0034] Writing Markush templates and free implementation distance vectors into the free implementation knowledge graph includes: representing Markush templates as Markush template nodes, establishing relational edges between molecular nodes and Markush template nodes, recording free implementation distance vectors in the attributes of relational edges, and providing a query interface for reading free implementation distance vectors based on relational edge attributes.
[0035] In one possible implementation, after completing the Markush template generation and the free implementation distance vector calculation, the present invention uniformly writes these two types of results into the free implementation knowledge graph, so that any subsequent module can directly obtain the free implementation metric result between a certain molecule and a certain Markush template through graph query, without having to repeatedly parse the claims or repeat the calculation.
[0036] Specifically, the free implementation of the knowledge graph is based on a graph database. In this database, a Markush template node is first created for each Markush template. Each Markush template node contains at least a template identifier, the source patent publication number, the associated claim number, the molecular skeleton code, and a structured description of the substituent positions and substituent constraints. The molecular skeleton code can be represented using a linear structure or a pre-generated skeleton fingerprint, while the substituent constraints are recorded in list or key-value format as allowed substituent types and combination constraints at each position. In this way, Markush templates are no longer merely combinations of text and images, but are abstracted into knowledge units that can participate in graph queries and algorithmic processing.
[0037] After the molecular node and Markush template node are created, this invention establishes a relation edge between them to represent the free implementation association between the molecule and the Markush template. The relation edge uses a single relation type and records the three components of the free implementation distance vector in its attributes: symbolic constraint distance, embedding space distance, and structural edit distance. To improve subsequent query efficiency, the relation edge attributes use fixed field names and data types. For example, various distance values are stored as numeric fields, the distance vector calculation time is recorded as a timestamp field, and a boolean field indicates whether it is the latest calculation result, thus supporting batch updates and validity checks.
[0038] To support different analytical scenarios, this invention defines several query modes at the graph database level. In the molecular retrieval scenario, the query interface receives a molecular identifier, first locates the corresponding molecular node in the free-form knowledge graph, then traverses the relational edges pointing from the molecular node to the Markush template node, reads the free-form distance vector attribute of each relational edge, and returns a list of Markush templates related to the molecule and their corresponding distance vectors. In the patent layout evaluation scenario, the query interface receives a Markush template identifier, locates the target Markush template node in the graph, traverses the relational edges pointing to the molecular node, reads the free-form distance vector of each relational edge, and filters out molecular nodes with smaller symbolic constraint distances or smaller structural edit distances to identify potentially high-risk molecule sets.
[0039] The aforementioned query interface can be implemented using the query language provided by the graph database, or it can be encapsulated as a service interface at the application layer, providing a unified calling method for upper-layer agents and workflow engines. Agents only need to call the interface based on molecular identifiers or Markush template identifiers to obtain pre-calculated and stored free-implementation distance vectors in a single query, without needing to concern themselves with the underlying data structure and computation process, thus decoupling the algorithm module from the data storage module. By explicitly writing the Markush template and free-implementation distance vectors into the free-implementation knowledge graph and organizing them in the form of relational edge attributes, this invention improves the traceability and reusability of free-implementation analysis results, facilitating verification and validation in subsequent causal analysis, parameter calibration, and workflow replay processes.
[0040] A target analysis workflow is constructed based on the free implementation knowledge graph and the free implementation distance vector. The candidate molecule generation agent and the free implementation evaluation agent are called to obtain the free implementation risk assessment results and recommendation results of the candidate molecules. In one embodiment, after the free-implementation knowledge graph and free-implementation distance vector are constructed, the present invention organizes the analysis process based on the target dimension and constructs a target analysis workflow. The system receives the target identifier, indication information, and constraints input by the R&D personnel, generates a workflow context, and correspondingly queries the free-implementation knowledge graph for molecular nodes, Markush template nodes, and their free-implementation distance vectors associated with the target node, and writes the query results into the workflow context to form an initial molecular candidate set and risk feature set for a single target.
[0041] The target analysis workflow consists of multiple agents invoked sequentially. The candidate molecule generation agent is responsible for expanding the structural space based on the initial candidate molecule set. This agent reads target identifiers, known active molecules, and relevant Markush template information from the workflow context. For each known active molecule, it performs structural editing operations to generate a set of structural neighborhood molecules. Simultaneously, it can retrieve molecules with known activity to the target but not yet appearing in the internal pipeline from external chemical structure libraries and add highly similar molecules to the candidate set through embedding space similarity screening. The newly generated candidate molecules are registered as molecular nodes and temporarily stored in the workflow context, marked as pending evaluation.
[0042] The free implementation evaluation agent is scheduled after the candidate molecule generation agent completes its execution. Based on the candidate molecule identifier, the free implementation evaluation agent invokes the relation query interface in the free implementation knowledge graph to batch read the free implementation distance vector between each candidate molecule and its relevant Markush templates. For each candidate molecule, the free implementation evaluation agent combines the symbolic constraint distance, embedding space distance, and structural editing distance into a free implementation risk score according to preset weighting rules. Simultaneously, it combines information such as the candidate molecule's target and whether it belongs to an internal research structure series to generate a structured free implementation risk assessment result. The assessment result includes at least the risk score, a list of major high-risk Markush templates, and their contribution descriptions for subsequent manual or automated review.
[0043] After obtaining the freedom of implementation risk assessment results, this invention further calculates recommendation results in the workflow. The system sorts candidate molecules from low to high according to their freedom of implementation risk scores, and can overlay scores from other dimensions such as activity prediction scores and synthetic feasibility scores to select a group of candidate molecules with high comprehensive scores and low freedom of implementation risk as recommendation results, which are written into the workflow context and persistently stored. The recommendation results record the candidate molecule identifier, recommendation level, main supporting evidence, and reference relationship with nodes in the freedom of implementation knowledge graph, so that subsequent R&D decision-making systems, patent teams, and project management systems can directly call them. Through the construction and execution of the above target analysis workflow, this invention transforms the static relationships and freedom of implementation distance vectors in the freedom of implementation knowledge graph into a dynamic decision-making process oriented towards specific targets, realizing the linkage between candidate molecule generation and freedom of implementation assessment, significantly reducing the workload of manually comparing claims line by line, and improving the accuracy and traceability of early patent risk identification.
[0044] The target analysis workflow based on the free implementation knowledge graph and the free implementation distance vector includes: receiving the target identifier; generating a target-associated molecule set based on the molecular nodes and claim nodes associated with the target node in the free implementation knowledge graph; generating a workflow context; the workflow context is used to record the target identifier and the target-associated molecule set; writing the target identifier and the target-associated molecule set into the workflow context; and configuring the call interface for accessing the free implementation distance vector in the workflow context.
[0045] In one embodiment, after constructing the free-implementation knowledge graph and free-implementation distance vector, the system organizes the analysis process using the target as the entry point, constructing a target analysis workflow. Researchers or upper-level systems input target identifiers, which can be gene names, protein names, or internally unified target codes. Upon receiving the target identifier, the workflow scheduling module locates the corresponding target node in the free-implementation knowledge graph. Based on a graph database query language, it traverses outwards along relational edges such as "function" and "involved targets" to obtain molecular nodes and claim nodes associated with the target node, and organizes the molecular identifiers corresponding to these molecular nodes into a target-associated molecule set.
[0046] To facilitate the sharing of the same batch of data among multiple agents, this invention defines a workflow context as a state container oriented towards a single target. The workflow context can be implemented using a key-value structure or a document structure, and at least includes a target identifier field and a target-associated molecule set field. It can also be extended to record auxiliary information such as indication information, screening conditions, and analysis timestamps. When generating the workflow context, the workflow scheduling module writes the input target identifier into the target identifier field, writes the target-associated molecule set obtained from the free implementation knowledge graph into the corresponding field, and simultaneously establishes a connection configuration with the free implementation knowledge graph, enabling subsequent agents to access molecular nodes, Markush template nodes, and their relationships within the knowledge graph through a unified entry point.
[0047] To enable the candidate molecule generation agent and the free implementation evaluation agent to directly utilize pre-calculated free implementation distance vectors, this invention configures a call interface for accessing these vectors within the workflow context. This call interface encapsulates graph database query logic and result format conversion logic. The interface input includes at least the molecule identifier and optional Markush template filtering conditions. The interface output is a list of free implementation distance vectors associated with the molecule, along with the corresponding Markush template identifiers. The workflow context only records the configuration parameters and access handle of the call interface; it does not directly copy the specific values of the free implementation distance vectors, thus avoiding data redundancy.
[0048] When the subsequent candidate molecule generation agent joins the target analysis workflow, it first reads the target identifier and the set of target-related molecules from the workflow context. Based on this, it performs structure editing and structure expansion to generate new candidate molecules. After generation, the new candidate molecules are added to the candidate molecule set in the workflow context, and the free implementation distance vectors of the new candidate molecules in the free implementation knowledge graph are queried through the API. For molecules that do not yet have distance vectors, the distance vector calculation module is triggered to perform supplementary calculations, and the results are returned in a unified format through the API. When the free implementation evaluation agent performs the evaluation task, it also obtains the set of target-related molecules and the set of candidate molecules through the workflow context, and uniformly obtains the free implementation distance vectors between all molecules and the relevant Markush templates in batches through the API for subsequent risk scoring and ranking.
[0049] By employing the above methods, this invention decouples the target analysis workflow from the free-implementation knowledge graph and the free-implementation distance vector. On one hand, the workflow context centrally records target-level inputs and intermediate results, ensuring that the candidate molecule generation agent and the free-implementation evaluation agent work collaboratively under the same data view. On the other hand, the interface for accessing the free-implementation distance vector shields the underlying graph database structure and query details, allowing new agents or external systems to reuse free-implementation metric results without altering the graph structure. This design improves the configurability and scalability of the target analysis process, facilitating the reuse of workflow templates across different targets and projects, while ensuring the consistency and traceability of free-implementation analysis results.
[0050] The candidate molecule generation agent generates a set of candidate molecules by performing structural editing operations on the molecular structures corresponding to the molecular nodes in the workflow context. The structural editing operations include at least one of adding substituents to the molecular backbone, replacing substituents, and changing the type of chemical bonds.
[0051] The candidate molecule generation agent is deployed within the target analysis workflow to automatically expand a set of structurally sound candidate molecules around molecular nodes in the workflow context, without deviating from existing chemical knowledge and patent constraints. The workflow context already records target identifiers and sets of target-associated molecules. Each target-associated molecule corresponds to a molecular node in the free-implementation knowledge graph and contains a machine-readable representation of its molecular structure. The candidate molecule generation agent first reads this set of molecular nodes from the workflow context and selects a starting molecule for structure editing according to preset rules.
[0052] For each starting molecule, the candidate molecule generation agent invokes cheminformatics tools to break down the molecular structure into a molecular backbone and substituent regions. The molecular backbone is used to maintain the overall skeletal type and core pharmacophore, while the substituent regions serve as editable areas. Based on the target type, historical active molecule characteristics, and Markush template information related to the molecule in the free-implementation knowledge graph, the agent determines the editable atomic positions and bonds, generating a "list of editable sites" to avoid editing in unsuitable areas such as metal coordination centers and key aromatic ring bridging positions.
[0053] After determining the editable sites, the candidate molecule generation agent executes structure editing operations one by one according to a predefined library of structure editing operations. For the operation of "adding substituents to the molecular skeleton," the agent selects substituent types that satisfy basic medicinal chemistry rules from the substituent library, such as halogens, alkyl groups, and nitrogen-containing heterocycles. It then filters out substituent combinations that clearly violate patent limitations or chemical rationality by combining these with the allowed substituent set in the Markush template, adding only substituents that conform to the rules at the editable sites to generate new molecular structures. For the operation of "replacing substituents," the agent identifies existing substituents in the starting molecule and replaces them with substituents that have similar physicochemical properties or are not covered in the Markush template to explore potential spaces for free implementation. For the operation of "changing chemical bond types," the agent performs limited switching between single, double, and aromatic bonds at allowed bond positions, ensuring the legality of valence states and bond numbers, and excluding obviously unstable or high-stress structures through simple rules.
[0054] After each structure editing operation, the candidate molecule generation agent performs basic validity checks on the newly generated molecules, including valence state checks, structural connectivity checks, and simple physicochemical property constraint checks. Only molecules that pass the checks are added to the candidate molecule set. To avoid a large number of duplicate structures, the agent converts the molecular structure into a uniform linear structural representation before adding it to the set. This representation is then compared with already generated candidate molecules using hashing or indexing. If a duplicate representation already exists, the current result is discarded, thereby controlling the size of the candidate molecule set and reducing the burden of subsequent evaluation.
[0055] Each generated candidate molecule is registered as a new molecular node in the free implementation knowledge graph, recording its molecular structure representation, originating molecule identifier, and target identifier. Simultaneously, a candidate molecule set field is added to the workflow context, recording the identifiers of these new molecules and their mapping relationships with the originating molecule. This facilitates the free implementation evaluation agent in directly calling the free implementation distance vector query interface based on the candidate molecule set in the next stage. Through this design, the candidate molecule generation agent, while maintaining the basic characteristics of the molecular skeleton, explores substituent and bond types in a controlled manner, expanding the optional chemical space and providing a diverse and chemically sound foundation of candidate molecules for subsequent free implementation risk assessment.
[0056] The free implementation evaluation agent, based on the candidate molecule set generated by the candidate molecule generation agent and the free implementation distance vector stored in the free implementation knowledge graph, performs weighted aggregation of the free implementation distance vectors corresponding to a set of Markush templates for each candidate molecule to obtain the free implementation risk assessment result of the candidate molecule.
[0057] After receiving the candidate molecule set written into the workflow context by the candidate molecule generation agent, the free implementation evaluation agent processes each candidate molecule as an evaluation unit. For a given candidate molecule, the free implementation evaluation agent first reads the molecule identifier and target identifier from the workflow context, and then accesses the free implementation knowledge graph through a pre-configured query interface. It then reads in batches the free implementation distance vectors related to the candidate molecule, as well as the corresponding Markush template identifiers and source claim information, from the edges of the relationships between molecule nodes and Markush template nodes. If a candidate molecule does not yet have a complete free implementation distance vector, the free implementation evaluation agent calls the distance calculation module to complete the missing distance components and writes the new calculation results back to the free implementation knowledge graph to ensure data consistency for subsequent evaluations.
[0058] After obtaining the free implementation distance vectors corresponding to candidate molecules and a set of Markush templates, the free implementation evaluation agent needs to convert the symbolic constraint distance, embedding space distance, and structural editing distance into comparable risk contribution values. To this end, the system performs a monotonic transformation on each distance component based on a preset threshold range or empirical distribution, mapping the distance values to risk contribution scores between zero and one, where smaller distances correspond to higher risk contributions, and larger distances correspond to lower risk contributions. For each combination of a candidate molecule and a Markush template, the free implementation evaluation agent calculates the single-template contribution value of the Markush template to the free implementation risk of the candidate molecule based on the risk contribution scores of the three types of distances and their respective importance weights.
[0059] After obtaining the single-template contribution values of all relevant Markush templates, the free implementation evaluation agent aggregates the risk contributions of the same candidate molecule to generate a free implementation risk score for that candidate molecule. The aggregation process can employ a linear weighted summation method, with the core calculation expression being:
[0060] in, as candidate molecules Risk assessment of free implementation as candidate molecules Number of associated Markush templates as candidate molecules Compared to the first The risk contribution value of a Markush template. The free implementation assessment agent can configure the weights of symbolic constraint distance, embedding space distance, and structural editing distance in the risk contribution value according to the risk sensitivity requirements of different drug projects, so that the risk score is more in line with the actual free implementation review focus.
[0061] After the risk score for free implementation is calculated, the free implementation evaluation agent classifies candidate molecules into three levels—high risk, medium risk, and low risk—based on preset grading rules. These grading rules can use fixed thresholds, such as statistically analyzing typical threshold ranges based on successful or failed free implementation cases in historical projects, or they can be adjusted appropriately based on target importance and project stage. For candidate molecules with risk scores falling in the high-risk range, the free implementation evaluation agent includes a list of Markush templates for the main risk sources and explanations of key distance components in the output, allowing the patent team to conduct targeted manual review. For candidate molecules with risk scores in the low-risk range, they are marked as priority recommendations in the output.
[0062] After the evaluation is completed, the free implementation evaluation agent writes the free implementation risk score, risk level, and a summary of the main risk sources for each candidate molecule into the workflow context. Simultaneously, it updates the molecule node attributes in the free implementation knowledge graph, recording the latest evaluation time and the most recent risk score. The upper-level R&D decision-making system or candidate molecule recommendation module can directly filter and sort based on the evaluation results in the workflow context without needing to access the underlying distance vector again. This significantly simplifies the integration of free implementation analysis into the overall R&D process and ensures the traceability and interpretability of the evaluation process.
[0063] A causal analysis model is established based on workflow logs, patent examination results, and infringement judgment results. The aggregation method of the free implementation distance vector and the free implementation risk assessment rules are then calibrated according to the causal analysis model.
[0064] The system constructs a causal analysis model through retrospective analysis of historical R&D project data to calibrate the aggregation method of the free implementation distance vector and the free implementation risk assessment rules. During workflow execution, the target analysis workflow, candidate molecule generation agent, and free implementation assessment agent continuously write to the workflow log. The workflow log records at least the candidate molecule identifier, target identifier, corresponding free implementation distance vector, aggregation weight used at that time, obtained free implementation risk score and risk level, whether it is recommended to enter the next R&D stage, and the assessment time. Over time, for the same candidate molecule or its series of molecules, patent application examination results and related infringement judgments will be generated successively. These results are imported into the system through an external interface and a correspondence is established with the candidate molecule identifier, patent application number, and claim scope.
[0065] To establish a causal analysis model, the system first aligns workflow logs, patent examination results, and infringement judgment results into sample data. Each sample corresponds to a freedom assessment decision and its subsequent actual outcome. The sample includes the values of each distance component in the freedom distance vector, the aggregation weights used at the time, the obtained freedom risk score and level, and the corresponding results such as whether the patent examination was rejected or whether an infringement judgment was established. Based on business experience, the system pre-sets some causal structures in the causal graph. For example, the freedom distance vector affects the freedom risk score; the freedom risk score and the project team's recommendation jointly affect whether a patent application is filed; patent application behavior and claim drafting strategy jointly affect the examination result; and certain components in the freedom distance vector are directly related to the infringement judgment result.
[0066] After the causal structure is determined, the system uses statistical learning methods to estimate the parameters of the causal analysis model. This enables the model to provide trends in the probability of patent rejection and the probability of infringement judgment under different free implementation distance vectors, different aggregation methods, and different risk threshold settings. By comparing the deviation between the model's predictions and historical actual results, the system calculates the systematic deviations of the current aggregation method and risk assessment rules. For example, if it finds that the impact of symbolic constraint distance on the rejection result is underestimated, while the impact of structural edit distance on the infringement judgment result is overestimated, the causal analysis model will output corresponding adjustment suggestions.
[0067] Based on the output of the causal analysis model, the system calibrates the aggregation method of the free-implementation distance vector, including adjusting the aggregation weights of each distance component, whether to introduce nonlinear transformations, and whether to set truncation or segmentation rules for certain distance components. Simultaneously, the free-implementation risk assessment rules are revised, including adjusting the thresholds for risk score classification and setting separate risk weight coefficients for specific targets or specific types of Markush templates. After calibration, the system writes the new aggregation weights, risk thresholds, and specific rules into the configuration library, which are then loaded and used by the free-implementation assessment agent in the next workflow execution. By continuously incorporating actual review results and infringement judgment results for causal feedback, this invention can dynamically correct the aggregation method of the free-implementation distance vector and the risk assessment rules, making the free-implementation risk score more consistent with real legal risks and reducing misjudgments and omissions caused by assessment bias.
[0068] The causal analysis model uses the results of the risk assessment of free implementation of candidate molecules, the recommendation results of candidate molecules, the patent examination results, and the infringement judgment results as nodes in the causal graph. Based on the workflow log, it constructs the statistical relationship between the results of the risk assessment of free implementation of candidate molecules and the results of patent examination rejection and infringement judgment.
[0069] In one embodiment, based on a long-term target analysis workflow, the present invention utilizes accumulated historical data to construct a causal analysis model to assess the true correlation between the freedom of implementation risk score and subsequent legal outcomes. To this end, the system first extracts structured evaluation records from the workflow log. Each record includes at least the candidate molecule identifier, the current freedom of implementation risk assessment result (specifically, a numerical score and risk level), whether it was adopted as a candidate molecule recommendation by the system or manually, the corresponding patent application number, and the evaluation time. Subsequently, the system synchronizes patent examination results and infringement judgment results from an external patent examination system and litigation management system, including information such as whether the application was rejected, a summary of the reasons for rejection, whether infringement was found, and the judgment time. Through the patent application number and the scope of the claims, the system establishes a one-to-one or one-to-many correspondence between these results and the candidate molecules.
[0070] After data alignment, the system constructs a causal graph based on these historical records. Nodes in the causal graph include at least the following: candidate molecule free implementation risk assessment result node, candidate molecule recommendation result node, patent examination result node, and infringement judgment result node. The candidate molecule free implementation risk assessment result node represents the risk score and risk level given by the system during the assessment; the candidate molecule recommendation result node indicates whether the candidate molecule was recommended to enter the subsequent R&D process during the assessment stage; the patent examination result node indicates whether the examination authority made a rejection decision for the patent application related to the candidate molecule; and the infringement judgment result node indicates whether a judicial result finding infringement occurred subsequently for the candidate molecule or its related compounds.
[0071] The directed edges in the causal graph are determined jointly by business logic and data statistics. For example, the results of the freedom to implement risk assessment directly affect the candidate molecule recommendation results, because project teams usually refer to risk scores when deciding whether to proceed with an application; the candidate molecule recommendation results have a prerequisite constraint on whether a patent examination result is generated, because candidates that are not recommended generally do not enter the patent application process; the correlation between the freedom to implement risk assessment results and the patent examination rejection results and infringement judgment results is modeled using the correlation strength statistically derived from historical data. When establishing the causal analysis model, the system not only records whether there is a causal relationship between these nodes, but also statistically analyzes the patent examination rejection rate and infringement judgment rate corresponding to different risk levels based on workflow logs, thereby forming a statistical relationship between the freedom to implement risk assessment results and legal results.
[0072] To avoid misleading models with simple correlations, this invention utilizes time information from workflow logs to strictly sort the risk assessment time, recommendation decision time, patent examination conclusion time, and infringement judgment time, retaining only causally possible temporal relationships. Multiple assessment records of the same candidate at different time points are merged or filtered to ensure each sample reflects a complete "assessment-decision-outcome" chain. Based on the processed samples, the system estimates the conditional distribution from the free implementation risk assessment result node to the patent examination rejection result node and the infringement judgment result node in the causal graph, describing the frequency of subsequent rejections or infringement determinations within a specific risk score range. Through the above causal analysis model, this invention can quantitatively characterize the consistency between the free implementation risk score and the actual legal risk, providing a reliable evidentiary basis for subsequent calibration of the free implementation distance vector aggregation method and risk assessment rules. This allows the system to gradually approach the risk judgment standards of actual examination practice and judicial decisions during continuous operation.
[0073] The calibration of the free implementation distance vector aggregation method and free implementation risk assessment rules based on the causal analysis model includes: adjusting the aggregation weight of each distance component in the free implementation distance vector and the judgment threshold of the free implementation risk assessment results based on the deviation between the candidate molecule free implementation risk assessment results output by the causal analysis model and the patent examination rejection results and infringement judgment results; and writing the adjusted aggregation weight and judgment threshold into the configuration library for subsequent target analysis workflow execution.
[0074] In one embodiment, after the causal analysis model is constructed, the present invention periodically calibrates the aggregation method of the free implementation distance vector and the free implementation risk assessment rules in an offline batch processing manner. The system first extracts samples that have completed closed-loop processing from the workflow log. Each sample includes at least the free implementation risk assessment result of the candidate molecule, the candidate molecule recommendation result, the corresponding patent examination result, and the corresponding infringement judgment result. Through the causal analysis model, the probability of patent rejection and infringement judgment being established under different free implementation risk assessment result intervals is calculated, and the deviation between this probability and the current risk level definition is quantified as an assessment deviation index, used to guide subsequent parameter adjustments.
[0075] During the aggregation calibration process, the system treats the symbolic constraint distance, embedding space distance, and structural editing distance in the free implementation distance vector as three distance components with adjustable weights. The causal analysis model, based on historical samples, assesses the sensitivity of each distance component to the impact of patent rejection and infringement judgments. For example, it finds that the symbolic constraint distance has a higher correlation with patent rejection results, while the structural editing distance plays a crucial role in infringement judgments. Based on this, the system generates weight adjustment suggestions, appropriately increasing the weight of distance components that contribute more to predicting patent rejection results and appropriately decreasing the weight of distance components that contribute less or have higher noise levels. Each candidate weight combination is recalculated on historical samples to determine the free implementation risk assessment results, and compared with the actual patent rejection and infringement judgment results. The combination with the smaller overall error is selected as the new aggregation weight.
[0076] Regarding the calibration of the judgment threshold for the free implementation of risk assessment results, the system, based on the output of the causal analysis model, groups historical samples according to the current risk level and statistically analyzes the proportion of patent rejections and infringement judgments at each level. If it is found that the proportion of rejections or infringements is too low in the high-risk level, while a large number of rejections or infringement cases appear in the medium and low-risk levels, it indicates that the current risk level classification threshold is set too broadly. By moving the level classification point on the risk scoring axis, the system explores the actual legal risk distribution corresponding to each level under different threshold combinations, prioritizing the selection of threshold combinations that allow the high-risk level to concentrate on covering most rejections and infringement cases, and the low-risk level to have a low misjudgment rate. The above weight adjustments and threshold adjustments can be performed together to form multiple candidate parameter schemes, and the scheme with the best performance is selected through replay evaluation on historical samples.
[0077] Once the new aggregation weights and decision thresholds are determined, the system writes them into the configuration library in a structured configuration format. The configuration library records the parameter version number, effective time, applicable target range, and the version information of the causal analysis model on which the generation is based. When the target analysis workflow starts, it reads the latest effective aggregation weights and decision thresholds from the configuration library, loads them by the free implementation evaluation agent, and applies them to the subsequent free implementation risk assessment process for candidate molecules. By continuously incorporating patent examination results and infringement judgment results for causal-driven parameter calibration, this invention enables the aggregation method of the free implementation distance vector and the free implementation risk assessment rules to be adaptively updated over time, making the assessment results closer to real patent examination practices and judicial decisions, and reducing the risk of systemic misjudgment.
[0078] In one specific embodiment, focusing on small molecules related to the kinase target HER2, a complete process was constructed from target input to free-implementation risk assessment and rule calibration. The project team selected a batch of 60 molecules from its internal R&D database, all with existing in vitro activity data for HER2. Of these, 12 were marketed or investigational clinical candidates, and 48 were early lead and candidate molecules. Simultaneously, 95 structural claims related to HER2 were screened using a patent search system. After structural identification and Markush parsing, 32 Markush templates were obtained. Importing the molecular data and patent text data into the free-implementation knowledge graph resulted in 60 molecular nodes, 95 claim nodes, 32 Markush template nodes, and approximately 420 relationship edges between molecular nodes and Markush template nodes. Each relationship edge records the corresponding free-implementation distance vector.
[0079] Based on this data, a target analysis workflow is initiated for the HER2 target. The workflow receives the target identifier "HER2" and searches the free implementation knowledge graph for molecular nodes and claim nodes that have a "function" or "involvement with target" relationship with the HER2 target node, resulting in a set of 60 target-related molecules. The target identifier and these 60 molecule identifiers are written into the workflow context, and an interface for accessing the free implementation distance vector is configured. Subsequent agents will collaborate within this context.
[0080] The candidate molecule generation agent performed one round of structural editing on the aforementioned 60 molecules. Based on the known structure-activity relationship of HER2 kinase inhibitors, the outer part of the aromatic ring and the replaceable side chain were designated as editable regions, while modifications to the core heteroaromatic ring and the position of key hydrogen bond acceptors were prohibited. At each editable site, five types of substituents—fluorine, chlorine, methoxy, cyano, and morpholino—were selected from a pre-defined substituent library for combination substitution, while limiting each molecule to a maximum of two editable sites to ensure that the structural changes were not excessive. After validity verification and duplicate removal, a total of 180 new candidate molecules were generated. Combined with the original 60 molecules, the total candidate molecule set reached 240, all of which were registered as molecular nodes in the knowledge graph.
[0081] like Figure 2 As shown, the free implementation analysis process of this invention is divided into four stages from left to right. In stage 1, the upper database stores chemical structure data, with various molecular structural formulas and numbers such as "60, 95, 32" indicating the number of different molecular samples on the side; the lower database stores patent text data, with multiple cards in front indicating the claim text and structural fragments. Stage 2 is the free implementation knowledge graph construction area. Circular nodes represent molecular nodes (labeled as A, B, etc.), square nodes represent claim nodes, and hexagonal nodes represent Markush template nodes (labeled as M1, M2, M3). The nodes are connected by lines, with values such as 0.10, 0.12, 0.25, 0.40, and 0.70 marked at some of the connections to represent different components of the free implementation distance vector. Phase 3 is the execution area of the agent workflow. The agent icon and gear symbol on the left represent the agent that generates candidate molecules. After receiving molecular information from the knowledge graph, it generates candidate molecular structures marked with the letters D and E. The fan-shaped dashboard in front of the agent icon on the right is divided into multiple arcs, with values such as 0.35, 0.75, and 1.10 marked on the arcs, representing different risk scores output by the free implementation evaluation agent. Phase 4 is located at the bottom. The small node diagram composed of H, M, L, and 1 illustrates the causal analysis model's modeling of the relationship between high, medium, and low risks and evaluation results. The large loop arrows from Phase 3 and Phase 4 point back to Phase 2, indicating a closed-loop calibration of the free implementation knowledge graph and evaluation rules based on feedback such as review results.
[0082] The free implementation evaluation agent invoked the query interface for each of the 240 candidate molecules to read their free implementation distance vectors with 32 Markush templates. Taking five representative candidate molecules as examples, the distance components of the three Markush templates most relevant to the coverage were statistically analyzed. Candidate molecule A had symbolic constraint distances of 0.10, 0.15, and 0.40 on the three Markush templates, embedding space distances of 0.25, 0.30, and 0.45, and structural editing distances of 2, 3, and 5 steps, which, after normalization, were converted to 0.20, 0.35, and 0.70, respectively. Candidate molecule B had relatively large distances across the three categories, with a symbolic constraint distance above 0.60, an embedding space distance around 0.70, and a structural editing distance of approximately 0.80 after conversion. Candidate molecule C had a symbolic constraint distance close to 0.05 on a certain Markush template, an embedding space distance of approximately 0.20, and a structural editing distance of approximately 0.15 after conversion, indicating a high degree of similarity to that template. Candidate molecules D and E fell between the above two categories.
[0083] The free implementation evaluation agent uses preset aggregation weights to convert the three types of distance components between each pair of candidate molecules and the Markush template into single-template risk contribution values. Then, it sums the contribution values of all relevant templates to obtain the free implementation risk score for each candidate molecule. Taking candidate molecule A as an example, the single-template risk contributions of symbolic constraint distance, embedding space distance, and structural editing distance after weighting are approximately 0.55, 0.50, and 0.35, respectively. After summing the three templates, the free implementation risk score is approximately 1.40, corresponding to a medium-high risk level. Candidate molecule B has low risk contributions on all templates, with a comprehensive score of approximately 0.35, and is classified as low-risk. Candidate molecule C is mainly "locked" by a core Markush template, with a contribution value close to 0.90. The contributions of other templates are relatively small, resulting in a final score of approximately 1.10, which is considered high-risk. Candidate molecules D and E have scores of approximately 0.75 and 0.65, respectively, falling into the medium-risk range.
[0084] The core calculation process of the evaluation module can be expressed by the following pseudocode (omitting exception handling and database connection details): defaggregate_fto_risk(dist_vectors,w_symbol,w_embed,w_edit): risk=0.0 forvindist_vectors:#v={"symbol":...,"embed":...,"edit":...} r_symbol=1.0-v["symbol"] r_embed=1.0-v["embed"] r_edit=1.0-v["edit"] r_single=w_symbol*r_symbol+w_embed*r_embed+w_edit*r_edit risk += max(r_single, 0.0) returnrisk In the pseudocode above, `dist_vectors` is the set of distance components corresponding to the candidate molecule and a set of Markush templates, `w_symbol`, `w_embed`, and `w_edit` are the aggregate weights of the three types of distance components, and the free implementation risk score is the sum of the risk contributions of each template. In the initial settings of this embodiment, the weights of symbolic constraint distance, embedding space distance, and structural editing distance are set to 0.5, 0.3, and 0.2, respectively.
[0085] To verify the effectiveness of the method of this invention, historical data from the HER2 project over the past three years were incorporated into the causal analysis model. In the historical data, a total of 38 molecules entered the patent application stage, of which 14 received examination opinions during examination rejecting the claims due to existing HER2-related Markush structures, and 6 were found to have structural coverage risks in foreign litigation. Figure 3 As shown, the three bars on the left represent the number of patent rejections or infringement judgments actually rendered in the high-risk (H), medium-risk (M), and low-risk (L) samples before calibration, with values of 62, 33, and 12 respectively. The three bars on the right represent the corresponding numbers adjusted to 79, 29, and 5 in the same group after calibrating the weights of the freedom of implementation distance vector aggregation and the risk threshold using a causal analysis model. This indicates that high-risk samples are further concentrated in the high-risk group, and the number of misjudged samples in the low-risk group has significantly decreased, verifying the accuracy improvement effect of the calibration mechanism of this invention on freedom of implementation risk assessment. Comparing the freedom of implementation risk scores calculated using the method of this invention for these molecules in that year with the actual examination and judgment results: under the initial weight settings, the proportion of high-risk molecules that were actually rejected or judged as infringing was approximately 62%, medium-risk was 33%, and low-risk was 12%, indicating a correlation between the scores and the actual risks. However, there were still some "false positives" in the high-risk group, and a certain proportion of missed judgments in the medium- and low-risk groups.
[0086] Based on the aforementioned biases, the causal analysis model optimizes the aggregation weights and risk thresholds. The model recommends increasing the weight of symbolic constraint distance to 0.6, slightly decreasing the weight of embedding space distance to 0.25, adjusting the weight of structural edit distance to 0.15, lowering the threshold for high-risk classification from 1.0 to 0.9, and increasing the threshold between medium and low risk from 0.6 to 0.7. Re-evaluating on the same batch of historical samples using the new parameters, the proportion of high-risk molecules rejected or found infringing increased to approximately 79%, medium-risk to 29%, and low-risk to approximately 5%. In other words, after calibration, high-risk molecules more effectively cover those with actual serious patent obstacles, while the number of actual infringement cases among low-risk molecules significantly decreases, better meeting the practical requirements of freedom of execution analysis.
[0087] After writing the optimized aggregation weights and risk thresholds into the configuration library, the current 240 candidate molecules for the HER2 project were re-evaluated for freedom of implementation, and a comprehensive ranking was performed by overlaying activity and druggability indicators. Ultimately, 14 molecules entered the quadrant of "high activity and low risk of freedom of implementation," of which 3 were original lead molecules and the remaining 11 were new structures automatically generated by the candidate molecule generation agent. Compared with the project team's traditional screening results based on manual novelty searches and experience-based judgment, there were 6 overlapping molecules, and another 8 molecules were additionally identified as "low-risk, high-value" candidates by the method of this invention. Subsequent manual searches and patent attorney reviews of these 8 molecules showed that their freedom of implementation conclusions were basically consistent with the evaluation results of this invention, proving that this invention can provide verifiable and traceable freedom of implementation risk analysis results in real-world R&D scenarios, and improve the patent compliance level of the candidate molecule screening stage while saving manpower and time costs.
[0088] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for constructing a free-implementation analytical knowledge base based on intelligent agents and workflows, characterized in that, Includes the following steps: A free-implementation knowledge graph is constructed, which establishes the relationships between molecular nodes, claim nodes, and target nodes based on chemical structure data and patent text data. Markush templates are generated based on the structured claims in the free implementation knowledge graph. The free implementation distance vector of each molecule relative to each Markush template is calculated, and the Markush template and the free implementation distance vector are written into the free implementation knowledge graph. A target analysis workflow is constructed based on the free implementation knowledge graph and the free implementation distance vector. The candidate molecule generation agent and the free implementation evaluation agent are called to obtain the free implementation risk assessment results and recommendation results of the candidate molecules. A causal analysis model is established based on workflow logs, patent examination results, and infringement judgment results. The aggregation method of the free implementation distance vector and the free implementation risk assessment rules are then calibrated according to the causal analysis model.
2. The method according to claim 1, characterized in that, The construction of a free-implementation knowledge graph includes: establishing molecular nodes based on chemical structure data, where molecular nodes record molecular structural formulas and molecular identifiers; establishing claim nodes based on patent text data, where claim nodes record claim texts and publication numbers; and establishing associations between molecular nodes, claim nodes, and target nodes through a graph database.
3. The method according to claim 1, characterized in that, Generating Markush templates based on claims with structured representations in a free-implementation knowledge graph includes: identifying structural reference markers in the claim text, extracting structural images corresponding to the structural reference markers, converting the structural images into molecular skeleton diagrams and substituent position sets using a structural recognition algorithm, and storing the molecular skeleton diagrams and substituent position sets as Markush templates.
4. The method according to claim 1, characterized in that, The free implementation distance vector includes symbolic constraint distance, embedding space distance, and structural editing distance. The symbolic constraint distance is calculated based on the substituent constraints in the Markush template and the matching results with the molecular structure. The embedding space distance is calculated based on the distance between the molecular vector generated by the molecular graph representation and the corresponding molecular vector in the Markush template. The structural editing distance is calculated based on the minimum number of editing steps required to perform structural editing operations on the molecular graph to meet the Markush template constraints.
5. The method according to claim 1, characterized in that, Writing Markush templates and free implementation distance vectors into the free implementation knowledge graph includes: representing Markush templates as Markush template nodes, establishing relational edges between molecular nodes and Markush template nodes, recording free implementation distance vectors in the attributes of relational edges, and providing a query interface for reading free implementation distance vectors based on relational edge attributes.
6. The method according to claim 1, characterized in that, The target analysis workflow based on the free implementation knowledge graph and the free implementation distance vector includes: receiving the target identifier; generating a target-associated molecule set based on the molecular nodes and claim nodes associated with the target node in the free implementation knowledge graph; generating a workflow context; the workflow context is used to record the target identifier and the target-associated molecule set; writing the target identifier and the target-associated molecule set into the workflow context; and configuring the call interface for accessing the free implementation distance vector in the workflow context.
7. The method according to claim 6, characterized in that, The candidate molecule generation agent generates a set of candidate molecules by performing structural editing operations on the molecular structures corresponding to the molecular nodes in the workflow context. The structural editing operations include at least one of adding substituents to the molecular backbone, replacing substituents, and changing the type of chemical bonds.
8. The method according to claim 1, characterized in that, The free implementation evaluation agent, based on the candidate molecule set generated by the candidate molecule generation agent and the free implementation distance vector stored in the free implementation knowledge graph, performs weighted aggregation of the free implementation distance vectors corresponding to a set of Markush templates for each candidate molecule to obtain the free implementation risk assessment result of the candidate molecule.
9. The method according to claim 1, characterized in that, The causal analysis model uses the results of the risk assessment of free implementation of candidate molecules, the recommendation results of candidate molecules, the patent examination results, and the infringement judgment results as nodes in the causal graph. Based on the workflow log, it constructs the statistical relationship between the results of the risk assessment of free implementation of candidate molecules and the results of patent examination rejection and infringement judgment.
10. The method according to claim 9, characterized in that, The calibration of the free implementation distance vector aggregation method and free implementation risk assessment rules based on the causal analysis model includes: adjusting the aggregation weight of each distance component in the free implementation distance vector and the judgment threshold of the free implementation risk assessment results based on the deviation between the candidate molecule free implementation risk assessment results output by the causal analysis model and the patent examination rejection results and infringement judgment results; and writing the adjusted aggregation weight and judgment threshold into the configuration library for subsequent target analysis workflow execution.