Formula performance prediction method based on component knowledge graph

By constructing an ingredient knowledge graph based method, a set of argumentation rules is generated, an attack relationship graph is constructed, a minimum inconsistency set is determined, and candidate solutions are solved and falsified. This solves the problem of stability and reliability of formulation performance in the R&D of daily chemical products, and reduces trial and error costs and mass production risks.

CN122089151APending Publication Date: 2026-05-26저장 아얀 바이오텍 컴퍼니 리미티드
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
저장 아얀 바이오텍 컴퍼니 리미티드
Filing Date
2026-02-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In the development of daily chemical products, existing technologies cannot guarantee the stability of formulation performance, batch-to-batch consistency and process scalability. Furthermore, when testing protocols, packaging and storage conditions or process windows change, the model predictions lack interpretable boundaries and counterexamples, causing candidate solutions to fail under extreme conditions.

Method used

A component knowledge graph-based approach is adopted. By constructing a component knowledge graph, generating a set of argumentation rules, and building an attack relationship graph, the minimum inconsistency set is determined, candidate solutions are solved, and falsification verification is performed to achieve reliable extrapolation and iterative optimization.

Benefits of technology

It achieves structured accumulation of component relationships and conditional dependencies under data sparsity and changing conditions, reducing trial-and-error costs and mass production failure risks, and improving the stability and reliability of candidate solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089151A_ABST
    Figure CN122089151A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of formula design and process optimization, in particular to a formula performance prediction method based on a component knowledge graph, and the method comprises the steps: obtaining formula batch data; constructing a component knowledge graph based on the batch data, and generating a demonstration rule set containing a trigger condition, a threshold judgment condition and a retrospect condition; constructing an attack relation graph based on the argumentation rule set, determining a minimum inconsistent set and generating a constraint set; solving in the candidate formula and the candidate process window according to the constraint set to obtain a candidate scheme and a performance prediction result; and carrying out security verification on the candidate schemes, generating disturbance samples, carrying out consistency verification, directionally searching counter-example records, updating the argumentation rule set and the constraint set based on the counter-example records, and outputting an updated performance prediction result. According to the method, a closed loop of conflict resolution, constraint solution and witness iteration is realized, and the extrapolation stability and extreme condition robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of formulation design and process optimization technology, specifically to a method for predicting formulation performance based on component knowledge graphs. Background Technology

[0002] In the R&D of daily chemical products, the stability of formulation performance, batch-to-batch consistency, and process scalability directly affect the R&D cycle, compliance verification, and mass production risks. Existing methods often rely on trial and error based on experience or black-box predictions based on regression models, which often fail to explicitly express "ingredient compatibility, substitutability, and process conditions" as auditable, structured knowledge. When testing protocols, packaging and storage conditions, or process windows change, models are prone to distribution shifts, and prediction conclusions lack interpretable boundaries and counterexamples, causing candidate solutions to fail under extreme conditions. At the same time, rule conflicts and inconsistent evidence are difficult to systematically locate and resolve, making it difficult for R&D teams to form reusable knowledge accumulation and closed-loop iteration mechanisms. Summary of the Invention

[0003] This invention provides a formulation performance prediction method based on component knowledge graphs, which at least addresses the problem of how to reliably extrapolate candidate formulations and process windows under conditions of data sparsity and changing conditions, and how to output stable performance prediction results that can be falsified iteratively.

[0004] This invention provides a method for predicting formulation performance based on component knowledge graphs, the method comprising: Obtain batch data for the formulation, including formulation composition data, process parameter data, packaging and storage condition data, testing protocol data, and test result data; A component knowledge graph is constructed based on batch data, and a set of argumentation rules is generated, which includes triggering conditions, threshold judgments, and rebuttal conditions. An attack relationship graph is constructed based on the set of argumentation rules, and a minimum inconsistency set is determined. A constraint set is generated based on the minimum inconsistency set. Candidate solutions are obtained by solving within the candidate formula and candidate process window based on the constraint set, and performance prediction results are obtained. The candidate solution is subjected to falsification verification, which includes generating perturbation samples based on the constraint set and performing consistency checks and targeted counterexample searches to obtain counterexample records, and updating the argumentation rule set and constraint set based on the counterexample records and outputting the updated performance prediction results.

[0005] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By constructing a component knowledge graph driven by batch data, the structured accumulation and traceable expression of component relationships and conditional dependencies are realized. By characterizing the conflict of the set of argumentation rules and the attack relationship graph, the inconsistency evidence is minimized and located and transformed into executable constraints, thereby obtaining feasible candidate solutions more stably within the candidate formulation and candidate process window. Through perturbation consistency testing and targeted counterexample search, the falsification verification of the boundary vulnerability of candidate solutions is realized, and the counterexample record drives the update of rules and constraints, realizing self-calibrating extrapolation and iterative output of results for extreme conditions, reducing trial and error costs and mass production failure risks. Attached Figure Description

[0006] Figure 1 This is a schematic diagram of the execution flow of the method of the present invention; Figure 2 This is an overview diagram of batch data and test results in a specific embodiment of the present invention; Figure 3 This is a schematic diagram of a partial structure of the component knowledge graph in a specific embodiment of the present invention; Figure 4 This is a feasible region distribution diagram of candidate solutions under the constraint set in a specific embodiment of the present invention; Figure 5 This is a graph showing the consistency of falsification verification perturbation and the search results for counterexamples in a specific embodiment of the present invention. Detailed Implementation

[0007] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation.

[0008] Ingredient knowledge graphs can be understood as a data organization and representation method for formulation development, used to explicitly represent the "ingredient-ingredient relationship-applicable conditions-evidence source" in a graph structure. Nodes in the graph correspond to ingredients or key processes and storage conditions, while edges describe compatibility, incompatibility, substitution, and other relationships. Triggering conditions and evidence strength are added to these relationships, enabling the application boundaries of the same relationship under different testing protocols, process windows, or packaging and storage conditions to be distinguished and recorded. Unlike simply storing formulation tables or training prediction models, ingredient knowledge graphs emphasize transforming verifiable facts obtained from experience, experiments, and statistical data into queryable and combinable structured knowledge, thereby supporting constrained reasoning and interpretable performance extrapolation within candidate formulations and process windows. Based on this, this invention further combines ingredient knowledge graphs with argumentation rules, conflict localization, constraint solving, and falsification iteration to achieve a closed-loop formulation performance prediction system for R&D scenarios.

[0009] like Figure 1As shown, a formulation performance prediction method based on component knowledge graph is proposed, which includes: Obtain batch data for the formulation, including formulation composition data, process parameter data, packaging and storage condition data, testing protocol data, and test result data; In one embodiment, formula batch data is acquired to form a computable experimental batch record, supporting subsequent component knowledge graph construction, validation rule generation, and performance prediction. Formula batch data uses the formula batch as the smallest management unit, covering the set of key facts for the same batch from formulation design to test results. Formula batch data includes at least formulation composition data, process parameter data, packaging and storage condition data, testing protocol data, and test result data. The formulation composition data characterizes the name and dosage of each component; the process parameter data characterizes the temperature, time, stirring, shearing, and feeding sequence of the preparation process; the packaging and storage condition data characterizes the packaging materials, sealing status, and storage environment; the testing protocol data characterizes sample preparation, testing instruments, and judgment caliber; and the test result data characterizes the performance index values ​​obtained under the testing protocol, thus obtaining a structured batch data record.

[0010] Acquiring batch data includes: generating batch identifiers; associating and storing formula composition data, process parameter data, packaging and storage condition data, testing protocol data, and testing result data with batch identifiers to form batch data; and performing field mapping and unit normalization on the batch data.

[0011] In one embodiment, acquiring batch data includes three steps: generating batch identifiers, associating storage and field mapping, and unit normalization. The added limitation is that the batch identifiers are used to aggregate multi-source data in a closed loop with the same key value, and the fields and units are unified before being stored in the database.

[0012] When generating batch identifiers, a unique string is generated as the batch identifier at the moment the formulation test or pilot production task is created. The batch identifier is obtained by concatenating the project code, date, and random sequence, with a check digit appended at the end to reduce manual input errors. The batch identifier is written into the experimental record system or production execution system and serves as the primary key for all subsequent data writes. During associated storage, formulation composition data, process parameter data, packaging and storage condition data, testing protocol data, and test result data are extracted into structured field sets and associated with the batch identifier to form batch data. The formulation composition data must at least include the standard name of the ingredients, grade information, unit of measurement, and dosage value. The process parameter data must at least include the feeding sequence, stirring speed or linear velocity, temperature setpoint, holding time, and cooling method. The packaging and storage condition data must at least include the type of packaging material, container volume, sealing method, light protection conditions, temperature, and humidity. The testing protocol data must at least include the sampling time point, sample preparation method, testing method number, and judgment threshold. The test result data must at least include the performance indicator name, value, number of repetitions, and statistical caliber.

[0013] During field mapping and unit normalization, a field mapping dictionary is maintained, containing the original field name, unified field name, data type, allowed unit set, and default value rules. Based on the field mapping dictionary, field names from different sources are mapped to unified field names, and synonymous component names are mapped to standard component names. Usage is standardized to mass fraction or mass percentage, with mass unit conversion based on weighing records and volume unit conversion based on component density information. Temperature is standardized to degrees Celsius, time to minutes, and rotation speed to revolutions per minute. For missing fields, they are filled or marked as missing according to default value rules. For values ​​exceeding the allowed range, anomaly markers are generated and written back to the batch data record; these markers are used to indicate removal or weight reduction in subsequent modeling stages. Simultaneously, the field mapping dictionary version number and normalization processing log are recorded to trace the processing source of the batch data.

[0014] During the associated storage process, batch data records are simultaneously written with the data source identifier, collection timestamp, and operator identifier, and a summary value is calculated for the batch data records for integrity verification; if an inconsistency in the summary value is found, the data entry is stopped and a list of fields that need to be collected is returned.

[0015] A component knowledge graph is constructed based on batch data, and a set of argumentation rules is generated, which includes triggering conditions, threshold judgments, and rebuttal conditions. In one embodiment, a component knowledge graph is constructed based on batch data, and a set of argumentation rules is generated to solidify the reusable relationships between "component combination—process—packaging and storage—testing protocol—testing results" into a computable structure. The component knowledge graph uses components as the core object, and describes the co-occurrence and substitution boundaries between components under specific conditions using compatibility, incompatibility, and substitution relationships. The set of argumentation rules consists of triggering conditions, threshold judgments, and rebuttal conditions. Triggering conditions limit the value combinations of formulation composition data, process parameter data, packaging and storage condition data, and testing protocol data. Threshold judgments indicate whether performance indicators meet the performance thresholds given by the testing protocol data. Rebuttal conditions characterize the combination of conditions where the threshold judgment changes from satisfied to unsatisfied after the triggering conditions change, thus providing an executable basis for the subsequent generation of constraint sets.

[0016] Constructing a component knowledge graph includes: using components as nodes in the component knowledge graph; using compatibility, incompatibility, and substitution relationships between components as edges in the component knowledge graph; determining compatibility, incompatibility, and substitution relationships based on batch data. Specifically, compatibility relationships indicate component combinations whose test results meet performance thresholds under the same testing protocol data, process parameter data, and packaging and storage conditions; incompatibility relationships indicate component combinations whose test results do not meet performance thresholds under the same testing protocol data, process parameter data, and packaging and storage conditions; and substitution relationships indicate that, provided the performance threshold is met, the difference in test results after replacing the first component with the second component does not exceed the tolerance defined by the testing protocol data. The performance threshold is given by the testing protocol data.

[0017] In one embodiment, the additional limitation of constructing the component knowledge graph is to use components as nodes, compatibility relationships, incompatibility relationships, and substitution relationships as edges, and to determine the value rules of the three types of edges based on batch data under the same testing protocol data, the same process parameter data, and the same packaging and storage conditions data, thereby solidifying the "reusable experience under consistent conditions" into a structured graph.

[0018] Specifically, the formulation composition data is first standardized. This standardization includes mapping ingredient names to unified standard names, mapping synonyms and product names to the same standard name, and retaining ingredient grade information and supplier information as node attributes to distinguish differences caused by different grades of the same name. Then, each ingredient is established as a node in an ingredient knowledge graph, and node attributes are written for each node. Node attributes include at least the ingredient category, allowable dosage range, and typical dosage interval. The allowable dosage range is obtained statistically from the ingredient dosage distribution in batch data, and the typical dosage interval is obtained statistically from a subset of batches that meet the performance threshold.

[0019] Then, using batch data as the sample source, the data is grouped according to testing protocol data, process parameter data, and packaging and storage condition data to obtain multiple sample groups with consistent conditions. Within each sample group, based on the performance threshold given by the testing protocol data, a pass / fail judgment is performed on the test results. The pass judgment indicates that the test results meet the performance threshold, and the fail judgment indicates that the test results do not meet the performance threshold. Based on the pass and fail judgments within the sample group, compatibility and incompatibility relationships are constructed: compatibility relationships indicate combinations of components that, within the same sample group with consistent conditions, occur repeatedly while meeting the minimum sample number requirement, and the corresponding test results all meet the performance threshold; incompatibility relationships indicate combinations of components that, within the same sample group with consistent conditions, occur repeatedly while meeting the minimum sample number requirement, and the corresponding test results all do not meet the performance threshold.

[0020] Component combinations can be generated at a two-level granularity: the first level is for component pairs, and the second level is for multi-component combinations. At the component pair level, compatible edges connect two component nodes, and incompatible edges connect two component nodes. At the multi-component combination level, multi-component combinations are treated as combination nodes and combined containment edges are established with each component node. Simultaneously, compatible or incompatible edges are established between combination nodes to preserve higher-order cooperative or conflicting information. To ensure feasibility and stability, the minimum number of samples can be set to no less than a preset number of iterations. The identifiers of sample groups whose number of iterations matches the conditions are recorded, and the number of iterations is written as an edge attribute into the component knowledge graph for subsequent rule selection. The determination of substitution relationships follows a "consistent conditions, satisfactory performance, and controlled differences" logic: Within a sample group with consistent conditions, sample pairs that change only at the position of one component are selected from the qualified samples. These sample pairs must satisfy the following conditions: the differences in the dosage of other components besides the target component are within a preset range, and the process parameter data, packaging and storage condition data, and testing protocol data are consistent. When the first component is replaced by the second component, if the difference in the test results does not exceed the tolerance specified in the testing protocol data, a substitution relationship is established between the first and second components. Differences in the test results can be compared item by item according to the performance indicators specified in the testing protocol data; if the difference in any performance indicator exceeds the tolerance, the substitution relationship is deemed invalid.

[0021] To avoid misjudgments caused by accidental substitution, the substitution relationship edges also record the number of times they are supported, and are only written into the component knowledge graph after the number of supports reaches a preset threshold. In this way, the component knowledge graph can provide an objective basis for judging compatibility, incompatibility, and substitution under clear conditional consistency constraints, and retain the source of evidence in the form of node attributes and edge attributes, which can be directly referenced when generating subsequent argumentation rules.

[0022] The generated set of argumentation rules includes: generating triggering conditions, threshold determination, and rebuttal conditions for each argumentation rule based on the component knowledge graph and batch data; the triggering conditions are jointly defined by the formulation composition data, process parameter data, packaging and storage condition data, and testing protocol data; the threshold determination is used to indicate whether the performance index in the performance prediction result meets or does not meet the performance threshold; the rebuttal conditions are used to indicate the combination of conditions that change the threshold determination from meeting the performance threshold to not meeting the performance threshold when at least one of the triggering conditions changes.

[0023] In one embodiment, the added limitation in generating the set of argumentation rules lies in generating triggering conditions, threshold judgments, and rebuttal conditions for each argumentation rule. The triggering conditions are jointly defined by formula composition data, process parameter data, packaging and storage condition data, and testing protocol data. The rebuttal conditions characterize the combination of conditions that cause a change in the triggering conditions to a change in the threshold judgment from being satisfied to not being satisfied, thus enabling the rules to be used for both inference and falsification. Specifically, firstly, performance thresholds and a set of performance indicators are determined based on the testing protocol data. The set of performance indicators includes the names and judgment methods of key indicators specified in the testing protocol data, and the test results data in the batch data are labeled as passing samples and failing samples according to the performance thresholds.

[0024] Subsequently, candidate trigger conditions are constructed, consisting of four categories of elements: First, formulation composition data elements, including whether an ingredient appears, whether its dosage falls within a certain range, and whether ingredient pairs appear simultaneously; second, process parameter data elements, including whether the temperature, time, stirring / shear, and feeding sequence meet a preset pattern; third, packaging and storage condition data elements, including packaging material type, sealing method, light protection conditions, storage temperature range, and storage time range; and fourth, testing protocol data elements, including the testing method number, sample preparation method, and judgment threshold version. Trigger conditions are formed by combining several candidate trigger conditions. The combination method prioritizes "combining a small number of elements" to ensure interpretability and feasibility, and a maximum upper limit is set for the number of elements to avoid overly detailed conditions. Next, the candidate triggering conditions are screened based on the component knowledge graph. The screening rules include: when the triggering condition contains component pairs, the component pairs are selected from the compatible and incompatible relationships first; when the triggering condition contains component substitution, the substitution pairs are selected from the substitution relationship first, so that the triggering conditions are consistent with the graph evidence.

[0025] Then, support is calculated for each trigger condition, where support is the number of batch samples that meet the trigger condition. A pass rate is also calculated, which is the percentage of samples that meet the trigger condition and pass the judgment. When both support and pass rate reach a preset threshold, a validation rule is generated, and a threshold judgment is given. The threshold judgment indicates whether the performance indicator meets or does not meet the performance threshold when the trigger condition is met. Threshold judgments can be generated separately for each performance indicator or based on a comprehensive judgment specified by the testing protocol data. The comprehensive judgment is based on the judgment criteria of the testing protocol data. The generation of rebuttal conditions follows the "minimum change from pass to fail" approach: for each generated trigger condition, samples in the batch data that are closest to the trigger condition but failed the judgment are retrieved. The closest judgment criterion is the minimum number of differences between the formulation composition data, process parameter data, packaging and storage condition data, and testing protocol data. Combinations of differences that cause the threshold judgment to flip are extracted from these samples, and these combinations are defined as rebuttal conditions.

[0026] The discrepancies can be one or more of the following: ingredient dosage exceeding the dosage range, key process parameters exceeding the range, changes in packaging materials, changes in storage temperature or duration, and changes in the testing protocol version. To avoid overly broad rebuttal conditions, the rebuttal conditions must also meet the minimum sample size requirement, and the number of times a rebuttal is supported must be recorded. The generated set of argumentation rules is stored in a structured manner. Each argumentation rule must include at least a rule identifier, trigger condition, threshold determination, rebuttal condition, support level, and pass rate, and a corresponding batch identifier list must be recorded to support traceability. Through the above process, the set of argumentation rules can form a closed loop under the "condition-conclusion-counter-evidence" ternary structure, enabling subsequent steps to directly use trigger conditions to generate constraints, use threshold determination to generate predicted conclusions, and use rebuttal conditions to trigger falsification and updates.

[0027] An attack relationship graph is constructed based on the set of argumentation rules, and a minimum inconsistency set is determined. A constraint set is generated based on the minimum inconsistency set. Candidate solutions are obtained by solving within the candidate formula and candidate process window based on the constraint set, and performance prediction results are obtained. In one embodiment, an attack relationship graph is constructed based on the set of argumentation rules, and a minimum inconsistency set is determined to identify conflicting rule combinations under the same triggering condition, thereby generating an executable constraint set. The constraint set limits the feasible range of candidate formulations and candidate process windows, enabling the solution to obtain candidate solutions and performance prediction results within the feasible range. When constructing the attack relationship graph, conflict determination is performed on rule pairs with consistent triggering conditions to form attack relationships between rules. When determining the minimum inconsistency set, the smallest subset of rules that prevents rules from simultaneously being valid is extracted from the attack relationship graph. When generating the constraint set, the rebuttal conditions corresponding to the minimum inconsistency set are transformed into conditional triggering constraints, and these constraints are applied to the component dosage values ​​and process parameter ranges. In the solution phase, candidate solutions are generated based on the component dosage value range of the candidate formulation and the process parameter range of the candidate process window. After constraint set verification, candidate solutions are determined, and the performance prediction results corresponding to the candidate solutions are output.

[0028] Constructing the attack relationship graph includes: when the threshold determination of the first argument rule and the threshold determination of the second argument rule are opposite to the same performance index in the performance prediction result, under the same triggering conditions, establishing the attack relationship between the first argument rule and the second argument rule in the attack relationship graph. The attack relationship is used to indicate that the first argument rule and the second argument rule cannot be true at the same time.

[0029] In one embodiment, the additional constraint in constructing the attack relationship graph is that the triggering conditions are consistent as a comparison premise, and the threshold judgments of the same performance index are opposite as conflict judgment criteria, thereby forming a verifiable attack relationship.

[0030] Specifically, the set of validation rules is first binned based on the equivalence of trigger conditions. Trigger condition equivalence indicates that two validation rules have consistent values ​​in terms of formulation composition data, process parameter data, packaging and storage condition data, and testing protocol data. Trigger condition equivalence can be determined by the consistency of the field set and the consistency of the field value range. Field set consistency indicates that the set of specified field names is consistent, and field value range consistency indicates that the value ranges of each field are the same or mutually inclusive, and the inclusion relationship satisfies the preset consistency rules. After binning, rule pairs are compared pairwise within the same trigger condition bin, with the comparison objects being threshold determination and performance indicators.

[0031] Threshold determination is based on the performance threshold given in the detection protocol data, and includes at least two categories: those that meet the performance threshold and those that do not. Performance metrics are based on the metric names specified in the detection protocol data; conflict determination is performed only when the performance metric names match during rule pair comparison. Conflict determination adopts an "opposite" criterion: if the threshold determination of the first argument rule meets the performance threshold and the threshold determination of the second argument rule does not meet the performance threshold, or if the threshold determination of the first argument rule does not meet the performance threshold and the threshold determination of the second argument rule meets the performance threshold, then the first and second argument rules are determined to conflict on that performance metric. For rule pairs determined to be conflicting, attack relationship edges are established in the attack relationship graph, and the attributes of the attack relationship edges are recorded as the conflicting performance metric name, the trigger condition bucket identifier, and the conflict determination timestamp. Attack relationship edges can be represented by undirected edges or bidirectional directed edges; when using directed edges, a bidirectional approach is adopted to maintain the symmetric semantics of "cannot be true simultaneously".

[0032] To avoid misjudgments due to low evidence strength, a minimum support constraint can be introduced when constructing the attack relationship graph. Minimum support indicates that the batch sample size corresponding to the triggering condition is not less than a preset number. When the support of either the first or second argument rule is lower than the minimum support, no attack relationship edge is established, and it is recorded as a rule pair to be confirmed. Through this process, the attack relationship graph uses consistent triggering conditions as boundaries and opposite threshold judgments as conflict criteria, ensuring that attack relationships have clear judgment basis and directly indicating that the first and second argument rules cannot be simultaneously true, thus providing structured input for the subsequent extraction of the minimum inconsistency set.

[0033] Determining the minimum inconsistency set includes: finding the minimum subset of argument rules in the attack relationship graph that prevents the set of argument rules from being simultaneously true as the minimum inconsistency set; generating the constraint set includes: converting the rebuttal conditions corresponding to the minimum inconsistency set into conditional trigger constraints; conditional trigger constraints are constraints that take effect when the trigger condition is met, and are used to tighten the allowable range of component dosage values ​​or tighten the range of process parameters, or to prohibit the combination of component dosage values ​​and process parameter ranges.

[0034] In one embodiment, the new limiting point for determining the minimum inconsistency set and the generating constraint set is to extract the minimum subset of rules from the attack relationship graph that makes it impossible for the rule set to be true at the same time, and to transform the rebuttal condition corresponding to the minimum subset of rules into a condition triggering constraint that takes effect when the triggering condition is met, thereby reducing the "conflict interpretation" to "executable constraint".

[0035] Specifically, when determining the minimum inconsistency set, firstly, conflict subgraphs are identified in the attack relationship graph. Conflict subgraphs are formed by connecting attack relationship edges within the same trigger condition bucket. For each conflict subgraph, a candidate conflict rule set is extracted, consisting of rule nodes from the conflict subgraph. Then, a minimization process is performed on the candidate conflict rule set. This minimization process removes rule nodes unrelated to the conflict from the candidate conflict rule set, resulting in the minimum rule subset where no two rules can be simultaneously true. The minimization process can employ a step-by-step elimination verification method: for each rule node in the candidate conflict rule set, the rule node is temporarily removed, and it is checked whether the remaining rule nodes still have conflicts with opposite threshold judgments. If the conflict disappears after removing a rule node, then that rule node belongs to the core of the conflict and is retained; if the conflict still exists after removing a rule node, then that rule node does not belong to the core of the conflict and can be eliminated.

[0036] The process is repeated until no further inconsistencies can be eliminated, resulting in a minimal inconsistency set. This minimal inconsistency set is stored by recording the trigger condition bucket identifier, the list of included rule identifiers, and the name of the conflict performance index, ensuring consistency of trigger conditions during subsequent constraint generation. When generating the constraint set, the rebuttal conditions of each argument rule in the minimal inconsistency set are transformed into conditional trigger constraints. Conditional trigger constraints consist of two parts: an activation condition, which is the trigger condition and indicates when the constraint is activated if the formulation composition data, process parameter data, packaging and storage condition data, and testing protocol data meet the corresponding trigger condition; and a constraint content, derived from the rebuttal conditions, used to limit the candidate formulation and candidate process windows.

[0037] Changes in formulation composition data in the rebuttal conditions are transformed into a tightening of the allowable range of component dosage values. Tightening methods include lowering the upper limit of component dosage, raising the lower limit of component dosage, or limiting the component dosage value to a preset range. Changes in process parameter data in the rebuttal conditions are transformed into a tightening of the process parameter range. Tightening methods include lowering the upper limit of process parameters, raising the lower limit of process parameters, or limiting the process parameters to a preset range. Changes in component combination in the rebuttal conditions are transformed into combination prohibition constraints. Combination prohibition constraints are used to prohibit the combination of a certain component dosage value with a certain process parameter range.

[0038] To ensure feasibility, conditional triggering constraints are expressed using structured fields, including at least a set of triggering fields, a range of triggering field values, the name of the restricted field, the allowed range of the restricted field, and a constraint type identifier, as well as recording the source rule identifier. The constraint set consists of multiple conditional triggering constraints, and the constraint set records its version number and generation time during storage for alignment with the update loop in the subsequent falsification and verification phase. Through this process, the minimum inconsistency set provides the smallest explanatory unit for the conflict, and the constraint set maps the explanatory unit to an executable boundary, thus providing a clear feasible domain constraint for the solution phase.

[0039] Solving within the candidate formulation and candidate process window based on the constraint set includes: generating candidate solutions within the range of component dosage values ​​defined by the candidate formulation and the range of process parameters defined by the candidate process window; performing constraint set verification on the candidate solutions, which includes determining whether the candidate solutions satisfy each constraint in the constraint set; and determining the candidate solutions that satisfy the constraint set as candidate schemes.

[0040] In one embodiment, the newly added constraint points obtained by solving within the candidate recipe and candidate process window based on the constraint set are first used to generate candidate solutions, then the constraint set is verified on the candidate solutions, and finally the candidate solutions that satisfy the constraint set are determined as candidate schemes, thereby decoupling the feasibility test from the scheme output, which is convenient for implementation and debugging.

[0041] Specifically, the boundaries of the candidate formulation and candidate process windows are first determined. The candidate formulation defines the set of components and the range of values ​​for each component's dosage, while the candidate process window defines the set of process parameter names and the range of values ​​for each process parameter. The candidate formulation and candidate process windows can be input by R&D personnel or obtained from historical batch data statistics, and in this embodiment, they are based on the scope that can cover the current R&D needs. Then, candidate solutions are generated within the range of component dosage values ​​and the range of process parameters. Each candidate solution includes at least one set of specific component dosage values ​​and one set of specific process parameter values.

[0042] Candidate solution generation can employ a hierarchical sampling approach: first, generate several formulation sub-candidates within the component dosage range of the candidate formulation at a preset step size; then, generate several process sub-candidates within the process parameter range of the candidate process window at a preset step size; finally, combine the formulation sub-candidates and process sub-candidates to form a candidate solution set. To reduce unnecessary combinatorial explosion, a component knowledge graph-based filtering approach can be introduced during the candidate solution generation stage: when a component pair has an incompatible relationship in the component knowledge graph and the number of times it is supported reaches a preset threshold, formulation sub-candidates containing that component pair are directly eliminated; when a substitution relationship exists and the number of times it is supported reaches a preset threshold, alternative versions are allowed to be generated in the same candidate solution set for subsequent comparison. After candidate solution generation is complete, constraint set verification is performed on each candidate solution.

[0043] Constraint set verification involves traversing each condition-triggered constraint in the constraint set and performing validity and constraint checks on candidate solutions: validity checks determine whether a candidate solution satisfies the validity condition of the condition-triggered constraint; when a validity check is true, constraint checks are performed to determine whether a candidate solution satisfies the constraint content of the condition-triggered constraint; when any constraint check is false, the candidate solution is marked as infeasible and a violation constraint identifier is recorded. After verifying all candidate solutions, the candidate solutions not marked as infeasible are determined as the candidate solution set. The candidate solution set can be further sorted according to the target performance index, and priority solutions are selected. The sorting is based on the margin between the performance index value and the performance threshold in the performance prediction results. The margin indicates the difference between the performance index value and the performance threshold.

[0044] The generation of performance prediction results can be performed after the candidate solutions have passed the constraint set verification, thus avoiding invalid predictions for infeasible candidate solutions. Performance prediction results are stored in a one-to-one correspondence with candidate solutions, and the corresponding trigger condition bucket identifier and constraint set version number are recorded to ensure that the constraint set on which the candidate solution was based can be traced during the subsequent falsification verification stage. Through the above process, the solution stage focuses on candidate solution generation, constraint set verification, and candidate solution output, ensuring consistency between the candidate solution and the constraint set, and establishing a clear correspondence with the performance prediction results.

[0045] The candidate solution is subjected to falsification verification, which includes generating perturbation samples based on the constraint set and performing consistency checks and targeted counterexample searches to obtain counterexample records, and updating the argumentation rule set and constraint set based on the counterexample records and outputting the updated performance prediction results.

[0046] In one embodiment, falsification verification is performed on candidate solutions to test whether they have stable threshold determination conclusions within the feasible region defined by the constraint set, and to form a traceable update loop when counterexamples are found. Falsification verification includes two parallel paths: one path generates perturbation samples based on the constraint set and performs consistency checks to verify the stability of performance prediction results under small feasible perturbations; the other path is a targeted counterexample search, which actively seeks counterexample records near the boundary of the constraint set that cause the threshold determination to flip. After falsification verification is completed, the argumentation rule set and constraint set are updated based on the counterexample records, and the updated performance prediction results are recalculated and output under the updated constraint set, ensuring that the prediction conclusions of the candidate solutions are consistent with the constraint boundaries and enabling the tracing of sources of inconsistency.

[0047] Generating perturbation samples includes: under the premise of satisfying the constraint set, increasing or decreasing the perturbation of the component dosage values ​​of the candidate scheme by a preset step size, and adjusting the dosage value of at least one unperturbed component to keep the total component dosage value unchanged, thus obtaining a formulation perturbation sample; under the premise of satisfying the constraint set, perturbing the process parameter range of the candidate scheme at the upper or lower boundary by a preset step size, thus obtaining a process perturbation sample; consistency verification includes: comparing the performance prediction results corresponding to the candidate scheme and the perturbation sample, and determining that the threshold judgment does not occur under the perturbation sample, and the difference in performance prediction results does not exceed the tolerance limit of the detection protocol data.

[0048] In one embodiment, the new constraint for generating perturbation samples and consistency testing is to perform controlled perturbation on the component dosage values ​​and process parameter ranges under the premise of satisfying the constraint set, and to use the threshold determination of no flipping and the prediction difference not exceeding the tolerance as the consistency judgment criteria, thereby measuring the local stability of the candidate scheme in a verifiable manner.

[0049] Specifically, firstly, the conditional triggering constraints applicable to the candidate schemes are parsed from the constraint set. This parsing yields the currently effective allowable ranges for component dosage values ​​and process parameters, which are then used as the boundaries of the perturbation. Subsequently, a formulation perturbation sample is generated. When generating the formulation perturbation sample, a target component from the candidate scheme is selected as the perturbed component. The target component can be selected based on the relationship strength in the component knowledge graph, or it can be selected based on the dosage proportion in the formulation composition data from high to low. The dosage value of the perturbed component is then increased or decreased according to a preset step size. This preset step size can be configured by the R&D personnel or obtained from the dosage resolution statistics of historical batch data.

[0050] To ensure consistency in the total formulation, after adding or subtracting perturbations to the disturbed components, at least one undisturbed component is selected as a compensation component. The dosage of this compensation component is then compensated in reverse to maintain the total dosage. When multiple compensation components exist, the compensation amount can be allocated according to preset weights, which can be determined based on the component's functional category or historical dosage percentage. For each formulation perturbation sample, a constraint set verification is performed; only samples that satisfy the constraint set are retained as valid formulation perturbation samples. Then, process perturbation samples are generated.

[0051] When generating process disturbance samples, a target process parameter from a candidate scheme is selected as the disturbed parameter. Under the premise of satisfying the constraint set, the disturbance is performed along the upper or lower boundary of the allowable range of the process parameter at a preset step size to obtain the process disturbance sample. When the process parameter is an interval parameter, the disturbance can change only the upper boundary of the interval, only the lower boundary, or both the upper and lower boundaries while keeping the interval width unchanged, depending on the physical meaning of the process parameter. For each process disturbance sample, constraint set verification is also performed; only samples that satisfy the constraint set are retained. After the disturbance samples are generated, performance prediction results are calculated for the candidate scheme and each disturbance sample, and performance index values ​​consistent with the detection protocol data are extracted from the performance prediction results. A consistency check is then performed.

[0052] The consistency test comprises two judgments: the first is a threshold judgment stability judgment, which determines whether the threshold judgment corresponding to the candidate solution does not flip under perturbation samples. Flipping includes changing from meeting the performance threshold to not meeting it, and vice versa. The second is a numerical difference judgment, which compares the performance index values ​​of the candidate solution and the perturbation samples and determines whether the difference does not exceed the tolerance limit set by the detection protocol data. The tolerance limit set by the detection protocol data can be an absolute tolerance or a relative tolerance. The absolute tolerance limits the upper limit of the allowable difference in performance index values, while the relative tolerance limits the upper limit of the proportion of the performance index value difference to the benchmark value. If any perturbation sample triggers a threshold judgment flip or the numerical difference exceeds the tolerance, the perturbation sample is marked as an inconsistent sample, and the triggered constraint entry and the name of the triggered performance index are recorded. If all perturbation samples do not trigger a flip and the differences do not exceed the tolerance, the candidate solution is marked as having passed the consistency test. Through the above process, the consistency test verifies the prediction stability within the feasible region using controlled perturbation, while also providing clues for subsequent counterexample search and rule updates.

[0053] The targeted counterexample search includes: taking the boundary point between the component dosage value and the process parameter range of the candidate solution as the search starting point, iteratively adjusting the component dosage value or process parameter value so that the adjusted component dosage value or process parameter value satisfies the rebuttal condition, and generating a counterexample record; the counterexample record includes the minimum trigger change, which is used to indicate the minimum variable change that causes the threshold judgment to change from satisfied to unsatisfied.

[0054] In one embodiment, the additional limiting point of the targeted counterexample search is to take the boundary point between the component dosage value of the candidate solution and the process parameter range as the search starting point, and to take the satisfaction of the rebuttal condition as the search target, so as to find the minimum trigger change that causes the threshold determination to be reversed in a counterexample-oriented manner.

[0055] Specifically, the search starting point set is first determined. This set consists of boundary points of the component dosage values ​​corresponding to the candidate solutions and boundary points of the process parameter ranges. The boundary points include the upper and lower bounds of the component dosage values ​​and the upper and lower bounds of the process parameter ranges. The values ​​of the boundary points are derived from the allowable ranges defined by the effective triggering constraints in the constraint set. Then, target rebuttal conditions are selected for each search starting point. These target rebuttal conditions are derived from the set of argumentation rules that are consistent with the triggering conditions of the candidate solutions, with priority given to rebuttal conditions that conflict with the threshold judgments corresponding to the candidate solutions. The targeted counterexample search is performed using an iterative adjustment method.

[0056] During iterative adjustments, variables directly related to the target counter-conditions are prioritized as adjustment variables. These include component dosage values ​​and process parameter values. When the target counter-condition indicates that a component dosage value exceeds the dosage range, the adjustment variable becomes the dosage value of the corresponding component. When the target counter-condition indicates that a process parameter value exceeds the range, the adjustment variable becomes the corresponding process parameter value. In each iteration, the adjustment variables are moved from the boundary point along a preset step size, forming new candidate counterexamples. Constraint set verification is performed on the new candidate counterexamples. If the constraint set is not satisfied, the step is rolled back and retried in the opposite direction or with a smaller step size. If the constraint set is satisfied, the performance prediction result is calculated, and it is determined whether the target counter-conditions are met. The criterion for satisfying the target counter-conditions is whether all the differences listed in the counter-conditions are true, such as whether the component dosage value enters the counter-condition range, whether the process parameter value enters the counter-condition range, or whether the component combination triggers the conflict mode corresponding to the combination prohibition constraint. When a new candidate counterexample satisfies the target counter-conditions, it is recorded as a counterexample record, and minimum trigger change extraction is performed on the counterexample record.

[0057] The minimum trigger change extraction employs a step-back reduction approach: starting from the found counterexample records, the change amount of the adjustment variable is reduced one by one according to a preset priority, and after each reduction, it is re-evaluated whether the rebuttal condition is still met; when further reduction would cause the rebuttal condition to no longer be met, the current change amount is determined as the minimum trigger change. The minimum trigger change is recorded in the form of variable name and change amount, where the change amount can be the difference between component dosage values ​​or process parameter values. The counterexample record includes at least the trigger condition identifier, target rebuttal condition identifier, adjustment variable name, minimum trigger change, component dosage values ​​and process parameter values ​​of the candidate counterexample, as well as the corresponding performance prediction results and threshold judgment reversal direction. Through the above process, the targeted counterexample search actively seeks counterevidence near the allowed boundaries of the constraint set and provides an executable boundary correction basis in the form of minimum trigger changes.

[0058] Updating the argument rule set and constraint set based on counterexample records includes: adding the rebuttal conditions corresponding to the counterexample records to the argument rule set, or replacing the rebuttal conditions corresponding to the same triggering conditions in the argument rule set; redetermining the minimum inconsistency set based on the updated argument rule set and generating the updated constraint set; and outputting the updated performance prediction results based on the updated constraint set.

[0059] In one embodiment, the new limitation of updating the argument rule set and constraint set based on the counterexample record is to write or replace the rebuttal condition corresponding to the counterexample record into the argument rule set, and regenerate the constraint set based on the updated argument rule set, so that the update action is traceable and consistent with the conflict structure.

[0060] Specifically, firstly, counterexample records are merged based on the consistency of triggering conditions and rebuttal conditions. For each merged group of counterexample records, the frequency of occurrence and minimum trigger change distribution are statistically analyzed to determine the stability of the counterexamples. Then, the argument rule set is updated. This update includes two types of operations: adding and replacing. When a counterexample record's triggering condition does not have a corresponding rule in the argument rule set, the corresponding rebuttal condition is written as the rebuttal condition for the new rule, and threshold and support fields are added to the new rule. When a counterexample record's triggering condition already has a corresponding rule in the argument rule set, the corresponding rebuttal condition replaces the rebuttal condition for the same triggering condition in that rule, retaining the original rule identifier and recording version increments during replacement.

[0061] When writing rebuttal conditions, the minimum triggering change is used as the boundary value source for the rebuttal conditions. For example, the boundary of the rebuttal interval for component dosage values ​​is set as the candidate solution value plus the minimum triggering change, or the boundary of the rebuttal interval for process parameter values ​​is set as the candidate solution value plus the minimum triggering change, to ensure that the update can directly tighten the boundaries. After completing the update of the argumentation rule set, the attack relationship graph construction, minimum inconsistency set determination, and constraint set generation are re-executed to ensure that the constraint set is consistent with the updated conflict structure. When generating the constraint set, the updated rebuttal conditions are preferentially converted into conditional triggering constraints, and tightening updates are performed on the conditional triggering constraints corresponding to the candidate solution triggering conditions; tightening updates include tightening the allowable range of component dosage values, tightening the range of process parameters, or adding prohibited combinations of constraints.

[0062] After the updated constraint set is formed, the performance prediction results of the candidate solutions are recalculated based on the updated constraint set, resulting in updated performance prediction results. These results are then compared and recorded with the performance prediction results before the update. The comparison record includes at least whether the threshold judgment has been reversed, changes in the main performance index values, and the identifier of the counterexample record that triggered the update. Finally, the updated argument rule set version number, the updated constraint set version number, and the updated performance prediction results are associated and stored, and traceability information is written into them. The traceability information includes the counterexample record identifier, the update timestamp, and the operator identifier. Through the above process, counterexample-driven updates can be implemented in a closed loop of "counterexample—rule—conflict—constraint—prediction," ensuring that the prediction conclusion converges as the evidence is updated, and each tightening can be traced back to the specific counterexample source.

[0063] In one specific embodiment, taking the application in the development of emulsion-type products as an example, a complete closed loop is presented, including constructing a knowledge graph from small-scale batch data, generating candidate formulations, and screening out highly robust solutions through "falsification-based verification." The implementation focuses on an aqueous-phase-dominant emulsion system, with core evaluation indicators including: viscosity at 25°C, pH, stratification height at 40°C after 28 days in the dark, and the logarithmic reduction in antibacterial activity. To facilitate reproducibility, this embodiment expresses the formulation as a mass percentage to ensure mass conservation. ,in For the first The mass fraction of each component in the target formulation The corresponding dosage of this component is specified in the "Component Dosage Value".

[0064] Seventy-two small-scale batches were conducted under laboratory conditions. The raw material ratios (glycerol, propylene glycol, nicotinamide, sodium hyaluronate, emulsifier A, thickener B, preservative system C, and water) and process parameters (emulsification temperature, homogenization speed, homogenization time, and cooling rate) for each batch were recorded. Two testing protocols were implemented: Protocol V1 (with a wider threshold) and Protocol V2 (with a threshold closer to the mass production quality threshold). Subsequent spectral analysis and candidate screening in this embodiment primarily used Protocol V2. The statistical results are as follows: Protocol V1 comprised 42 batches, with 33 passing; Protocol V2 comprised 30 batches, with 3 passing. The main failure modes of the failed batches were concentrated in "viscosity exceeding limits" and "stratification exceeding limits," while pH and antibacterial indicators met the thresholds within the range of these batches, consistent with the engineering common sense that "pH / antibacterial activity is more stable within the formulation window, and rheology and stability are more sensitive."

[0065] like Figure 2As shown, the horizontal axis represents viscosity, and the vertical axis represents stratification height. The dashed lines mark the upper and lower limits of viscosity and the upper limit of stratification for Protocol V2, respectively. It can be observed that in the region where viscosity is near the upper limit, increasing the thickener by a small amount can lead to exceeding the limit; in the region where stratification height is near the upper limit, increasing the process cooling rate or decreasing the emulsifier significantly increases the risk of stratification. This phenomenon provides data support for subsequent "relationship extraction" and "counterexample search".

[0066] When constructing the knowledge graph, "component / process factors" are used as nodes, and "substitutable, compatible, incompatible" are used as the relationship types for edges, with conditions assigned to the edges (such as threshold ranges and linkage constraints). Relationship learning uses two types of signals: Co-occurrence-failure attribution: In protocol V2 batches, the pass rate of component pairs co-occurring in the "high-high interval" is statistically analyzed, and attribution is made in conjunction with the failure item (viscosity or stratification); Falsification-based counterexamples: The candidate rules are subjected to "minimum perturbation triggering" and "directed counterexample search". If there is a very small change that can flip the decision from pass to fail, then the condition corresponding to the change is solidified as "edge condition" or "constraint condition".

[0067] Figure 3 A partial structural diagram is provided: for example, "glycerol-propylene glycol" indicates a "substitutable" relationship, which does not mean simple equal substitution, but rather "substitutable while keeping the total amount of polyol constant," that is... When kept within the empirical safety window, the water activity and rheological contribution of the system can be approximately equivalent. For example, "emulsifier A - thickener B" is an "incompatible" relationship, indicating that when both are at high levels, it is easier to trigger viscosity overflow (not chemical incompatibility, but an incompatible window at the index level). "Propylene glycol - thickener B" is an "incompatible" relationship to express that after high shear homogenization and rapid cooling, propylene glycol enhances the solubility of the aqueous phase, which changes the thickening network recovery behavior and amplifies the stratification sensitivity, which needs to be mitigated through process / formulation constraints.

[0068] Candidate search is performed under the constraints of the relationships and conditions output by the knowledge graph. Example of candidate variable range: Emulsifier A [0.95, 1.55]%, Thickener B [0.25, 0.45]%, Glycerin [5.0, 6.2]%, Propylene Glycol [0, 2.5]%, Nicotinamide [3.0, 4.2]%, Sodium Hyaluronate [0.15, 0.45]%, Preservative System C [0.75, 1.05]%, Water is a supplementary item and requires ≥85%; Example of process parameters: Emulsification temperature [72, 78]℃, Homogenization speed [3800, 5200] r / min, Homogenization time [3.0, 5.5] min, Cooling rate [1.0, 1.8]℃ / min.

[0069] The screening target adopts the "robust margin maximization". For each candidate, four indicators are predicted and their respective margins are calculated: viscosity margin is the minimum distance from the upper and lower limits; pH margin is the minimum distance from the upper and lower limits; stratification margin is "the upper limit of stratification minus the predicted stratification height"; and antibacterial margin is "the logarithmic decrease value minus the threshold".

[0070] Comprehensive margin This ensures that even the weakest point has a safety margin. This is a minimum value operation used to select the smallest value from the candidate values ​​within the parentheses; The comprehensive margin is used to characterize the weakest safety margin of a candidate scheme under multiple performance indicators; Viscosity margin represents the minimum safety margin of the predicted viscosity of a candidate solution relative to the viscosity threshold range given by the detection protocol. The unit of viscosity is mPa. s; pH margin represents the minimum safety margin of the predicted pH of the candidate scheme relative to the pH threshold range given by the detection protocol. The layer margin represents the difference between the upper limit of the layer height given by the detection protocol and the predicted layer height of the candidate scheme. The unit of layer height is mm. The inhibition margin represents the safety margin of the predicted log reduction in inhibition value of the candidate protocol relative to the given inhibition threshold of the detection protocol. The unit of the log reduction in inhibition value is the log reduction value. This is a subscript identifier for viscosity indices; This is a subscript identifier for the layer height indicator; This is a subscript identifier for the log reduction value index of antibacterial activity.

[0071] Among 8,000 randomly generated candidate points, the prediction pass rate is about 30.6% without graph constraints. After adding graph constraints, the candidate space shrinks to about 36.7% of the point set, but the prediction pass rate increases to about 39.4%, which reflects the engineering benefits of "first compressing the search space with knowledge graph and then performing robust optimization in the feasible region". Figure 4 The feasible region distribution with “thickener B - emulsifier A” as the two-dimensional projection is given: the circle indicates that the constraint is satisfied and the prediction is passed, the cross indicates that the constraint is eliminated or the prediction is rejected, and the asterisk is the finally selected candidate.

[0072] The candidate formulation (mass percentage) and process parameters obtained in this embodiment are as follows: glycerin 5.764%, propylene glycol 0.000%, nicotinamide 3.754%, sodium hyaluronate 0.160%, emulsifier A 1.434%, thickener B 0.300%, preservative system C 0.322%, water 86.266%; emulsification temperature 76.5℃, homogenization speed 4140 r / min, homogenization time 4.6 min, cooling rate 1.00℃ / min. Mass conservation check: the sum of all components except water is 13.734%, therefore water is 100% - 13.734% = 86.266%, which satisfies the balance logic.

[0073] The predicted index for this candidate is: viscosity approximately 6398 mPa. s (protocol V2 upper limit 6500, lower limit 5200), pH approximately 5.804 (range 5.2–6.2), stratification height approximately 0.187 mm (upper limit 0.6), inhibition log reduction value approximately 5.00 (threshold 3.0). The corresponding margins are approximately: , , , ,therefore The weakest point is pH margin, indicating that the main risk of this candidate is not "stability / antibacterial", but "long-term drift of pH window", which makes it easier for researchers to carry out targeted reinforcement (such as fine-tuning of the buffer system).

[0074] To prove that the candidate solution is not a "chance hit", this embodiment introduces falsification verification: without changing the original method framework, perturbation consistency test and targeted counterexample search are performed on the candidate points to form a two-way evidence chain of "passing sample - counterexample sample".

[0075] Perturbation Consistency Test: Construct 19 groups of perturbations (e.g., small ± perturbations to nicotinamide, thickener B, emulsifier A, corrosion inhibitor system C, and key process parameters, with moisture compensation to maintain a total of 100%), calculate the difference between each group of perturbations and the baseline, and normalize according to the V2 tolerance protocol: for example, viscosity normalization difference. pH normalized difference Stratified normalization differences When all Furthermore, if the candidate is determined not to flip, it is considered robust to small perturbations. This is the viscosity normalization difference, used to characterize the relative deviation of the current sample viscosity from the reference sample viscosity. The viscosity of the current sample, in mPa. s. The viscosity of the reference sample, in mPa. s. Viscosity tolerance, limited by the testing protocol data, unit: mPa. s. The pH normalization difference is used to characterize the relative deviation of the current sample's pH from the baseline sample's pH. This represents the pH level of the current sample. The pH value is the reference sample. The pH tolerance is limited by the testing protocol data. This is a stratified normalization difference used to characterize the relative deviation of the current sample stratification height from the baseline sample stratification height. This represents the layer height of the current sample, in mm. This represents the layer height of the baseline sample, in mm. Layer height tolerance is limited by the testing protocol data, in mm. This is a general term for normalized differences, and its values ​​can be: , or .

[0076] The results showed that 18 out of 19 perturbations met the consistency and non-reversal criteria, while only 1 group triggered viscosity exceedance and reversed to fail when "thickener B was increased by 0.02%", indicating that the candidate was close to the viscosity limit and a stricter process control band for thickener metering deviation needed to be set during mass production. The above results are in... Figure 5 The results show that the normalized difference curves of most perturbation samples are below the threshold line of 1.0, while the "triggered flip" samples show obvious over-limit.

[0077] Targeted Counterexample Search: To verify the necessity of the graph constraints, a minimum-cost counterexample search is performed. Two types of key counterexamples are obtained: Firstly, a counterexample regarding viscosity: with the baseline formulation unchanged, only the thickener B is increased from 0.300% to 0.312% (an increase of approximately 0.012%), with the remainder compensated with water. The predicted viscosity is approximately 6509 mPa. The value of s exceeded the V2 limit of 6500, resulting in a failure to pass the test. This counterexample demonstrates that the sensitivity of "thickener B - viscosity limit" is objectively real, and viscosity margin must be explicitly considered in the constraints.

[0078] Secondly, a counterexample of stratification: While maintaining other indicators within the threshold, reducing emulsifier A by 0.25% (1.434% → 1.184%) and increasing the cooling rate to 1.8℃ / min resulted in a stratification height of approximately 0.604mm, slightly exceeding the 0.6mm upper limit, thus failing the test. This counterexample illustrates a linkage boundary between "cooling rate and emulsifier A": when cooling is too rapid, the lower limit of emulsifier concentration must be raised, or compensation must be achieved by extending the homogenization time / increasing the homogenization speed.

[0079] Based on the above counterexamples, this embodiment solidifies the edge conditions in the knowledge graph into executable constraints: (1) Total polyol content window: 4.5% ≤ ≤6.5%; (2) Viscosity risk limit: Set an upper limit for the combination of "emulsifier A + thickener B" to avoid exceeding the limit; (3) Cooling-emulsification linkage: When the cooling rate is at the high end, emulsifier A is required to be no less than the empirical threshold, or the stratification margin is met through process compensation items (homogenization speed / time). This forms a closed loop of "data - spectrum - candidate - disturbance verification - counterexample correction - re-screening".

[0080] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for predicting formulation performance based on component knowledge graph, characterized in that, The method includes: Obtain batch data for the formulation, including formulation composition data, process parameter data, packaging and storage condition data, testing protocol data, and test result data; Based on the batch data, a component knowledge graph is constructed and a set of argumentation rules is generated, which includes triggering conditions, threshold judgments, and rebuttal conditions. Based on the set of argumentation rules, an attack relationship graph is constructed and a minimum inconsistency set is determined. A constraint set is generated based on the minimum inconsistency set. Candidate solutions are obtained by solving the constraint set within the candidate formula and candidate process window, and performance prediction results are obtained. The candidate solution is subjected to falsification verification, which includes generating perturbation samples based on the constraint set and performing consistency checks and targeted counterexample searches to obtain counterexample records, and updating the argumentation rule set and the constraint set based on the counterexample records and outputting the updated performance prediction results.

2. The method according to claim 1, characterized in that, Obtaining the batch data includes: Generate batch identifier; The batch data is formed by associating and storing the formula composition data, process parameter data, packaging and storage condition data, testing protocol data, and testing result data with the batch identifier. Perform field mapping and unit normalization processing on the batch data.

3. The method according to claim 1, characterized in that, Constructing the component knowledge graph includes: The components are used as nodes in the component knowledge graph; the compatibility, incompatibility, and substitution relationships between components are used as edges in the component knowledge graph. Based on the batch data, compatibility, incompatibility, and substitution relationships are determined. The compatibility relationship indicates the component combination whose test result data meets the performance threshold under the same testing protocol data, the same process parameter data, and the same packaging and storage conditions data. The incompatibility relationship indicates the component combination whose test result data does not meet the performance threshold under the same testing protocol data, the same process parameter data, and the same packaging and storage conditions data. The substitution relationship indicates that, under the premise of meeting the performance threshold, the difference in the test result data after the first component is replaced by the second component does not exceed the tolerance defined by the testing protocol data. The performance threshold is given by the detection protocol data.

4. The method according to claim 3, characterized in that, Generating the set of argument rules includes: Based on the component knowledge graph and the batch data, the triggering condition, the threshold determination and the rebuttal condition are generated for each argumentation rule; The triggering conditions are jointly defined by the formula composition data, process parameter data, packaging and storage condition data, and testing protocol data; The threshold determination is used to indicate whether the performance indicators in the performance prediction results meet or do not meet the performance threshold. The rebuttal condition is used to indicate a combination of conditions in which the threshold determination changes from satisfying the performance threshold to not satisfying the performance threshold when at least one of the trigger conditions changes.

5. The method according to claim 1, characterized in that, Constructing the attack relationship graph includes: When the triggering conditions are consistent, and the threshold determination of the first argument rule and the threshold determination of the second argument rule are opposite to the same performance index in the performance prediction result, an attack relationship is established between the first argument rule and the second argument rule in the attack relationship graph. The attack relationship is used to indicate that the first argument rule and the second argument rule cannot be true at the same time.

6. The method according to claim 5, characterized in that, Determining the minimum inconsistency set includes: finding the minimum subset of argument rules in the attack relationship graph that makes it impossible for the set of argument rules to be true at the same time as the minimum inconsistency set; Generating the constraint set includes: converting the rebuttal conditions corresponding to the minimum inconsistency set into condition-triggered constraints; the condition-triggered constraints are constraints that take effect when the triggering conditions are met, and the condition-triggered constraints are used to tighten the allowable range of component dosage values ​​or tighten the range of process parameters, or to prohibit the combination of component dosage values ​​and process parameter ranges.

7. The method according to claim 1, characterized in that, Solving the constraint set within the candidate formulation and candidate process window includes: Candidate solutions are generated within the range of component dosage values ​​defined by the candidate formulation and the range of process parameters defined by the candidate process window; Perform constraint set verification on the candidate solution, the constraint set verification including determining whether the candidate solution satisfies each constraint in the constraint set; The candidate solutions that satisfy the set of constraints are determined as the candidate schemes.

8. The method according to claim 1, characterized in that, Generating the perturbation sample includes: under the premise of satisfying the constraint set, increasing or decreasing the perturbation of the component dosage values ​​of the candidate scheme by a preset step size, and adjusting at least one unperturbed component dosage value so that the sum of the component dosage values ​​remains unchanged, thereby obtaining a formulation perturbation sample; under the premise of satisfying the constraint set, perturbing the process parameter range of the candidate scheme at the upper or lower boundary by the preset step size, thereby obtaining a process perturbation sample; The consistency check includes: comparing the performance prediction results of the candidate scheme with those of the perturbation sample, and determining that the threshold determination does not flip under the perturbation sample, and that the difference in performance prediction results does not exceed the tolerance limit defined by the detection protocol data.

9. The method according to claim 1, characterized in that, The targeted counterexample search includes: taking the boundary point between the component dosage value and the process parameter range of the candidate solution as the search starting point, iteratively adjusting the component dosage value or process parameter value so that the adjusted component dosage value or process parameter value satisfies the rebuttal condition, and generating the counterexample record; The counterexample record includes a minimum trigger change, which indicates the minimum amount of variable change that causes the threshold determination to change from satisfied to unsatisfied.

10. The method according to claim 1, characterized in that, Updating the set of argumentation rules and the set of constraints based on the counterexample records includes: Add the rebuttal condition corresponding to the counterexample record to the argument rule set, or replace the rebuttal condition corresponding to the same triggering condition in the argument rule set; The minimum inconsistency set is redefined based on the updated set of argumentation rules, and an updated set of constraints is generated; the updated performance prediction result is output based on the updated set of constraints.