A method and system for small scale screening and invalid formulation elimination in a research and development process

CN122598801APending Publication Date: 2026-08-18TONGYI PETROLEUM CHEM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610656612.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

而且传统方法只输出一个孤立的合格配方,没有记录从原配方一步步调整到最优区间的移动轨迹,研发人员无法理解调整逻辑,也难以积累经验

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598801A_ABST
    Figure CN122598801A_ABST
Patent Text Reader

Abstract

The application discloses a kind of small sample screening and invalid formula exclusion method and system in research and development process, utilize NLP from oil product research and development basic data extraction entity, form triad based on chemical knowledge database and store into graph database and set weight, solve semantic inconsistency, relationship is fuzzy, provide basis for priori reasoning and antagonism / coordination search.It is collected in controlled environment to small sample data, and is preprocessed into standardized data set;Import thermokinetics model, calculate heat value promotion contribution rate and stability attenuation, substitute into conflict deviation degree formula, and then move into incompatible formula list if it is more than 0.75;Based on the performance correlation path of priori logic of knowledge graph, provide intermediate variable chain;Compare path with quality standard, identify over-limit logic node, and output compatible / incompatible formula list;For improvable formula, according to antagonism / coordination, generate feasible interval constraint, find optimal interval by discrete boundary enumeration and backtracking search, and generate optimization path containing moving track.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oil product research and development, and in particular to a method and system for small-sample screening and elimination of ineffective formulations during the research and development process. Background Technology

[0002] Small-scale sample screening and elimination of ineffective formulations are crucial steps in fuel product development. The core objective is to quickly identify effective formulations from a large pool of candidate formulations that meet performance standards, demonstrate reliable stability, and have mass production potential, while simultaneously eliminating ineffective formulations that exhibit performance conflicts or insufficient environmental adaptability. The traditional approach involves relying on the experience and knowledge of R&D personnel, combined with limited historical experimental data, to formulate formulations through trial and error. Several key indicators (such as calorific value, emissions, and stability) are then tested using combustion bench experiments, and finally, the effectiveness of the formulation is judged based on a single indicator threshold.

[0003] However, in the traditional approach, a large amount of knowledge, such as the effects of olefins on calorific value, humidity constraints, stability, and the antagonism of certain additives against certain aromatics, recorded in research literature, component safety data, and industry standards, is stored in scattered, unstructured text. When faced with a new formulation, researchers can only rely on memory or manual searching, unable to form machine-readable triple relationships. This results in any subsequent causal reasoning lacking prior logical constraints, forcing them to blindly search for variable relationships within the data.

[0004] The inability to quantify the conflict between positive benefits and negative costs leads to missed detection of formulations with performance exceeding limits. Traditional methods test the calorific value and stability under standard operating conditions separately and compare each to a threshold. However, while high-olefin components increase calorific value, they often cause a sharp decline in combustion stability under temperature and humidity fluctuations. This conflict between benefits and costs cannot be identified by single-index threshold methods.

[0005] Conventional causal analysis cannot capture nonlinear bifurcation abrupt changes, making it impossible to predict the stability critical point. Traditional data-driven models (regression, neural networks) provide linear approximations or black-box mappings, which cannot reflect the causal chain of olefin ratio → ignition delay period → combustion isochoricity → calorific value, nor can they detect the bifurcation critical point where the combustion pressure signal phase space trajectory transitions from the limiting cycle to the chaotic attractor when the olefin ratio changes continuously. It is precisely near this critical point that combustion stability deteriorates precipitously.

[0006] Traditional trial-and-error methods simply label incompatible formulations as unqualified without providing a traceable optimization path. After a formulation is deemed invalid, the traditional method relies solely on experience to manually adjust component ratios and repeat the experiment. When the component-performance mapping exhibits non-convex, multi-peak characteristics (e.g., the calorific value initially increases, then decreases, and then increases again when the olefin ratio increases), gradient descent is prone to getting trapped in local optima. Furthermore, traditional methods only output an isolated, qualified formulation, without recording the step-by-step adjustment trajectory from the original formulation to the optimal range. Researchers cannot understand the adjustment logic and find it difficult to accumulate experience. Summary of the Invention

[0007] The purpose of this invention is to provide a method and system for small-sample screening and invalid formula elimination in the research and development process, which solves the above-mentioned technical problems pointed out in the prior art.

[0008] This invention provides a method for small-sample screening and elimination of ineffective formulations during the research and development process, comprising the following steps:

[0009] Entities are extracted from basic data of oil product research and development using natural language processing. Relationships between these entities are established based on a pre-set chemical knowledge database, forming a structured set of triples and storing it in a graph database, thus constructing an oil product research and development knowledge graph.

[0010] Real-time acquisition of sample test data, and preprocessing of the sample test data to form a standardized experimental dataset;

[0011] The standardized experimental dataset is imported into a thermodynamic analysis model. Energy conservation equations and mass conservation equations are constructed using the thermodynamic analysis model with the imported standardized experimental dataset. The oxidation reaction rate of high olefin components under different oxygen content environments is simulated. The contribution rate of high olefin components to the improvement of overall combustion heat value within a specific proportion range is calculated. At the same time, variance analysis is used to identify the amount of combustion stability decay caused by the high olefin components under the conditions of fluctuating ambient temperature and humidity. The conflict deviation between the improvement contribution rate and the amount of combustion stability decay is calculated.

[0012] Based on the prior logic provided by the oil product R&D knowledge graph, the standardized experimental dataset is analyzed by applying a causal reasoning algorithm to extract the nonlinear mapping relationship between component ratio, environmental variables and performance output, and to construct a performance association path from the input end to the output end.

[0013] The performance-related path is compared with the preset performance quality standards to identify out-of-limit logic nodes. Based on the out-of-limit logic nodes, the formula is divided into a suitable formula list, a formula list to be improved, and an unsuitable formula list. At the same time, the formula list to be improved is used to find the optimal component ratio range that meets the performance constraints based on the conflict deviation degree, discrete boundary enumeration, and backtracking search, thereby generating an optimized formula path.

[0014] Preferably, the relationships between entities include relationship type labels and corresponding confidence scores; the relationship type labels include antagonistic relationships and cooperative relationships.

[0015] The standardized experimental dataset includes component ratios, environmental variables, intermediate combustion states, and performance indicators. Component ratios include mass fraction values.

[0016] Preferably, based on the prior logic provided by the oil product R&D knowledge graph, a causal reasoning algorithm is applied to analyze the standardized experimental dataset, extracting the nonlinear mapping relationship between component ratios, environmental variables, and performance outputs, and constructing a performance association path from the input end to the output end, including:

[0017] Using structural equation modeling, the standardized experimental dataset is used as a variable to perform causal search under the prior logical constraints provided by the oil product R&D knowledge graph.

[0018] In the causal search process, in view of the coupling effect between the change in the proportion of high olefin components and the environmental humidity variable, based on the triplet constraint relationship between the olefin entity and the moisture content entity stored in the oil product research and development knowledge graph, combustion state bifurcation analysis is introduced. By calculating the bifurcation parameters of the phase space trajectory of the combustion pressure signal under different high olefin proportion steps and environmental humidity combinations, the critical conditions for the nonlinear transformation of combustion mode caused by the change in high olefin proportion are identified.

[0019] Based on the critical conditions for the nonlinear transition of the combustion mode and the edge weights stored in the oil product R&D knowledge graph, hypothetical environmental variable perturbations are applied to each formulation sample in the standardized experimental dataset through intervention experiment simulation. The changes in performance output are collected, and the nonlinear mapping relationship between component ratio, environmental variables and performance output is extracted to construct a performance correlation path including direct influence path, indirect influence path and interactive influence path.

[0020] Preferably, the prior logic constraints include triplets of influence relationships between components and performance, triplets of antagonistic or synergistic relationships between components, and triplets of constraint relationships between environmental factors and performance indicators stored in the oil product R&D knowledge graph. Each triplet is associated with a confidence score and edge weight. The interaction influence path includes the obstacle path of high olefin components and environmental humidity on combustion efficiency through the emulsification tendency variable.

[0021] Preferably, by simulating intervention experiments, hypothesized environmental variable perturbations are applied to each formulation sample in the standardized experimental dataset, and the changes in performance output are collected, including:

[0022] Obtain the component proportion vector, environmental variable vector, and performance output vector corresponding to each formulation sample in the standardized experimental dataset. The component proportion vector includes the mass fraction of high olefins, the mass fraction of aromatics, the mass fraction of saturated hydrocarbons, and the mass fraction of each additive. The environmental variable vector includes the temperature value, humidity value, and intake pressure value. The performance output vector includes the calorific value of combustion, the standard deviation of combustion pressure oscillation, and the concentration of exhaust gas emissions.

[0023] For each formulation sample, based on the edge weights of the triplet constraint relationship between olefin entities and moisture content entities stored in the oil product R&D knowledge graph and the critical conditions, the environmental humidity perturbation response function of the formulation sample is constructed.

[0024] Based on the environmental humidity disturbance response function, a disturbance is applied to each formulation sample, and the change in the standard deviation of combustion pressure oscillation under each disturbance is calculated. The degree to which the change in the standard deviation of combustion pressure oscillation exceeds the threshold of the bifurcation parameter corresponding to the critical condition of the nonlinear transition of the combustion mode is taken as the change in performance output.

[0025] Preferably, the change in the standard deviation of combustion pressure oscillation under each disturbance is calculated, and the degree to which the change in the standard deviation of combustion pressure oscillation exceeds the bifurcation parameter threshold corresponding to the critical condition for the nonlinear transition of the combustion mode is used as the change in performance output, including:

[0026] Obtain the baseline combustion pressure oscillation standard deviation of the formulation sample under undisturbed state, and the perturbed combustion pressure oscillation standard deviation of the formulation sample under a positive humidity perturbation state, and calculate the deviation of the perturbed combustion pressure oscillation standard deviation relative to the baseline combustion pressure oscillation standard deviation.

[0027] Obtain the critical stability change corresponding to the critical condition for the nonlinear transition of the combustion mode of the formula sample;

[0028] The deviation of the standard deviation of combustion pressure oscillation is compared with the change in critical stability. The excess of the deviation exceeding the change in critical stability is calculated, and the excess is mapped to the dimensionless change in performance output within the interval [0,1].

[0029] Preferably, based on the conflict deviation degree combined with discrete boundary enumeration and backtracking search, the optimal component ratio range that satisfies the performance constraints is found, and an optimized formulation path is generated, including the following steps:

[0030] The list of incompatible formulations is traversed, and formulations with a conflict deviation greater than the conflict deviation threshold but less than the second conflict deviation threshold, and which simultaneously meet the preset performance and quality standards, are selected as formulations to be improved.

[0031] For each formulation to be improved, obtain the names of each component in the component ratio vector; in the oil product R&D knowledge graph, query the influencing components that have antagonistic or synergistic relationships with the formulation to be improved based on each component name; and form a set of feasible interval constraints for component ratios based on all influencing components.

[0032] For each component in the feasible interval constraint set of component proportions, a candidate proportion list for that component is generated based on the upper and lower boundaries and coupling adjustment step rules corresponding to each component; an initial grid point set is constructed based on all candidate proportion lists.

[0033] For each candidate ratio list in the initial grid point set, the predicted value is calculated by importing the thermodynamic analysis model of the standardized experimental dataset; the predicted value is compared and screened according to the preset performance quality standards to obtain a feasible grid subset consisting of multiple feasible grid points;

[0034] Multiple starting grid points are selected based on the Euclidean distance between each feasible grid point in the feasible grid subset and the component ratio vector of the formulation to be improved; the starting grid points and their corresponding neighboring grid points are traversed to analyze and obtain the candidate optimal component ratio range for each starting grid point and the corresponding movement trajectory of each starting grid point.

[0035] The target optimal component ratio range is generated by perturbing the ambient humidity based on the candidate optimal component ratio range; the target optimal component ratio range is then connected with the movement trajectory to obtain the optimized formulation path.

[0036] Preferably, a feasible set of component proportion constraints is formed based on all influencing components, including the following operational steps:

[0037] Extract the original component names of each component from the component ratio vector of the formulation to be improved, and obtain the set of original component names;

[0038] Using each component name in the original component name set as the query key, retrieve all triples with the original component name as the head or tail entity in the oil product R&D knowledge graph; take the other component in the retrieved triple whose name is different from the original component name as the influencing component, record the relationship type and confidence score of the triple, and determine the adjustment direction information according to the order of the head and tail entities in the triple.

[0039] Constraint parameters are extracted from the attribute constraint layer of the knowledge graph of oil product R&D by the relationship type of the triples corresponding to each influencing component.

[0040] Calculate the upper and lower boundaries of the feasible proportion interval for each influencing component based on the constraint parameters.

[0041] Remove all names of influencing components from the original set of component names to obtain a set of non-influencing component names; for each component in the set of non-influencing component names, calculate the upper boundary and lower boundary of the feasible non-influencing ratio interval by combining the upper and lower floating ratios of the default adjustment range in the attribute constraint layer with the current quality score of each component.

[0042] By summarizing the upper and lower boundaries of the feasible proportion intervals for all influencing components, along with the upper and lower boundaries of the non-influencing feasible proportion intervals, we obtain the set of component proportion feasible interval constraints for the formulation to be improved.

[0043] Preferably, the constraint parameters include antagonistic constraint parameters and cooperative constraint parameters. Specifically, the antagonistic constraint parameters include a preset upper limit scaling factor and a minimum effective addition amount, while the cooperative constraint parameters include a synchronous adjustment step value, a maximum allowable adjustment offset, a minimum effective addition amount, and a maximum allowable addition amount.

[0044] Accordingly, this invention also proposes a small-sample screening and invalid formula elimination system in the research and development process, including a knowledge graph construction module, a data acquisition module, a deviation analysis module, a correlation analysis module, and an optimization module;

[0045] Among them, the knowledge graph construction module is used to extract entities from basic data of oil research and development using natural language processing, establish the relationship between the entities based on a preset chemical knowledge database, form a structured set of triples and store them in a graph database, and construct an oil research and development knowledge graph.

[0046] The data acquisition module is used to collect sample test data in real time and preprocess the sample test data to form a standardized experimental dataset.

[0047] The deviation analysis module is used to import the standardized experimental dataset into the thermodynamic analysis model. By importing the standardized experimental dataset into the thermodynamic analysis model, energy conservation equations and mass conservation equations are constructed to simulate the oxidation reaction rate of high olefin components under different oxygen content environments. The contribution rate of high olefin components to the overall combustion heat value improvement within a specific proportion range is calculated. At the same time, variance analysis is used to identify the amount of combustion stability decay caused by the high olefin components under the conditions of fluctuating ambient temperature and humidity, and the conflict deviation between the improvement contribution rate and the combustion stability decay is calculated.

[0048] The association analysis module is used to analyze the standardized experimental dataset based on the prior logic provided by the oil product R&D knowledge graph, and to extract the nonlinear mapping relationship between component ratio, environmental variables and performance output by applying causal reasoning algorithms, and to construct a performance association path from the input end to the output end.

[0049] The system also includes an optimization module, which compares the performance-related path with preset performance quality standards, identifies out-of-limit logic nodes, and divides the formula into a suitable formula list, a formula list to be improved, and an unsuitable formula list based on the out-of-limit logic nodes. At the same time, for the formula list to be improved, the system uses conflict deviation degree combined with discrete boundary enumeration and backtracking search to find the optimal component ratio range that meets the performance constraints, and generates an optimized formula path.

[0050] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:

[0051] Analysis of the above-mentioned method and system for small-sample screening and invalid formula elimination in the research and development process provided by the present invention shows that, in specific applications, natural language processing is used to extract entities related to hydrocarbons, additives, physicochemical properties and performance indicators from basic data of oil product research and development, such as historical research and development literature, experimental reports, component safety instructions and industry standards; the association relationship between entities is established based on a preset chemical knowledge database to form a structured set of triplets, and stored in a graph database. Simultaneously, dynamic initial weights are assigned to each triple edge based on historical frequency and source authority. Furthermore, the triples stored in the knowledge graph provide prior logical constraints for subsequent causal inference algorithms and offer retrieval basis for antagonistic / synergistic component adjustments. Next, in a controlled experimental environment, ignition and combustion experiments are conducted on multiple oil formulations with different component ratios. Real-time small-sample test data, including combustion calorific value, intake pressure, exhaust gas composition, temperature fluctuations, and humidity changes, are collected using a multi-point sensor cluster. The collected data undergoes data cleaning, noise smoothing, and normalization to form a standardized experimental dataset containing component ratios, environmental variables, intermediate combustion states, and performance indicators. This dataset provides calibrated and validated basic input data for the thermodynamic analysis model and also provides computable experimental samples for the causal inference algorithm. The controlled environment ensures that the collected test data primarily reflects the impact of changes in formulation components, rather than the impact of random environmental fluctuations.

[0052] Furthermore, standardized experimental datasets were imported into a thermodynamic analysis model to construct energy and mass conservation equations, simulating the oxidation reaction rate of high-olefin components under different oxygen contents, and calculating the contribution rate of high-olefin components to the overall combustion calorific value improvement. Simultaneously, variance analysis was used to identify the combustion stability degradation caused by fluctuations in ambient temperature and humidity. The contribution rate and stability degradation were substituted into the conflict deviation formula; when the deviation exceeded the threshold of 0.75, it was determined to be a performance over-limit conflict. Positive benefits (calorific value improvement contribution rate) and negative costs (combustion stability degradation) were incorporated into a unified quantitative framework, solving the problem that traditional methods, which analyze calorific value or stability separately, cannot identify the internal imbalance between positive and negative benefits. A conflict deviation greater than 0.75 directly determined that the formulation had a conflict where the calorific value benefit was offset by the cost of stability deterioration, and this formulation was moved to the list of unsuitable formulations as a screening basis for subsequent formulations to be improved. The contribution increment and degradation were normalized using the baseline calorific value and standard operating condition stability as normalization benchmarks, respectively, to ensure comparability between different formulations.

[0053] Furthermore, based on the prior logic provided by the knowledge graph, a causal reasoning algorithm is applied to the standardized experimental dataset to extract the nonlinear mapping relationship between component ratios, environmental variables, and performance outputs. This constructs a performance-related path from input to output. The prior logic in the knowledge graph constrains the direction of the causal search, reduces the complexity of the search space, and ensures the physical interpretability of the discovered causal path. The constructed performance-related path provides a structured chain of intermediate variables for identifying out-of-limit logic nodes, allowing each intermediate variable to be compared with the quality standard interval one by one. The interactive influence path specifically captures the hindering effect of high-olefin components and environmental humidity on combustion efficiency, compensating for the inability of traditional linear analysis to capture coupling effects.

[0054] Finally, the performance-related paths are compared item by item with preset performance quality standards to identify out-of-limit logical nodes that pose a risk of exceeding performance limits under extreme operating conditions. Based on this, a set of suitable formulations and a list of unsuitable formulations are divided and output. For formulations with improvement potential among the unsuitable formulations, a set of feasible component ratio constraint sets is generated based on antagonistic / synergistic relationships in the knowledge graph. The optimal component ratio range that satisfies all performance constraints is found through discrete boundary enumeration and backtracking search, and an optimized formulation path from the original formulation to the optimal range is generated. Through point-by-point comparison of out-of-limit logical nodes, the system accurately locates which intermediate variable or final indicator exceeds the limit in the performance-related path of the formulation. It overcomes the deficiency of relying solely on the final performance threshold to pinpoint the failure point; it generates a constraint set based on antagonistic and synergistic relationships in the knowledge graph, making the search space more consistent with chemical mechanisms and avoiding meaningless grid points; it adopts discrete boundary enumeration and backtracking search to solve the problem that conventional gradient descent methods are prone to getting trapped in local optima under non-convex and multi-peak mappings, and can find the globally optimal component ratio range; the output optimized formulation path fully records every step of the movement trajectory from the original unsuitable formulation to the optimal range, and R&D personnel can adjust step by step according to the trajectory, and the reliability of the range is ensured by robustness verification (humidity disturbance test did not trigger nonlinear transition of combustion mode). Attached Figure Description

[0055] Figure 1 This is a schematic diagram of the main process of a method for small-sample screening and elimination of ineffective formulations in the research and development process;

[0056] Figure 2 This is a schematic diagram of a triplet set simulation in a sample screening and invalid formulation elimination method during the research and development process.

[0057] Figure 3 This is a schematic diagram of a nonlinear S-curve simulation of humidity-high olefin coupling in a small-sample screening and invalid formulation elimination method during the research and development process.

[0058] Figure 4 This is a schematic diagram simulating the S-shaped growth relationship between excess quantity and performance output in a sample screening and invalid formula elimination method during the research and development process.

[0059] Figure 5 This is a schematic diagram of the overall architecture of a sample screening and invalid formula elimination system in the research and development process.

[0060] Attached image labels: Knowledge graph construction module 10, data acquisition module 20, deviation analysis module 30, association analysis module 40, optimization module 50. Detailed Implementation

[0061] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.

[0063] Example 1 like Figure 1 As shown, Embodiment 1 of the present invention provides a method for small-sample screening and elimination of ineffective formulations during the research and development process, including the following steps:

[0064] Step S10: Using natural language processing, extract entities related to hydrocarbons, additives, physicochemical properties, and performance indicators from basic oil product R&D data (which includes historical documents on oil product R&D, experimental reports on oil product R&D, safety instructions for oil product components, and industry standards for oil product R&D; the above basic data is mainly text data). Establish the relationships between the entities based on a preset chemical knowledge database, form a structured set of triplets, and store them in a graph database to construct an oil product R&D knowledge graph (covering the logic of component characteristics and performance).

[0065] It should be noted that in the above embodiments of this application, the preset chemical knowledge database is such as a publicly available chemical knowledge database, or some preferred industry databases from PubChem, NISTChemistryWebBook, and Reaxys in the embodiments of this invention; establishing the association between the entities is done by calling the publicly available chemical knowledge database and then obtaining the predefined relationships between components through API (e.g., "hexadecane number" and "CN value" are synonyms; "olefin content" and "calorific value" have a positive correlation); these relationships are directly mapped into a triplet format and assigned a default confidence level of 0.9;

[0066] Then, identical entities extracted from different sources are aligned (cosine similarity is used to calculate the similarity of entity names and attributes, and entities exceeding the threshold of 0.95 are merged) to eliminate redundancy and ensure the uniqueness of entities. Next, the finally confirmed triples are stored in a graph database (such as Neo4j), and dynamic initial weights are set for each edge according to the frequency of occurrence in historical experiments and the authority of the source: w(e)=0.5×freq(e)+0.5×aauth(e), where freq(e) is the normalized value of the frequency of occurrence of the knowledge relation in historical documents, and aauth(e) is the source authority score (1 for Q1 journals, 0.8 for Q2 journals, and 0.6 for internal experimental reports). Finally, w(e)∈[0,1].

[0067] Specifically, the above relationship types include: influence (for example, the ternary group is <olefin content, influence, calorific value>, indicating that changes in the proportion of olefins in the formulation will have a positive or negative effect on the calorific value), antagonism (for example, the ternary group is <detergent dispersant X, antagonism, aromatic Y>, indicating that detergent dispersant X and aromatic Y have a mutually weakening effect on cleaning performance, and the adjustment direction information is: head entity (detergent dispersant X) antagonizes tail entity (aromatic Y), that is, the adjustment direction is to prioritize limiting the upper limit of the proportion of head entity), synergy (for example, the ternary group is <antioxidant A, synergy, antioxidant B>, indicating that the effect of using two antioxidants in combination is better than the sum of their individual uses, and the adjustment direction information is that both should be adjusted in the same direction and synchronously), and constraint (for example, the ternary group is <ambient humidity, constraint, combustion stability>, indicating that ambient humidity is an external condition limiting factor affecting combustion stability).

[0068] A structured set of triples refers to a collection of data units in the form of <head entity, relation, tail entity>. Each triple represents a knowledge fragment, and all triples together constitute the data layer of a knowledge graph.

[0069] Structured sets of triples are presented in graph structure form (e.g.) Figure 2 (As shown) Stores knowledge related to oil product research and development, supports efficient association queries and path retrieval, and provides prior logical constraints for subsequent causal reasoning algorithms. For example, when searching for the relationship between components and the environment, existing influence or constraint edges in the knowledge graph can be traversed first. The confidence score and edge weight attached to the triple can be used to dynamically adjust the weight coefficients in subsequent calculations.

[0070] For example: The graph database of a knowledge graph stores the following triples:

[0071] <High olefin content, impact, calorific value, confidence level 0.92>

[0072] <High olefin content, impact on combustion stability, confidence level 0.85>

[0073] <Ambient humidity, constraint, combustion stability, confidence level 0.78>

[0074] <Base oil PAO6, synergistic effect, antioxidant A, confidence level 0.63>

[0075] These ternary groups together constitute a structured knowledge fragment describing the conflict between the calorific value contribution and stability risk of high olefin components;

[0076] The aforementioned construction of a knowledge graph for oil product R&D, encompassing component characteristics and performance logic, begins by using an NLP model to identify all named entities related to hydrocarbons, additives, physicochemical properties, and performance indicators from over 500,000 historical R&D documents, experimental reports, MSDS, and industry standards. Then, by analyzing these entities in conjunction with a chemical knowledge database, the relationships between them are obtained, forming triples, which are then stored in a graph database to construct the oil product R&D knowledge graph (i.e., a graph database based on triples that stores relationships between components, performance, and the environment).

[0077] Finally, since the extracted entities are themselves divided into component characteristic categories (such as high olefin ratio and kinematic viscosity) and performance categories (such as calorific value and emission concentration), and the relationship types clearly include influence, constraint, antagonism, and synergy, the stored map naturally expresses the logical relationship chain between component characteristics and performance, that is, it covers the logic of component characteristics and performance.

[0078] Step S20: (In a controlled experimental environment, ignition and combustion experiments are conducted on oil formulations with different component ratios. A cluster of multi-point sensors deployed inside and outside the combustion chamber is used to) collect small sample test data in real time (including combustion calorific value, intake pressure, exhaust gas composition, temperature fluctuations, and humidity changes). The small sample test data is then preprocessed (preprocessing includes data cleaning, noise smoothing, and normalization) to form a standardized experimental dataset (the standardized experimental dataset includes component ratios, environmental variables, intermediate combustion states, and performance indicators, where the component ratios include mass fraction values).

[0079] It should be noted that, in the embodiments of this application described above, the controlled experimental environment refers to a combustion laboratory capable of precisely adjusting and maintaining environmental variables such as temperature, humidity, intake pressure, and intake oxygen content. This laboratory is equipped with a high-precision temperature control system (fluctuation range controlled within ±0.1℃), a humidity control system, a pressure regulation system, and an exhaust emission analysis system. This ensures that in each small-sample ignition and combustion experiment, the deviation between the set values ​​and actual values ​​of the environmental variables is within an acceptable range, thereby ensuring that the collected test data primarily reflects the impact of changes in the formulation components, rather than the impact of random environmental fluctuations.

[0080] Examples of oil formulations with different component ratios:

[0081] The oil formulation is made by mixing base oil and various additives in a certain mass fraction. Researchers use an automatic dispensing pump to precisely control the amount of each component added, forming a series of test samples with gradient changes.

[0082] Example formulation set (as shown in Table 1, used to study the effect of changes in the proportion of high olefins):

[0083] F-01 85 5 8 1 1 F-02 82.5 7.5 8 1 1 F-03 80 10 8 1 1 F-04 77.5 12.5 8 1 1 F-05 75 15 8 1 1

[0084] The above is a set of oil formulations with different component ratios, wherein the olefin ratio increases in increments of 0.5% or 2.5%.

[0085] Example of real-time sample test data:

[0086] During the combustion experiment, a multi-point sensor cluster was deployed to collect data at a frequency of 50kHz. After preprocessing, a standardized dataset was obtained. Example data snippets (as shown in Table 2, snapshots of data at a certain moment):

[0087] Calorific value 43.2 MJ / kg Calorific value analyzer Intake pressure 101.5 kPa Intake manifold pressure sensor Exhaust gas composition (CO) 0.15%vol Infrared Spectrometer Exhaust Gas Analyzer Exhaust gas components (HC) 120ppm Infrared Spectrometer Exhaust Gas Analyzer Temperature fluctuations 25.1℃ (mean ± 0.2℃) Combustion chamber multi-point thermocouple Humidity changes 55.3%RH (mean ± 0.5%RH) Humidity sensor

[0088] Step S30: Import the standardized experimental dataset into the thermodynamic analysis model. (The main purpose of importing the standardized experimental dataset is to correct the key uncertain parameters in the thermodynamic analysis model. The correction targets are the Arrhenius parameters of each elementary reaction (especially the olefin oxidation reaction), namely the pre-exponential factor (A) and activation energy (Ea), because these parameters are often incomplete or have large deviations in the literature. Specifically, the correction involves selecting some measured data from the standardized experimental dataset (e.g., the calorific value of combustion, pressure versus time curves, and concentration of key products of formulations F-01 to F-03 under standard operating conditions) as the training set. Then, the objective function is defined as the root mean square error (RMSE) between the model prediction and the measured value. Next, optimization algorithms (such as particle swarm optimization, Bayesian inversion, or genetic algorithms) are used to repeatedly adjust the Arrhenius parameters. After each adjustment, the thermodynamic model solver is called to calculate the simulated value under the current parameters, and the mean square error is calculated.) The root mean square error (RMSE) is used to calculate the performance degradation rate. Optimization stops when the RMSE is less than a preset threshold (e.g., a system-stored preset threshold of 1%) or the number of iterations reaches its limit, and the corrected parameter set is output. The corrected model can then be used for subsequent simulations. Energy and mass conservation equations are constructed using a thermodynamic analysis model imported from a standardized experimental dataset. The oxidation reaction rate of high-olefin components under different oxygen content environments is simulated, and the contribution rate of high-olefin components to the overall combustion heat value within a specific proportion range is calculated. Simultaneously, variance analysis is used to identify the combustion stability degradation caused by fluctuations in ambient temperature and humidity. The performance exceedance conflict caused by changes in the high-olefin proportion is quantified by calculating the conflict deviation between the contribution rate and the combustion stability degradation (i.e., the formulations corresponding to the performance exceedance conflict are moved to an internal conflict risk list. This list is stored independently and only contains formulations with a conflict deviation greater than 0.75).It should be noted that this internal conflict risk list is only used to mark formulations with a risk of imbalance between calorific value gains and stability costs, and is not the final formulation classification result; the final formulation classification is carried out in step S50 by combining the performance quality standard comparison results for unified classification; specifically, step S30 of the above-mentioned embodiment of this application firstly uses the partial sample test data obtained in step S20 (partial sample test data is a standardized experimental dataset, such as the measured data of formulations F-01 to F-03 with gradient changes in olefin ratio under standard operating conditions) to calibrate the parameters of the thermodynamic analysis model, that is, to determine the key reactions in the model (such as olefins) through the inversion method. The Arrhenius parameters (pre-factor, activation energy) of hydrocarbon oxidation are calculated. Then, based on the calibrated model (i.e., the thermodynamic analysis model with the standardized experimental dataset imported above), the combustion process is simulated for formulation ratios that are not directly tested (such as F-04 and F-05) and extreme environmental conditions (such as 45℃ and 90%RH), and the outputs such as combustion calorific value and pressure oscillation under each condition are calculated. Finally, based on the simulation results (the simulation results specifically include combustion calorific value, pressure oscillation, etc. under each condition), the contribution rate of high olefin components to the improvement of combustion calorific value and the amount of combustion stability decay caused by environmental fluctuations are calculated, thereby quantifying the performance over-limit conflict.

[0089] It should be noted that, in the embodiments of this application described above, the thermodynamic analysis model is a set of mathematical equations based on the principles of chemical reaction kinetics and thermodynamics, used to describe the chemical changes that occur in the fuel within the combustion chamber and the accompanying energy release. This model originates from classical chemical reaction engineering and combustion theory, specifically including the Arrhenius equation (describing the relationship between the reaction rate constant and temperature), the law of mass action (describing the relationship between the reaction rate and reactant concentration), the energy conservation equation, and the component mass conservation equation. In this application, the model is implemented as a numerical solver, running on a high-performance computing cluster, to simulate the combustion process of different formulations (especially when the proportion of high olefins changes) under given environmental conditions; and to calculate the formation rates of key intermediate products and final combustion products; simultaneously quantifying the theoretical contribution of high olefin components to the calorific value of combustion and the decay of combustion stability due to changes in the reaction pathway.

[0090] The above energy conservation equation is based on the first law of thermodynamics within the control volume of the combustion chamber:

[0091] ;

[0092] in, In order to control the body's energy, This represents the rate of change of internal energy over time (unit: J / s or W). This item reflects the speed at which energy accumulates or is released within the combustion chamber; The heat released by the combustion reaction (calculated from the enthalpy of each elementary reaction). It is the specific enthalpy (unit: J / kg) of the substance (such as fuel or air) entering the combustion chamber. It is the specific enthalpy (unit: J / kg) of the exhaust gas discharged from the combustion chamber, and the two are multiplied by the mass flow rate. and Then, the enthalpy flow rate (J / s) of the incoming and outgoing airflow is obtained. In order to do good for external purposes, For the inlet and outlet enthalpy flow, in a closed constant-volume incendiary bomb testing environment, it can be simplified as follows:

[0093] ;

[0094] Where T is temperature. It is the rate of change of temperature over time (unit: K / s). For density, For constant volume specific heat, and Let be the molar enthalpy of reaction and the net reaction rate of the j-th reaction, respectively. The mass conservation equation is established for each chemical component i participating in the reaction:

[0095] ;

[0096] in Let i be the mass fraction of component i. It is the rate of change of the mass fraction of the component over time (unit: ), Let be the stoichiometric coefficient of component j in the j-th reaction. Let i be the molar mass of component i. Let be the net reaction rate of the j-th reaction;

[0097] The aforementioned laws of conservation of mass and energy are common knowledge to those skilled in the art, and will not be elaborated upon in this application.

[0098] The above simulation of the oxidation reaction rate of high-olefin components under different oxygen content environments was first conducted by obtaining detailed oxidation reaction mechanism documents (containing hundreds of elementary reactions) of high-olefin components (such as 1-hexene) from literature and knowledge graphs. Then, by changing the ratio of nitrogen to oxygen in the inlet gas, different initial oxygen molar fractions were set in the model. Finally, the reaction rate of each elementary reaction j was calculated. The consumption rate of olefins and the formation rate of intermediate products at each time point are obtained by calculating the Arrhenius equation (which is common knowledge to those skilled in the art and will not be described in detail here).

[0099] The above calculation of the contribution rate of high olefin components to the overall calorific value improvement within a specific proportion range first obtains the baseline calorific value of a benchmark formulation without high olefin components (or with extremely low content) under standard operating conditions through simulation or experiment. Then, the proportion of high olefins in the formulation is gradually increased in increments of 0.5%, and the corresponding change in calorific value is recorded at each step under the same operating conditions. The difference between the two is used to calculate the calorific value increment (or contribution increment) at each step. Finally, the contribution rate is obtained by dividing the calorific value increment by the baseline calorific value. This indicator quantifies the pure gain effect of olefin components on calorific value within a specific high olefin proportion range.

[0100] The above-mentioned use of analysis of variance to identify the decrease in combustion stability of high-olefin components under fluctuating ambient temperature and humidity conditions employs a two-factor factorial experimental design of temperature and humidity, with the standard deviation of combustion pressure oscillation as the metric. As the response variable, the sum of squares, mean squares, and F-values ​​of each effect were calculated and subjected to empirical tests to confirm that the main effects and interaction effects of temperature and humidity influence the stability of the high-olefin formulation. This part follows the standard variance analysis procedure and is common knowledge, so it will not be elaborated here. Conventional variance analysis only shows that environmental fluctuations do indeed worsen stability, but it does not provide a quantitative indicator that can be directly used in subsequent formulas. This application defines the combustion stability degradation amount as: degradation amount. ;in, The standard deviation of combustion pressure oscillation under extreme operating conditions; The standard deviation of combustion pressure oscillation under standard operating conditions; standard operating conditions and extreme operating conditions (i.e. The selection of formulation F is based on the specific needs of the oil product R&D scenario, rather than general statistical practices; for example, formulation F was measured under standard operating conditions (25℃, 55%RH). =2.10 kPa; measured under extreme conditions (45℃, 90% RH) =5.30kPa, then the combustion stability attenuation amount =5.30−2.10=3.20kPa;

[0101] On the other hand, conventional analysis only analyzes calorific value or stability separately, while this method calculates the conflict deviation between positive benefits (calorific value improvement contribution rate) and negative costs (stability degradation), thereby determining whether there is a conflict in the formulation where the calorific value benefit is offset by the cost of stability deterioration. For example, if the calorific value improvement contribution rate of formulation F is 8% (contribution increment = 3.44 MJ / kg), the baseline calorific value is 43.0 MJ / kg, and the average combustion pressure is 2.8 kPa, then substituting these values ​​into the conflict assessment formula:

[0102] If the value exceeds the threshold of 0.75, it is determined to be a performance over-limit conflict, and the formula is moved to the list of incompatible formulas.

[0103] The above-mentioned method of quantifying the performance over-limit conflict caused by changes in the proportion of high olefins by calculating the conflict deviation between the contribution rate of improvement and the amount of combustion stability decline refers to the state of imbalance between calorific value gains and stability costs. That is, the contribution rate of improvement reflects the positive benefits (increased calorific value) brought by high olefin components; the amount of combustion stability decline reflects the negative costs (deterioration of stability) brought by high olefin components when the environment changes.

[0104] The conflict deviation is calculated as follows:

[0105] ,in, (Unit: MJ / kg) The incremental contribution of the high-olefin component to the calorific value is equal to the calorific value of the formulation at the current high-olefin ratio minus the calorific value of the baseline formulation (0% high-olefin mass fraction) under the same standard operating conditions. (Unit: MJ / kg) is the calorific value of the baseline formulation, used as the normalization benchmark for the calorific value dimension. (Unit: kPa) represents the amount of combustion stability degradation. , It is the standard deviation of combustion pressure oscillation measured under extreme conditions (e.g., 45°C, 90%RH). It is the standard deviation of combustion pressure oscillation measured under standard operating conditions (e.g., 25°C, 55%RH) for the same formulation (this parameter has a dual role in the formula, in the numerator, ...). Used as a benchmark value in calculating stability decay. This characterizes the degree of deterioration in absolute stability caused by environmental fluctuations; in the denominator, As a normalization benchmark, the absolute attenuation is converted into a relative attenuation rate. This makes formulations with different basic stability levels comparable; this approach, which uses relative decay rate as a normalization strategy, can effectively identify formulations that are sensitive to the operating environment, even if a formulation exhibits excellent stability under standard operating conditions. (If the stability is relatively small, but the stability is extremely sensitive to temperature and humidity fluctuations, even small environmental changes can lead to a significant decrease in stability, and the degree of conflict deviation will also be highlighted due to the large relative decrease). and These are dimensionless weighting coefficients, representing the relative importance of calorific value gains and stability costs in conflict determination, respectively. The default value for both is 0.5, which can be adjusted according to product design preferences (e.g., for high-power demand scenarios). =0.7, taken as 0.7 in scenarios requiring high stability. =0.7); The above formula uses the standard deviation of combustion pressure oscillation under standard operating conditions. As a normalization benchmark, it aims to measure the degree of relative stability degradation caused by environmental fluctuations. This method can effectively identify formulations sensitive to the operating environment: even if a formulation exhibits excellent stability under standard operating conditions, if it is extremely sensitive to temperature and humidity fluctuations, even small environmental changes will lead to a significant decrease in stability, and its conflict deviation will be highlighted due to the large relative decrease. Conversely, a formulation with a small stability margin under standard operating conditions has limited allowable degradation space, which is consistent with the actual usability of the formulation. This normalization strategy makes formulations with different basic stability levels comparable; the above and The value of satisfies α + β = 1. Its value is determined by looking up a table based on the R&D scenario. This table is pre-set by the system based on expert experience and does not require real-time calculation; typical mapping relationships are shown in Table 3 below:

[0106] Standard Equilibrium Scenario 0.5 0.5 The default value is suitable for most routine formulation development. High power priority scenarios 0.7 0.3 Suitable for benchmarking R&D of competing products with stringent upper limits on calorific value. High stability priority scenarios 0.3 0.7 Suitable for scenarios with strict regulatory requirements on combustion cycle variation rate (COV).

[0107] The mapping table is stored in the system as a configuration file. Users can directly select scenarios according to project requirements, and the system automatically calls the corresponding weight coefficients, thus transparently and repeatably converting design preferences into calculation parameters. When the deviation exceeds 0.75 (the value of 0.75 is obtained by performing ROC curve analysis on 1000 groups of labeled ineffective and effective formulations in the historical sample database, selecting the point corresponding to the maximum Youden index), the system will automatically convert the deviation to the maximum value. The value, 0.75, has been verified to have a screening accuracy of 92% at this threshold. This threshold is a reference value calculated based on a historical general sample database. In actual R&D projects, if there is a small amount of project-specific pre-experimental data (e.g., 5-10 proven effective or ineffective formulations), it is recommended to recalculate the threshold specific to the current project, use this pre-experimental data as a sample, plot the ROC curve, and reselect the point corresponding to the maximum Youden index. If no such data is available, the default value of 0.75 can be used, and it can be dynamically adjusted in subsequent verification. This means that while increasing the olefin ratio can still slightly improve the calorific value, the resulting stability degradation exceeds the acceptable range, creating a conflict where the calorific value gain is offset by the cost of stability degradation. If this conflict is quantified and identified at the small-scale stage, the formulation can be judged as having a risk of exceeding performance limits.

[0108] The above calculation of the conflict deviation degree uses the square root of the sum of squares (Euclidean norm) because the heat value gain and the stability cost are two orthogonal dimensions. The sum of squares can avoid the cancellation of positive and negative values ​​and reflect the nonlinear accumulation of conflict.

[0109] Step S40: Based on the prior logic provided by the oil product R&D knowledge graph (the prior logic originates from the known chemical mechanism relationships provided in the oil product R&D knowledge graph and is used to constrain causal search), the standardized experimental dataset is analyzed using a causal reasoning algorithm to extract the nonlinear mapping relationship between component ratio, environmental variables and performance output, and to construct a performance association path from the input end to the output end (i.e., a causal chain from input (components, environment) to output (performance)).

[0110] It should be noted that, in the embodiments of this application described above, prior logic refers to logical relationships that have been verified through historical experience, chemical theory, or previous experiments and stored in the knowledge graph, in addition to the current sample test data. These are known truths that existed before the causal reasoning algorithm began searching for patterns in the data. For example, the triple <olefin content, influence, calorific value> in the knowledge graph is a prior logic, which tells the algorithm that olefin content is a causal variable that should be considered; the triple <ambient humidity, constraint, combustion stability> tells the algorithm that ambient humidity is a candidate variable that cannot be ignored when looking for the cause of changes in combustion stability; and the triple <olefin, cause, high-temperature cracking to produce enol> provides a priori clue for an intermediate reaction path.

[0111] Performance correlation paths refer to a directed causal chain that, under the action of causal reasoning algorithms, starts from the component ratio and environmental variables at the input end, passes through several intermediate states (such as ignition delay period, flame propagation speed, and peak in-cylinder pressure), and finally reaches the performance indicators (such as calorific value, emission concentration, and stability indicators) at the output end. It is a complete and explainable logical path from cause to intermediate to result. The above types include: direct impact path: olefin content to calorific value; indirect impact path: olefin content to extended ignition delay period to decreased combustion isochoricity to decreased calorific value; interactive impact path: olefin content + ambient humidity to microemulsification to increased atomized particle size to decreased combustion efficiency.

[0112] Step S50: Compare the performance correlation path with the preset performance quality standards to identify (those at risk of exceeding performance limits under extreme conditions) out-of-limit logical nodes (the identification method for out-of-limit logical nodes is as follows: for each performance correlation path constructed in step S40 (e.g., olefin ratio to ignition delay period to flame propagation speed to combustion isochoricity to calorific value), pre-set corresponding quality standard intervals (derived from the same industry standards or experimental experience) for each intermediate variable (such as ignition delay period, flame propagation speed, combustion isochoricity) on the path, calculate the values ​​of each intermediate variable of the formulation on the path through a causal reasoning algorithm, and compare them one by one with the corresponding quality standard intervals. If a certain intermediate variable value exceeds its allowable range, the node corresponding to the variable is marked as an out-of-limit logical node. If multiple nodes exceed the limit, all of them are marked). Based on the out-of-limit logical nodes, divide the formulation into a list of suitable formulations, a list of formulations to be improved, and a list of unsuitable formulations (suitable formulations...). The list of formulas is divided into three categories: Formula List: Formulas with conflict deviation ≤ 0.75 and all measurable performance indicators meet the preset performance quality standards; Formula List to be Improved: Formulas with conflict deviation > 0.75 but all measurable performance indicators still meet the preset performance quality standards. These formulas have internal performance conflict risks but have not yet exceeded the standards and have improvement potential; Incompatible Formula List: Formulas with at least one performance indicator exceed the preset performance quality standards and are judged as invalid formulas and directly excluded. Meanwhile, the list of formulas to be improved uses discrete boundary enumeration and backtracking search to find the optimal component ratio range that meets the performance constraints (specifically, it divides and outputs the set of compatible formulas and the list of incompatible formulas based on the out-of-limit logic node; then, for formulas in the incompatible formula list that still have improvement potential (e.g., formulas with conflict deviation higher than the threshold but without catastrophic failure), it uses discrete boundary enumeration and backtracking search to find the optimal component ratio range that meets all performance constraints, generating an optimized formula path).

[0113] It should be noted that the optimized formulation path in the above embodiments of this application fully records the entire process of transforming an unsuitable formulation (i.e., a formulation with a conflict deviation greater than the threshold and a risk of performance exceeding limits) into an optimal component ratio range that satisfies all performance constraints through a series of component ratio adjustment steps.

[0114] Specifically, the optimized formulation path consists of the following three interconnected parts:

[0115] Starting point: The component ratio vector of the original unsuitable formulation (e.g., high olefins 12.5%, aromatics 8.0%, antioxidant A 1.0%, detergent dispersant X 1.0%).

[0116] Movement trajectory: Every adjustment operation from the starting point to the end point, clearly recording by how much the mass fraction of each component changed (e.g., detergent dispersant X decreased from 1.0% to 0.7%; antioxidant A and antioxidant B increased simultaneously by 0.2% to 1.2% and 0.7% respectively; high olefins decreased from 12.5% ​​to 11.5%).

[0117] Endpoint: The optimal component ratio range finally confirmed after the robustness verification of the candidate optimal component ratio range in step S56 (i.e., passing the environmental humidity disturbance test without triggering the nonlinear transition of combustion mode) (e.g., high olefins 11.0%~12.0%, aromatics 7.5%~8.5%, antioxidant A 1.1%~1.3%, detergent dispersant X 0.6%~0.8%).

[0118] The optimized formulation path does not simply provide an isolated final formulation ratio, but rather offers a traceable and reproducible adjustment route. Researchers can follow each step of the movement path to gradually adjust the original formulation to the final optimal component ratio range, which has been verified to be stable and reliable under environmental disturbances. Specifically, the optimized formulation path is a data structure that encapsulates the adjustment process, recording the starting point (the ratio vector of the original unsuitable formulation), each step of the movement (such as the action sequence of "adjusting component A from X% to Y%)", and the endpoint (a verified optimal ratio range that meets all performance constraints and is robust). Its key value lies in providing a traceable and reproducible formulation iteration process, rather than just providing the final result.

[0119] The aforementioned preset performance quality standards are a set of quantitative evaluation criteria stored in the system for the test results of small-scale fuel samples. These standards are derived from industry standards (such as the national standard GB17930-2016 "Automotive Gasoline"), internal control indicators of enterprises, and target values ​​set by R&D projects; for example, see Table 4 below (in Table 4 below, the standard deviation of combustion pressure oscillation is shown). The relationship between (unit: kPa) and combustion pressure cyclic variation rate COV (unit: %) is as follows: ;in, The average pressure inside the combustion chamber (unit: kPa). In this application, for the sake of calculation accuracy and clarity of physical meaning, all intermediate calculations use the standard deviation of combustion pressure oscillations. As a stability indicator; it is only used when finally compared with the preset performance quality standard, according to the specific provisions of the quality standard. Convert to COV or use directly Threshold judgment is performed; for the quality standards listed in Table 4, the indicator COV≤5% needs to be converted and compared, while other indicators can be directly compared using the original values.

[0120] Calorific value ≥42.8MJ / kg Standard operating conditions Enterprise internal control Cyclic variation rate of combustion pressure (COV) ≤5% All working conditions Industry-wide HC emissions from exhaust gas ≤150ppm Cold start at room temperature National VI Standard NOx emissions from exhaust gas ≤80ppm High load conditions National VI Standard Ignition success rate ≥98% Low temperature -20℃ Enterprise internal control

[0121] In the above embodiments of this application, in step S50, the system compares the performance association path constructed in step S40 with the specific numerical thresholds in the table item by item. When the predicted performance index of a certain formula under simulated extreme working conditions exceeds the allowable range, the formula is identified as a logical node with performance over-limit risk and is thus assigned to the unsuitable formula list.

[0122] Specifically, in step S40, based on the prior logic provided by the oil product R&D knowledge graph, a causal reasoning algorithm is applied to analyze the standardized experimental dataset, extracting the nonlinear mapping relationship between component ratios, environmental variables, and performance output, and constructing a performance association path from the input end to the output end, including:

[0123] Step S41: Using structural equation modeling, the component ratios and environmental variables in the standardized experimental dataset are taken as exogenous variables, the intermediate combustion state is taken as the mediating variable, and the performance indicators are taken as endogenous variables. Causal search is performed under the prior logical constraints provided by the oil product R&D knowledge graph. The prior logical constraints include the triplet of influence relationship between components and performance, the triplet of antagonistic or synergistic relationship between components, and the triplet of constraint relationship between environmental factors and performance indicators stored in the oil product R&D knowledge graph. Each triplet is associated with a confidence score and edge weight.

[0124] It should be noted that in the above embodiments of this application, step S41 uses structural equation modeling for causal search and uses the prior logic in the oil product R&D knowledge graph as a constraint, so that the causal search process does not blindly exhaust all possible variable relationships, but explores along directions supported by prior chemical mechanisms, thereby reducing the complexity of the search space and improving the physical interpretability of the discovered causal path; at the same time, the introduction of confidence scores and edge weights allows for the quantitative differentiation of the guiding force of prior knowledge from different sources and frequencies on the causal search, ensuring that high-quality prior knowledge plays a leading role in constructing performance-related paths.

[0125] Step S42: In the causal search process, regarding the coupling effect between the change in the high olefin component ratio and the environmental humidity variable, based on the triplet constraint relationship between the olefin entity and the moisture content entity stored in the oil product R&D knowledge graph, combustion state bifurcation analysis is introduced. (Combustion state bifurcation analysis is an analysis method based on nonlinear dynamics theory. It reconstructs the phase space of the combustion pressure signal and extracts quantitative indicators that characterize the topological structure of the system's dynamic characteristics, thereby identifying the critical conditions for the system to transition from one stable mode to another or a chaotic state. In this scheme, this analysis is specifically implemented by calculating the bifurcation parameters of the phase space trajectory of the combustion pressure signal under different high olefin ratio steps and environmental humidity combinations. Combustion state bifurcation analysis is used in nonlinear dynamics systems based on bifurcation theory to detect abrupt changes in system state. When control parameters (such as olefin ratio and environmental humidity) change continuously, the type of the system's stable solution (attractor) may change abruptly at a certain critical point, for example, from a stable focus to a polarity.) Bifurcation occurs when a system transitions from a limiting cycle to chaos. In this scheme, the bifurcation parameter of the geometric dispersion of the phase space trajectory of the combustion pressure signal and its first derivative with respect to the olefin ratio are calculated. The inflection point where the derivative changes from positive to negative is detected; this inflection point corresponds to the critical condition for the system to bifurcate from a stable combustion mode to an unstable combustion mode. The bifurcation parameter of the phase space trajectory of the combustion pressure signal is calculated under different high olefin ratio steps and ambient humidity combinations. (This parameter characterizes the square of the average Euclidean distance of all points in the reconstructed phase space trajectory of the combustion pressure signal relative to the geometric center of the trajectory, i.e., the dispersion of the phase space trajectory. In stable combustion (such as periodic limiting cycles), the trajectory is concentrated in a low-dimensional subspace, and the bifurcation parameter value is small; when combustion becomes unstable and enters a chaotic state, the trajectory fills a high-dimensional space, and the bifurcation parameter value increases significantly. Therefore, the bifurcation parameter can be used to quantify the degree of evolution of the system from a stable periodic state to an unstable chaotic state, and the point of sign change of its derivative corresponds to the bifurcation critical point of the dynamic system.) The critical condition for the nonlinear transition of the combustion mode caused by changes in the high olefin ratio is identified.

[0126] It should be noted that in the above embodiments of this application, step S42 introduces combustion state bifurcation analysis into the causal search process because in the complex thermodynamic system of oil combustion, the coupling of changes in the high olefin ratio and ambient humidity may cause the combustion state to suddenly change from a stable combustion mode to an unstable combustion mode. This sudden change is manifested in the data as the bifurcation of the phase space trajectory of the pressure signal. Traditional linear correlation analysis cannot capture such nonlinear bifurcation behavior. By calculating the bifurcation parameters, the critical point of the formulation that causes the discontinuous deterioration of the combustion mode can be accurately located in the small sample testing stage, which solves the problem of the search blind zone of the causal inference algorithm in nonlinear sudden change scenarios.

[0127] Combustion state bifurcation analysis, by calculating bifurcation parameters (The physical meaning of the BF value can be intuitively understood as: a measure of the average dispersion of the combustion pressure signal trajectory points around its geometric center in the d-dimensional reconstructed phase space. During stable combustion, the phase space trajectory exhibits a regular limit ring structure, with all trajectory points tightly clustered on a low-dimensional ring surface, resulting in a small BF value (unit: kPa²). When combustion becomes unstable and tends towards chaos, the trajectory points are diffusely distributed and fill the high-dimensional space, leading to a sharp increase in the BF value. Therefore, the magnitude of the BF value directly reflects the stability state of the combustion system, and the sign change of its rate of change with control parameters (such as the high olefin ratio α) (the first derivative changes from positive to negative) indicates the boundary of the qualitative change in the system's dynamic characteristics from order to disorder.) This is used to quantify the abrupt change in the topological structure of the combustion pressure signal phase space trajectory. The specific calculation process is as follows: First, the combustion pressure signal... (Sampling frequency 50kHz, duration 2s), phase space trajectory is reconstructed using time delay embedding method: ;

[0128] in, To delay the time, the first minimum point is determined using the average mutual information method. At discrete time points Instantaneous combustion pressure values ​​collected at all times (unit: kPa or Pa). It is the i-th point of the sampled time series, with a sampling frequency of 50kHz, and a typical value of ( =0.02ms), The embedding dimension is determined using the pseudo-nearest neighbor method. When the proportion of pseudo-nearest neighbors is less than 1%, the current dimension is taken, with a typical value of 5.

[0129] Then, calculate the bifurcation parameters. : ;

[0130] in, Total number of time points; The geometric center of the trajectory For the Euclidean norm, The larger the value, the more dispersed the phase space trajectory and the less stable the combustion.

[0131] In addition, the threshold of the bifurcation parameter It is along the high olefin mass fraction (0.5% increment) Calculate different Below Value, drawing Curve. The first-order derivative sign detection method is used to locate the bifurcation critical point: And satisfy ;

[0132] That is, the inflection point where the first derivative changes from positive to negative corresponds to... The value, which physically corresponds to the boundary of the transition of the combustion attractor from the stable limiting loop to the chaotic attractor;

[0133] The above uses a one-dimensional combustion pressure timing signal The time-delay embedding method is used to reconstruct the dynamic trajectory (attractor) of the system in a high-dimensional space (such as 5D), and then the square of the average Euclidean distance of all points on this trajectory relative to its geometric center is calculated. Simply put, it measures the dispersion or complexity of the combustion state in a high-dimensional space. During stable combustion, the trajectory is confined to a low-dimensional torus, and the BF value is small. When combustion becomes unstable and tends towards chaos, the trajectory fills the high-dimensional space, and the BF value increases sharply. The bifurcation parameter BF reflects the degree of discrete distribution of the trajectory points of the combustion pressure signal in the high-dimensional phase space. When the combustion process is stable and exhibits periodic characteristics, the phase space trajectory is constrained to a low-dimensional manifold, the trajectory points are concentrated, and the BF value is small. When combustion becomes unstable and tends towards chaos, the trajectory points are dispersed in the high-dimensional space, and the BF value increases sharply. Therefore, the BF curve shows an inflection point with the change of control parameters, marking the transition of combustion dynamics from order to disorder.

[0134] Step S43: Based on the critical condition of the nonlinear transition of the combustion mode and the edge weights stored in the oil product R&D knowledge graph, apply hypothetical environmental variable perturbations to each formulation sample in the standardized experimental dataset through intervention experiment simulation, collect the change in performance output, and extract the nonlinear mapping relationship between component ratio, environmental variables, and performance output (the nonlinear mapping relationship describes the non-proportional dependence between input and output variables, which is determined by the intervention experiment simulation results in the causal inference algorithm, specifically manifested as: when the input variable changes within a specific value range, the rate of change of the output variable is not constant, and at least one of the following characteristics may appear: threshold effect: exist Below the threshold, it approximates zero; above the threshold, it displays non-zero; saturation effect: Follow Increases and gradually decreases to zero; Interaction: The performance correlation path is constructed, which includes direct influence path, indirect influence path and interactive influence path. The interactive influence path includes the hindrance path of high olefin component and ambient humidity on combustion efficiency through emulsification tendency variable.

[0135] It should be noted that, in the above embodiments of this application, the intervention experiment simulation refers to: based on the existing data of each formulation sample in the standardized experimental dataset, by setting hypothetical environmental variable perturbations (such as changes in humidity or temperature values), and using an established model (such as an environmental humidity perturbation response function, describing the saturated influence relationship of humidity changes on combustion stability, expressed as...). This involves calculating the counterfactual change in performance output under the perturbation, thus observing the causal effect without conducting actual physical experiments. Nonlinear mapping relationships refer to relationships between component proportions, environmental variables, and performance output that are not simple proportional relationships, but rather involve complex dependencies such as threshold effects, interactive enhancement or weakening, and saturation. For example, the contribution of a high olefin proportion to combustion calorific value is approximately linear in the low proportion range, but exceeding a certain threshold may lead to a sharp deterioration in combustion stability, resulting in a decrease in overall benefits. Similarly, the hindering effect of ambient humidity on combustion efficiency is stronger at high olefin proportions, exhibiting a coupled nonlinear S-shaped pattern (e.g., ...). Figure 3 (As shown) or exponential input-output relationship; the above-mentioned construction of performance correlation paths including direct influence paths, indirect influence paths, and interactive influence paths is based on the causal reasoning algorithm framework. According to the mediating variable setting of the structural equation model in step S41, the causal effect is decomposed into: direct influence path, which directly points from exogenous variables (such as high olefin ratio) to endogenous performance indicators (such as calorific value) without any mediating variables; indirect influence path: exogenous variables act on performance indicators through mediating variables (such as ignition delay period, flame propagation speed); interactive influence path: two exogenous variables (such as high olefin ratio and ambient humidity) first couple (such as through the intermediate state of micro-emulsification tendency) and then jointly affect performance indicators (such as combustion efficiency); by calculating the coefficients of each path and their significance levels, the above three types of paths are extracted from the standardized experimental data and integrated into a complete performance correlation path network. Among them, the interactive influence path specifically captures the hindering effect of high olefin components and ambient humidity on combustion efficiency through emulsification tendency.

[0136] Specifically, in step S43, hypothetical environmental variable perturbations are applied to each formulation sample in the standardized experimental dataset through intervention experiment simulation, and the changes in performance output are collected, including:

[0137] Step S431: Obtain the component ratio vector, environmental variable vector, and performance output vector corresponding to each formulation sample in the standardized experimental dataset. The component ratio vector includes the mass fraction of high olefins, the mass fraction of aromatics, the mass fraction of saturated hydrocarbons, and the mass fraction of each additive. The environmental variable vector includes the temperature value, humidity value, and intake pressure value. The performance output vector includes the calorific value of combustion, the standard deviation of combustion pressure oscillation, and the concentration of exhaust gas emissions.

[0138] Step S432: For each formulation sample, based on the edge weights of the triplet constraint relationship between the olefin entity and the moisture content entity stored in the oil product R&D knowledge graph and the critical condition (the nonlinear transition of the combustion mode calculated in step S42), construct the environmental humidity perturbation response function of the formulation sample. The environmental humidity perturbation response function is used to characterize the mapping relationship between the rate of change of the standard deviation of combustion pressure oscillation with environmental humidity and the bifurcation parameter under a given high olefin mass fraction.

[0139] It should be noted that the above S432 constructs the environmental humidity perturbation response function by combining the edge weights in the knowledge graph with the identified bifurcation critical conditions. This is because the different affinity of different olefin isomers for water molecules is reflected in the knowledge graph with different edge weights. At the same time, this affinity is further manifested as the severity of combustion stability bifurcation when humidity changes under combustion conditions. Using this function, the perturbation response characteristics of the formulation under unmeasured humidity conditions can be extrapolated with only a small number of experimental samples, providing a perturbation propagation model with chemical mechanism support for intervention experiment simulation.

[0140] The environmental humidity disturbance response function Used to characterize a given high olefin mass fraction Under these conditions, the standard deviation of combustion pressure oscillation varies with ambient humidity. The response relationship to changes in relative humidity (%RH) is expressed as follows: ;

[0141] in, For the formulation at standard humidity =5 Standard deviation of combustion pressure oscillation at actual pressure (unit: kPa); The bifurcation parameters of the formulation at standard humidity, calculated in step S42; The humidity sensitivity coefficient is calculated by extracting the edge weights corresponding to the ternary group <olefins, affected by humidity, combustion stability> from the knowledge graph. The weights range from 0.2 to 1.5, with a default value of 0.8. The calculation method is as follows: ,in =0.2, =1.5 are the lower and upper bounds of the humidity sensitivity coefficient set according to prior knowledge, respectively. w(e) is the normalized edge weight of the triple <olefin, affected by humidity, combustion stability> in the knowledge graph, with a value range of [0,1]. This linear mapping ensures that λ always falls within a reasonable range, and the larger the edge weight, the higher the sensitivity coefficient. When there is no such triple, λ is taken as 0.8. The humidity response saturation rate (unit: 1 / %RH) is the rate at which the response of the combustion pressure oscillation standard deviation to humidity changes approaches saturation as humidity deviation increases. This parameter reflects the diffusion limitation effect of water molecules participating in the combustion reaction: when the humidity deviation from the baseline is small, the combustion system is sensitive to humidity changes. Follow The increase is approximately linear; as... As the concentration of water molecules increases further, the side reactions involving water molecules gradually reach diffusion saturation. The growth rate gradually decayed to zero, by fitting experimental data under at least five different humidity levels (e.g., 30%, 50%, 70%, 90%, 95% RH). The typical value is obtained as follows: =0.05 (fitted using least squares method); apply positive humidity perturbation ( ) or negative humidity disturbance ( After that, the change in the standard deviation of combustion pressure oscillation is calculated as follows: .

[0142] The above environmental humidity disturbance response function use The coefficient is used because the bifurcation parameter itself quantifies the system's fundamental instability under undisturbed humidity; the impact of humidity disturbance should be proportional to this fundamental instability. The form is because humidity affects stability by having water molecules participate in the combustion reaction, absorb heat, and alter fuel atomization characteristics. This effect gradually saturates as humidity deviates from the baseline value, conforming to the saturation law in diffusion and reaction kinetics; the absolute value in the exponent. This indicates that regardless of whether the humidity is high or low, the greater the deviation, the stronger the impact, but the direction of the impact varies. and The comparison determines (usually, increased humidity worsens stability, i.e.) > );

[0143] The above uses the humidity response function This form, because the rate at which water molecules participate in the combustion reaction is limited by diffusion and there is a saturation effect, conforms to a combination of Fick's law and the Arrhenius equation;

[0144] Based on the edge weights of the temperature-sensitive triplets stored in the oil product R&D knowledge graph, the environmental temperature perturbation response function of the formulation sample is constructed:

[0145] ;

[0146] Where T is the ambient temperature (unit: °C). The standard temperature is 25℃. The temperature sensitivity coefficient is determined by extracting the edge weights corresponding to the triple <olefins, temperature-dependent, combustion stability> from the knowledge graph. The temperature response saturation rate was determined by fitting experimental data; a positive temperature perturbation was applied ( =45℃) or negative temperature disturbance ( After reaching -20℃, calculate the change in performance output under temperature disturbance by referring to steps S4331 to S4333.

[0147] Step S433: Based on the environmental humidity disturbance response function, apply positive humidity disturbance, negative humidity disturbance, positive temperature disturbance and negative temperature disturbance to each formulation sample in sequence, calculate the change in the standard deviation of combustion pressure oscillation under each disturbance, and take the degree to which the change in the standard deviation of combustion pressure oscillation exceeds the threshold of the bifurcation parameter corresponding to the critical condition of the nonlinear transition of the combustion mode as the change in performance output.

[0148] It should be noted that the above-described embodiments of this application first obtain component proportion vectors, environmental variable vectors, and performance output vectors to provide input data sources for constructing the environmental humidity disturbance response function in step S432, clarifying the variable dimensions required for subsequent analysis, and ensuring a traceable correspondence between the component proportions, environmental conditions, and performance outputs of each formulation sample; then, based on the edge weights of the ternary group <olefins, affected by humidity, combustion stability> in the oil product R&D knowledge graph, the humidity sensitivity coefficient is calculated. Combined with the bifurcation parameters of the formulation under standard humidity calculated in step S42, and the humidity response saturation rate κ obtained by fitting experimental data under at least 5 different humidity levels, a function expression is established. This function describes the change law of the standard deviation of combustion pressure oscillation when humidity deviates from the standard value in the form of a saturation exponent, and the change amplitude The system gradually saturates as |H-H0| increases, and BF(α) is used as a coefficient to make the influence of humidity disturbance proportional to the system's inherent instability. Finally, based on the environmental humidity disturbance response function constructed in step S432, positive humidity disturbance, negative humidity disturbance, positive temperature disturbance, and negative temperature disturbance are applied to each formulation sample in sequence, and the change in the standard deviation of combustion pressure oscillation under each disturbance is calculated. Then, the change in the standard deviation of combustion pressure oscillation is compared with the bifurcation parameter threshold corresponding to the critical condition of nonlinear transition of combustion mode in step S42. The degree to which the change in the standard deviation of combustion pressure oscillation exceeds the critical change obtained by mapping the bifurcation parameter threshold is used as the change in performance output. This quantitative mapping from environmental disturbance to combustion stability deterioration provides numerical basis for the instability risk of formulation under extreme conditions for subsequent small-sample screening.

[0149] Specifically, in step S433, the change in the standard deviation of combustion pressure oscillation under each disturbance is calculated, and the degree to which the change in the standard deviation of combustion pressure oscillation exceeds the bifurcation parameter threshold corresponding to the critical condition for the nonlinear transition of the combustion mode is used as the change in performance output, including:

[0150] Step S4331: Obtain the standard deviation of the baseline combustion pressure oscillation of the formulation sample under undisturbed state, and the standard deviation of the combustion pressure oscillation of the formulation sample under a positive humidity disturbance state, and calculate the deviation of the standard deviation of the combustion pressure oscillation of the disturbance state relative to the standard deviation of the baseline combustion pressure oscillation.

[0151] Step S4332: Obtain the critical stability change corresponding to the critical condition for the nonlinear transition of the combustion mode of the formulation sample. (Unit: kPa). The method for determining the critical stability change is as follows: When plotting the bifurcation parameter BF(α) curve in step S42, simultaneously record the change in the standard deviation of combustion pressure oscillation before and after a positive humidity disturbance (from 55%RH to 90%RH) at different high olefin mass fractions α. (α); when the bifurcation parameter BF(α) reaches the bifurcation parameter threshold. When the BF value (i.e., the inflection point where the first derivative of the BF curve changes from positive to negative) is reached, the corresponding high olefin mass fraction is α*, and the α* corresponds to... (α*) is the critical stability change Δ The physical meaning of this critical stability change is: when the formulation is subjected to humidity disturbance at a high olefin ratio α, if its stability decreases by more than Δ... If this happens, the combustion system will transition from a stable attractor to a chaotic attractor, resulting in a nonlinear transition of the combustion mode.

[0152] Step S4333: The standard deviation value D of the combustion pressure oscillation obtained in step S4331 (i.e. (Unit: kPa) and the change in critical stability determined in step S4332 (Unit: kPa) Compare the values ​​and calculate the excess amount E = max(0, D-) of the deviation value exceeding the critical stability change. The excess quantity is mapped to a dimensionless change in performance output within the interval [0,1]. This mapping employs a nonlinear mapping function based on a triplet of constraints between combustion stability decay and emission degradation stored in the oil product R&D knowledge graph, resulting in an S-shaped growth relationship between the excess quantity and the change in performance output. Figure 4As shown in the figure, this embodiment uses an S-shaped nonlinear mapping function to map the excess quantity and the change in performance output in an S-shaped growth relationship, reflecting the nonlinear transition characteristics of the combustion system from stability to chaos. The mapping formula is as follows: Where E is the excess amount (unit: kPa). The excess amount corresponding to the midpoint of the S-curve is denoted as [value]. , k is the steepness coefficient of the S-curve, determined by fitting the relationship between excess quantity and measured performance degradation in historical experimental data. The default value sets E from 0.1× To 0.9× Within the range (From 0.1 to 0.9).

[0153] In the specific implementation process of the above-described embodiments of this application, the technicians found that in the actual screening of small-scale oil samples, the mapping relationship between component ratios and performance indicators often exhibits non-convex, multi-peak, and discontinuous characteristics (for example, when the olefin ratio gradually increases from 5% to 15%, the calorific value may fluctuate by first increasing, then decreasing, and then increasing again, rather than changing monotonically; there are antagonistic effects between certain additives, resulting in steeply decreasing boundaries on the performance surface). In this case, conventional gradient descent processing relies solely on the local derivative information of the current point for iteration, easily getting trapped in a locally optimal component ratio range, and failing to find the truly globally optimal range that satisfies all performance constraints. More seriously, due to limited experimental data and the presence of measurement noise, the calculated gradient direction may be distorted, causing the algorithm to oscillate or even diverge in the infeasible region, and the final output optimized formulation path may point to a practically infeasible or even worse performance formulation range.

[0154] Specifically, in step S50, the optimal component ratio range that satisfies the performance constraints is found based on the conflict deviation degree combined with discrete boundary enumeration and backtracking search, and an optimized formulation path is generated, including the following operation steps:

[0155] Step S51: Traverse the list of formulas to be improved and select formulas with a conflict deviation greater than the conflict deviation threshold (i.e., 0.75 above) and less than the second conflict deviation threshold (e.g., 1.0) as formulas to be improved;

[0156] It should be noted that the second conflict deviation threshold is greater than the aforementioned conflict deviation threshold. The method for determining this threshold is as follows: Obtain the conflict deviation distribution of all formulations in the historical sample database that meet performance standards under standard operating conditions but exceed performance limits under extreme high-temperature and high-humidity conditions. Take the upper quartile of this distribution as the second conflict deviation threshold. This threshold represents the upper limit of the conflict deviation that allows for the possible restoration of performance robustness through component fine-tuning. If historical data is insufficient, a default value of 1.0 is used. This value corresponds to a state where the weighted Euclidean distance between the relative increase in calorific value and the relative decrease in stability is close to 1, indicating a severe imbalance where benefits and costs are nearly equally weighted. When the conflict deviation of a formulation exceeds this second threshold, it means that its internal performance conflicts have exceeded the adjustable range of component fine-tuning, and it is determined to be unimprovable and directly excluded.

[0157] Step S52: Obtain the names of each component in the component proportion vector for each formulation to be improved; in the oil product R&D knowledge graph, query the influencing components that have antagonistic or synergistic relationships with the formulation to be improved based on each component name (for example, if the formulation contains aromatic Y, and there is a ternary group <detergent-dispersant X, antagonist, aromatic Y, confidence 0.85> in the graph, then according to the antagonistic direction recorded in the ternary group, limit the upper limit of the feasible proportion range of detergent-dispersant X to 70% of the current value (i.e., it must not exceed 0.7 times the original proportion), otherwise it will aggravate the performance conflict. If there is a synergistic ternary group... For groups (e.g., <antioxidant A, synergistic, antioxidant B, confidence 0.78>), it is permissible to adjust the proportions of both components simultaneously in the same direction (e.g., simultaneously increase or simultaneously decrease), with an adjustment step of 0.2% mass fraction; a set of feasible range constraints for component proportions is formed based on all influencing components (this set is a data structure, each element containing: component name, upper boundary (numerical value) of feasible proportion range, lower boundary (numerical value) of feasible proportion range, whether it participates in coupling (yes / no), and if it participates in coupling, the name of the coupled object component and the step value (e.g., step 0.2%) are recorded);

[0158] It should be noted that, in the above embodiments of this application, for example, the set of formulas F-04 is as follows:

[0159] Detergent dispersant X: Upper boundary = 0.7 × 1.0% = 0.7%, lower boundary = 0.1%, no coupling;

[0160] Antioxidant A: Upper boundary = 1.0% + 0.4% = 1.4%, Lower boundary = 1.0% - 0.4% = 0.6%, Coupled object = Antioxidant B, Step = 0.2%;

[0161] Antioxidant B: Upper boundary = 0.5% + 0.4% = 0.9%, Lower boundary = 0.5% - 0.4% = 0.1%, Coupled object = Antioxidant A, Step = 0.2%;

[0162] Aromatic hydrocarbon Y: Upper boundary = 8.0% + 1.0% = 9.0%, Lower boundary = 8.0% - 1.0% = 7.0%, no coupling;

[0163] High olefins: Upper boundary = 12.5% ​​+ 1.0% = 13.5%, Lower boundary = 12.5% ​​- 1.0% = 11.5%, no coupling (Note: the current values ​​are for example only).

[0164] Step S53: For each component in the feasible range constraint set of component proportions, generate a candidate proportion list for that component based on its upper and lower boundaries and coupling adjustment step rules; construct an initial grid point set based on all candidate proportion lists (each grid point in the initial grid point set is a complete candidate proportion list; after constructing the initial grid point set, perform mass score normalization constraint filtering on the grid point set. Specifically, for each candidate grid point in the initial grid point set, calculate the sum of the mass scores of all its components; if one or more components are not explicitly enumerated (e.g., base oil components), the mass score of the unenumerated components must first be calculated by subtracting the sum of the mass scores of the enumerated components from 100%, and it must be determined whether the mass score of the unenumerated components is within its feasible range. Remove candidate grid points where the sum of the mass scores of all components is not equal to 100%, and candidate grid points where the mass scores of the unenumerated components exceed their feasible range. After mass score normalization constraint filtering, the remaining grid points constitute the final initial grid point set of this step).

[0165] It should be noted that in the above embodiments of this application, the candidate ratio list of each component is generated based on the upper and lower boundaries of each component and the coupling adjustment step rule. Specifically, for uncoupled components (such as detergent dispersant X, aromatic Y, and high olefins), the candidate ratio list is generated in steps of 0.5% mass fraction between the lower and upper boundaries; if the step cannot divide the interval evenly, the values ​​of the two endpoints are included.

[0166] For coupled component pairs (such as antioxidant A and antioxidant B), with the current values ​​of both components as the center, within the range from their respective lower to upper boundaries, the recorded coupling adjustment step (0.2%, which is taken from the minimum resolution of small sample blending of oil products in industry standard GB / T28768-2012) is simultaneously increased or decreased simultaneously. Specifically, the offset step number k is enumerated (k=-2,-1,0,+1,+2, corresponding to offsets of -0.4%, -0.2%, 0,+0.2%,+0.4% respectively), then the candidate value of antioxidant A = antioxidant B. The current value of oxidant A is increased by k × 0.2%, and the candidate value of antioxidant B is increased by k × 0.2%. Here, it is required that the candidate values ​​do not exceed their respective upper and lower boundaries. The candidate ratio lists of all components are combined using a Cartesian product to generate a list of all possible candidate ratios. However, it is important to note that for coupled component pairs, only those combinations where the offsets relative to the current values ​​are equal (i.e., k is the same) are retained. For uncoupled components, all combinations are retained. Finally, an initial grid point set is obtained, where each grid point represents a complete candidate ratio list.

[0167] For example, the initial grid point set of formulation F-04 contains one grid point: high olefins 12.0%, aromatics 8.0%, antioxidant A 1.0%, antioxidant B 0.5%, detergent dispersant X 0.5% (assuming the detergent dispersant is taken from the 0.7% step enumeration of 0.5%).

[0168] Step S54: For each candidate ratio list in the initial grid point set, calculate the predicted value by importing the thermodynamic analysis model of the standardized experimental dataset (this predicted value is the performance index value directly related to the over-limit logic node, such as the predicted value of combustion heat value, COV predicted value, or the predicted value of exhaust gas HC emission, etc.); compare and filter the predicted values ​​according to the preset performance quality standards to obtain a feasible grid subset consisting of multiple feasible grid points.

[0169] It should be noted that the above-described embodiment of this application traverses each candidate ratio list in the initial grid point set. For each candidate ratio list, a simplified and rapid prediction mode of the thermodynamic analysis model of the standardized experimental dataset is imported (fixed ambient temperature 25℃, relative humidity 55%RH). Only the performance index values ​​directly related to the over-limit logic node under that combination are calculated. If the over-limit logic node of the formulation to be improved includes the calorific value of combustion, the calorific value of combustion is predicted. If it includes the cyclic variation rate of combustion pressure (COV), the standard deviation of combustion pressure oscillation is predicted and converted into COV. The predicted values ​​are then compared with the preset performance... The quality standards (such as Table 4) are compared one by one (for example, if the predicted calorific value is <42.8MJ / kg, or the predicted COV is >5%, or the predicted HC emission in the exhaust gas is >150ppm, then the grid point is marked as a hard violation point and deleted from the initial grid point set. After screening all grid points, the remaining grid points constitute a feasible grid subset). The feasible grid subset is obtained by screening (the feasible grid subset consists of multiple feasible grid points that are not hard violations. If the feasible grid subset is empty, the formulation to be improved is determined to have no feasible optimization direction, and the original state is maintained without generating an optimized formulation path).

[0170] Step S55: (First, calculate the Euclidean distance between each feasible grid point in the feasible grid subset and the component proportion vector of the formulation to be improved, then) select multiple starting grid points based on the Euclidean distance between each feasible grid point in the feasible grid subset and the component proportion vector of the formulation to be improved; perform traversal processing based on each starting grid point and its corresponding neighboring grid points, and analyze to obtain the candidate optimal component proportion range for each starting grid point and the corresponding movement trajectory of each starting grid point;

[0171] It should be noted that in the above embodiments of this application, the Euclidean distance (square root of the sum of squares of the differences in mass fractions of each component) between each grid point in the feasible grid subset and the original component proportion vector of the formulation to be improved is calculated, and the three points with the smallest distance are selected as the starting grid points; for each starting grid point, its neighborhood in the feasible grid subset is defined: all other grid points that differ from the current grid point by one step (0.5%) in the mass fraction of a single component and are still within the feasible grid subset; the comprehensive performance score of the current grid point is calculated, and the score formula is: comprehensive performance score = α' × (predicted calorific value of the grid point - calorific value of the baseline formulation) / calorific value of the baseline formulation - β' × (standard deviation of combustion pressure oscillation of the grid point - reference standard deviation under standard conditions) / reference standard deviation under standard conditions, wherein the calorific value and reference standard deviation of the baseline formulation are obtained from step S30; each neighborhood grid point of the current grid point is checked in turn. If a neighboring grid point has a higher overall performance score than the current grid point, then a move is made (this move is not a physical movement, but rather an operation within the discrete set of grid points in the feasible grid subset, where one component ratio combination (i.e., one grid point) is moved to another component ratio combination (another grid point). The calorific value of the aforementioned baseline formulation is the calorific value of the baseline formulation (high olefins 0%) in step S30. The formula for calculating the overall performance score is essentially a linearized approximation of the conflict deviation formula in the optimization search scenario. It retains the trade-off structure between the two core dimensions of calorific value improvement and stability decay, while discarding the square root operation in the original formula to reduce the computational complexity during discrete enumeration search. Searching for the grid point that maximizes the overall performance score is equivalent to finding the component ratio combination with the minimum conflict deviation within the feasible domain that satisfies the performance constraints. For example, suppose the feasible grid subset contains the following three grid points:

[0172] Grid point A: High olefins 12.0%, aromatics 8.0%, detergent dispersant 0.6%;

[0173] Grid point B: High olefins 11.5%, aromatics 8.0%, detergent dispersant 0.6%;

[0174] Grid point C: High olefins 12.0%, aromatics 7.5%, detergent dispersant X 0.5%;

[0175] A shift refers to moving from grid point A to grid point B (reducing the mass fraction of high olefins by 0.5%), or from grid point A to grid point C (reducing aromatics by 0.5% and detergent-dispersant X by 0.1%). Each shift corresponds to a change in the mass fraction of one or more components according to a preset step (0.5% or 0.1%).

[0176] The above movement aims to find a grid point with the highest overall performance score (i.e., a locally optimal grid point) in a feasible grid subset, starting from an initial grid point and gradually changing the component proportions.

[0177] The reason for not directly selecting the highest-scoring grid point from all grid points, but instead starting from a certain initial grid point and gradually moving forward, is that although the number of grid points in the feasible grid subset is controllable (usually tens to hundreds), directly calculating the scores of all grid points and taking the maximum value, while feasible, cannot generate a movement trajectory that gradually adjusts the original formula ratio to the optimal ratio. The optimized formula path output in step S50 is essentially a set of continuous adjustment steps, i.e., a movement trajectory. Without a movement trajectory, only an isolated final optimal interval can be output, unable to tell the developers how to gradually change the original formula ratio to the optimal ratio. The movement is recorded as follows: from which component ratio combination to which component ratio combination, and the direction of movement (e.g., clean). Dispersant X is reduced from 0.6% to 0.5%); the moved grid point is used as the new current grid point, and the above process is repeated until there are no grid points with higher overall performance scores in all neighborhoods of the current grid point. At this point, the current grid point is marked as a local optimum grid point, and all movement records from the starting grid point to this point are sequentially connected into a movement trajectory; the above search is performed on the three starting grid points respectively to obtain a maximum of three local optimum grid points and their corresponding movement trajectories; for each local optimum grid point, the region enclosed by the mass fraction of each component fluctuating by one step (0.5%) with the grid point as the center (i.e., the interval of each component taking [center value -0.5%, center value +0.5%], and intersecting with the feasible proportion interval of the component generated in step S52) is defined as a candidate optimal component proportion interval.

[0178] To ensure that combinations other than the center point within this interval also meet performance constraints, uniformity verification is performed for each candidate optimal component ratio interval: k verification points are extracted within the interval using the Latin hypercube sampling method (k is min(20, 30% of the total number of grid points in the interval)). The predicted performance index values ​​for each verification point are calculated using a simplified rapid prediction mode of the thermodynamic analysis model (fixed ambient temperature 25℃, relative humidity 55%RH). If the predicted performance index values ​​of all verification points meet the preset performance quality standards, the candidate optimal component ratio interval passes the uniformity verification; if any verification point does not meet the performance quality standards, the range of the candidate optimal component ratio interval is narrowed until all verification points meet the performance quality standards.

[0179] Specifically, in the specific processing of the above-described embodiments of this application, the one with the highest comprehensive performance score among the three starting grid points is taken as the optimization starting point. Based on this, the original component ratio vector of the formulation to be improved is taken as the path starting point. The component changes from the original component ratio vector to the optimization starting point are recorded as the first step of the movement trajectory. Each change generated in the subsequent neighborhood search movement process is added to the movement trajectory in sequence. The final generated movement trajectory completely includes all adjustment steps from the original formulation to the local optimal grid point.

[0180] Step S56: Generate the target optimal component ratio range based on the candidate optimal component ratio range by perturbing the ambient humidity; connect the target optimal component ratio range with the movement trajectory to obtain the optimized formulation path;

[0181] It should be noted that in the above embodiments of this application, generating the target optimal component ratio interval involves selecting the component ratio combination corresponding to the center point of each candidate optimal component ratio interval as the verification formula (if the center point of the interval is not within the feasible grid subset, then the grid point closest to the center within the interval is selected); then, using the environmental humidity disturbance response function, the verification formula is calculated under positive humidity disturbance ( =90%RH) and negative humidity disturbance ( Change in the standard deviation of combustion pressure oscillations at 30%RH The specific calculation method is as follows:

[0182] First, the high olefin mass fraction α is extracted from the component ratio of the validation formulation; then, it is calculated according to the formula in step S432. ,in This formula is used at standard humidity. =55%RH combustion pressure oscillation standard deviation (obtained from the experiment in step S30 or predicted by the model), BF(α) is the bifurcation parameter of the formulation calculated in step S42, λ is the humidity sensitivity coefficient extracted from the knowledge graph (read from the edge weights of the ternary <olefins, affected by humidity, combustion stability>), κ is the humidity response saturation rate (typical value 0.05); then calculate Take the absolute value; then, calculate the result... The value is compared with a pre-calibrated critical value, which is calibrated in step S42 when the bifurcation parameter BF(α) equals the bifurcation parameter threshold. When the value is 0.80 (e.g., 0.80), the corresponding change in the standard deviation of combustion pressure oscillation is recorded as the critical change. (Obtained through simulation or experimentation, for example) =1.2 kPa, specifically, the standard deviation of combustion pressure oscillation recorded simultaneously when plotting the BF(α) curve in step S42. Find BF(α) = The corresponding α is then used to extract the humidity before and after the disturbance at that α from experimental or simulation data. Difference, as );like > If the verification formula is determined to trigger a nonlinear transition in combustion mode (from a stable limit cycle to a chaotic attractor) under humidity disturbance, the candidate optimal component ratio range is deemed to lack robustness and is therefore eliminated. Next, all candidate optimal component ratio ranges that pass the robustness verification (i.e., the aforementioned target optimal component ratio ranges) are retained. For each retained target optimal component ratio range, its corresponding movement trajectory is concatenated with the range description to form a complete optimized formula path. The specific format is: from the original component ratio vector of the original formula, through the movement trajectory [movement step 1; movement step 2; ...], to the optimal component ratio range [component A: X1%~X2%, component B: Y1%~Y2%, ...];

[0183] For example, formulation F-04, starting from its original component ratio (12.5% ​​high olefins, 8.0% aromatics, 1.0% antioxidant A, 0.5% antioxidant B, and 1.0% detergent-dispersant X), underwent the following trajectory: detergent-dispersant X decreased from 1.0% to 0.7%; antioxidant A and antioxidant B simultaneously increased by 0.2% to 1.2% and 0.7% respectively; and high olefins decreased from 12.5% ​​to 11.5%, reaching the optimal component ratio range: high olefins 11.0%~12.0%, aromatics 7.5%~8.5%, antioxidant A 1.1%~1.3%, antioxidant B 0.6%~0.8%, and detergent-dispersant X 0.6%~0.8%. Verification showed that within this range, under a 90% RH disturbance, Δσ_stab = 1.1 kPa < 1.2 kPa, no nonlinear transition of combustion mode was triggered, satisfying all performance constraints.

[0184] Specifically, in step S52, a set of feasible interval constraints for component proportions is formed based on all influencing components, including the following steps:

[0185] Step S521: Extract the original component names of each component from the component ratio vector of the formulation to be improved, and obtain the set of original component names (e.g., {aromatic hydrocarbon Y, antioxidant A, detergent dispersant X, high olefin, saturated hydrocarbon}).

[0186] Step S522: Using each component name in the original component name set as the query key, retrieve all triples (including antagonistic triples and synergistic triples) with the original component name as the head or tail entity in the oil product R&D knowledge graph; take the other component in the retrieved triple whose name is different from the original component name as the influencing component, and record the relationship type (antagonistic or synergistic) and confidence score of the triple. Determine the adjustment direction information according to the order of the head and tail entities in the triple (if it is an antagonistic relationship, the head entity is the restricted party; if it is a synergistic relationship, the two are the coupled parties).

[0187] Step S523: Extract constraint parameters from the attribute constraint layer of the oil product R&D knowledge graph by the relationship type of the triples corresponding to each influencing component (constraint parameters include antagonistic constraint parameters and synergistic constraint parameters; specifically, antagonistic constraint parameters include the preset upper limit scaling factor and the minimum effective addition amount, and synergistic constraint parameters include the synchronous adjustment step value, the maximum allowable adjustment offset, the minimum effective addition amount, and the maximum allowable addition amount).

[0188] It should be noted that the extraction constraint parameters in the above-described embodiments of this application are obtained by traversing each influencing component and determining whether the relationship type of the triples corresponding to each influencing component is antagonistic. If so, the preset upper limit scaling factor (e.g., a stored value of 0.7) corresponding to the antagonistic relationship is read from the attribute constraint layer of the oil product R&D knowledge graph. At the same time, the minimum effective addition amount of the influencing component in the industry standard (e.g., the minimum effective addition amount of detergent dispersant X is 0.1%) is read from the attribute constraint layer. If not (then the relationship type of the triple is synergistic), the preset coupling adjustment step rule (including the synchronous adjustment step value (e.g., 0.2% mass fraction) and the maximum allowable offset (e.g., fluctuating ±0.4% around the current value)) corresponding to the synergistic relationship is read from the attribute constraint layer of the oil product R&D knowledge graph. At the same time, the minimum effective addition amount and the maximum allowable addition amount of each influencing component are read from the attribute constraint layer.

[0189] Step S524: Calculate the upper boundary and lower boundary of the feasible proportion interval for each influencing component based on the constraint parameters;

[0190] It should be noted that, specifically, the calculations in the above embodiments of this application are as follows: for the antagonistic influence component, the upper boundary of the feasible proportion range of the influence component is calculated as the current quality score of the influence component multiplied by a preset upper limit scaling factor; the lower boundary of the feasible proportion range is calculated as the minimum effective addition amount of the influence component; since it is an antagonistic relationship, the influence component does not participate in the coupling adjustment, and its feasible proportion range is only constrained by the upper and lower boundaries, without step coupling rules;

[0191] For synergistic influence components, since the synergistic relationship involves two components, they need to be processed in pairs. For each pair of synergistic influence components (e.g., antioxidant A and antioxidant B), calculate the upper boundary of their respective feasible proportion intervals as the current mass score plus the maximum allowable offset, and the lower boundary as the current mass score minus the maximum allowable offset. Additionally, check if the lower boundary is lower than the minimum effective addition amount; if so, use the minimum effective addition amount as the lower boundary. Check if the upper boundary is higher than the maximum allowable addition amount (if any); if so, use the maximum allowable addition amount as the upper boundary. Because it is a synergistic relationship, there will be coupling effects; therefore, the synchronous adjustment step value for this pair of influence components also needs to be recorded as a coupling rule in subsequent steps.

[0192] Step S525: Remove all names of influencing components (from step S522) from the original set of component names to obtain a set of non-influencing component names; for each component in the set of non-influencing component names, calculate the upper boundary and lower boundary of the feasible non-influencing ratio interval by combining the upper and lower floating ratios of the default adjustment range in the attribute constraint layer with the current quality score of each component.

[0193] It should be noted that the calculation in the above embodiment of this application is to read the upper and lower fluctuation ratios (e.g., ±1.0%) of the default adjustment range from the attribute constraint layer, and calculate the upper boundary of the feasible ratio range = the current quality score value × (1 + upper and lower fluctuation ratio), and the lower boundary of the feasible ratio range = the current quality score value × (1 - upper and lower fluctuation ratio).

[0194] Step S526: Summarize the upper and lower boundaries of the feasible proportion intervals of all affected components and the upper and lower boundaries of the non-affected feasible proportion intervals to obtain the set of component proportion feasible interval constraints for the formulation to be improved.

[0195] It should be noted that in the above embodiments of this application, the original component name set is extracted from the formula to be improved; antagonistic / synergistic triples are retrieved in the knowledge graph using each name as a key to obtain the influencing components and relationship types; constraint parameters are extracted from the attribute constraint layer; the upper and lower boundaries of the feasible proportion ranges of antagonistic components, synergistic components, and non-influencing components are calculated respectively, and the constraint set is obtained by summarizing them. This addresses the problem that the adjustment range setting in the background technology is arbitrary, with too narrow a range failing to escape local inferior solutions and too wide a range violating chemical constraints; the antagonistic relationship in the knowledge graph automatically limits the upper limit of the head entity proportion, and the synergistic relationship ensures synchronous adjustment in the same direction, so that the search space of subsequent discrete enumeration strictly conforms to the chemical mechanism, avoids meaningless grid points, and improves optimization efficiency.

[0196] Example 2

[0197] like Figure 5As shown, this second embodiment is based on the method for small sample screening and invalid formula exclusion in the research and development process provided in the first embodiment of the invention, and also provides a system for small sample screening and invalid formula exclusion in the research and development process, including a knowledge graph construction module 10, a data acquisition module 20, a deviation analysis module 30, a correlation analysis module 40, and an optimization module 50;

[0198] Among them, the knowledge graph construction module 10 is used to extract entities from basic data of oil research and development using natural language processing, establish the relationship between the entities based on the preset chemical knowledge database, form a structured set of triples and store it in the graph database, and construct an oil research and development knowledge graph.

[0199] The data acquisition module 20 is used to acquire sample test data in real time and preprocess the sample test data to form a standardized experimental dataset.

[0200] The deviation analysis module 30 is used to import the standardized experimental dataset into the thermodynamic analysis model. By importing the standardized experimental dataset into the thermodynamic analysis model, energy conservation equations and mass conservation equations are constructed to simulate the oxidation reaction rate of high olefin components under different oxygen content environments. The contribution rate of high olefin components to the overall combustion heat value improvement within a specific proportion range is calculated. At the same time, variance analysis is used to identify the amount of combustion stability decay caused by the high olefin components under the conditions of fluctuating ambient temperature and humidity, and the conflict deviation between the improvement contribution rate and the amount of combustion stability decay is calculated.

[0201] The association analysis module 40 is used to analyze the standardized experimental dataset based on the prior logic provided by the oil product R&D knowledge graph, and to extract the nonlinear mapping relationship between component ratio, environmental variables and performance output, and to construct a performance association path from the input end to the output end.

[0202] And optimization module 50, used to compare the performance association path with the preset performance quality standard, identify the out-of-limit logic node, divide the formula into a suitable formula list, a formula list to be improved and an unsuitable formula list according to the out-of-limit logic node, and at the same time, find the optimal component ratio range that meets the performance constraints based on the conflict deviation degree combined with discrete boundary enumeration and backtracking search of the formula list to be improved, and generate an optimized formula path.

[0203] In summary, the present invention presents a method and system for small-sample screening and invalid formula elimination in the R&D process. By integrating knowledge graph construction, controlled experimental data acquisition, thermodynamic conflict quantification, causal reasoning path construction, and discrete enumeration backtracking optimization, a complete small-sample screening and invalid formula elimination process is formed, from prior knowledge to experimental verification and then to formula optimization.

[0204] In the specific execution process, the prior logic in the knowledge graph (the triplet of influence, antagonism, synergy, and constraint, as well as their confidence and edge weights) is used as constraints. Causal search is carried out through structural equation modeling. Combustion state bifurcation analysis is introduced to identify the critical conditions for nonlinear transition of combustion mode caused by changes in the proportion of high olefins. Then, nonlinear mapping relationships are extracted by simulating the application of environmental disturbances through intervention experiments, and direct, indirect, and interactive influence paths are constructed.

[0205] Furthermore, by acquiring the component ratios, environmental variables, and performance output vectors for each formulation sample, and based on the edge weights and bifurcation critical conditions of olefin and moisture content in the knowledge graph, an environmental humidity perturbation response function is constructed. Humidity and temperature perturbations are applied to each formulation, and the change in the standard deviation of combustion pressure oscillations is calculated. The degree exceeding the bifurcation parameter threshold is used as the change in performance output. Under a small number of experimental sample conditions, the perturbation response of untested operating conditions is extrapolated, providing a counterfactual analysis capability with chemical mechanism support for causal reasoning. By combining the bifurcation parameter with the humidity sensitivity coefficient, the saturated response of high-olefin formulations to performance deterioration under humidity fluctuations is quantified, making up for the shortcomings of traditional methods that cannot predict the impact of environmental perturbations.

[0206] Furthermore, by obtaining the deviation value of the standard deviation of combustion pressure oscillation before and after the disturbance; determining the bifurcation parameter threshold based on the inflection point where the first derivative of the bifurcation parameter curve changes from positive to negative; comparing the deviation value with the bifurcation parameter threshold to calculate the excess and map it as a dimensionless change in performance output (using a nonlinear S-shaped mapping function), the measurement of the change in performance output is highly consistent with the actual combustion physics instability boundary (bifurcation critical point), rather than simply relying on empirical fixed values; through the mapping of the excess to the S-shaped function, the nonlinear transition characteristics from stability to chaos are reflected, solving the defect that traditional linear quantification cannot characterize the degree of abrupt change;

[0207] The optimization process involves selecting formulations with moderate conflict deviations and no excessive intermediate variables; generating a set of feasible component ratio constraints based on antagonistic / synergistic relationships in a knowledge graph; generating an initial grid point set through discrete enumeration; quickly predicting and filtering hard violation points using a thermodynamic model to obtain a feasible grid subset; finding locally optimal grid points within the subset through local neighborhood climbing and recording their movement trajectories to form candidate optimal intervals; verifying robustness through humidity perturbation response; and finally outputting a complete optimized formulation path including the starting point, movement trajectory, and endpoint. This addresses the shortcomings of previous optimization techniques, such as only judging unsuitable formulations as unqualified, not providing traceable optimization paths, and gradient descent easily getting trapped in local optima under non-convex multimodal mappings.

[0208] Furthermore, the original component name set is extracted from the formula to be improved; antagonistic / synergistic triples are retrieved in the knowledge graph using each name as a key to obtain the influencing components and relationship types; constraint parameters are extracted from the attribute constraint layer; the upper and lower boundaries of the feasible proportion ranges of antagonistic components, synergistic components, and non-influencing components are calculated respectively, and the constraint set is obtained by summarizing them. This addresses the problem in the optimization background technology where the adjustment range is set arbitrarily, with too narrow a range failing to escape local inferior solutions and too wide a range violating chemical constraints; the upper limit of the head entity proportion is automatically limited by the antagonistic relationship in the knowledge graph, and the synchronous adjustment in the same direction is ensured by the synergistic relationship, so that the search space of the subsequent discrete enumeration strictly conforms to the chemical mechanism, avoiding meaningless grid points and improving optimization efficiency.

[0209] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; those skilled in the art can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for small-sample screening and elimination of ineffective formulations during the research and development process, characterized in that, The following steps are included: Entities are extracted from basic data of oil product research and development using natural language processing. Relationships between these entities are established based on a pre-set chemical knowledge database, forming a structured set of triples and storing it in a graph database, thus constructing an oil product research and development knowledge graph. Real-time acquisition of sample test data, and preprocessing of the sample test data to form a standardized experimental dataset; The standardized experimental dataset is imported into a thermodynamic analysis model. Energy conservation equations and mass conservation equations are constructed using the thermodynamic analysis model with the imported standardized experimental dataset. The oxidation reaction rate of high olefin components under different oxygen content environments is simulated. The contribution rate of high olefin components to the improvement of overall combustion heat value within a specific proportion range is calculated. At the same time, variance analysis is used to identify the amount of combustion stability decay caused by the high olefin components under the conditions of fluctuating ambient temperature and humidity. The conflict deviation between the improvement contribution rate and the amount of combustion stability decay is calculated. Based on the prior logic provided by the oil product R&D knowledge graph, the standardized experimental dataset is analyzed by applying a causal reasoning algorithm to extract the nonlinear mapping relationship between component ratio, environmental variables and performance output, and to construct a performance association path from the input end to the output end. The performance-related path is compared with the preset performance quality standards to identify out-of-limit logic nodes. Based on the out-of-limit logic nodes, the formula is divided into a suitable formula list, a formula list to be improved, and an unsuitable formula list. At the same time, the formula list to be improved is used to find the optimal component ratio range that meets the performance constraints based on the conflict deviation degree, discrete boundary enumeration, and backtracking search, thereby generating an optimized formula path.

2. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 1, characterized in that, The relationships between entities include the relationship type label and the corresponding confidence score; the relationship type label includes antagonistic relationships and cooperative relationships; The standardized experimental dataset includes component ratios, environmental variables, intermediate combustion states, and performance indicators, where component ratios include mass fraction values.

3. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 2, characterized in that, Based on the prior logic provided by the oil product R&D knowledge graph, a causal reasoning algorithm is applied to analyze the standardized experimental dataset, extracting the nonlinear mapping relationship between component ratios, environmental variables, and performance outputs, and constructing a performance association path from the input to the output, including: Using structural equation modeling, the standardized experimental dataset is used as a variable to perform causal search under the prior logical constraints provided by the oil product R&D knowledge graph. In the causal search process, in view of the coupling effect between the change in the proportion of high olefin components and the environmental humidity variable, based on the triplet constraint relationship between the olefin entity and the moisture content entity stored in the oil product research and development knowledge graph, combustion state bifurcation analysis is introduced. By calculating the bifurcation parameters of the phase space trajectory of the combustion pressure signal under different high olefin proportion steps and environmental humidity combinations, the critical conditions for the nonlinear transformation of combustion mode caused by the change in high olefin proportion are identified. Based on the critical conditions for the nonlinear transition of the combustion mode and the edge weights stored in the oil product R&D knowledge graph, hypothetical environmental variable perturbations are applied to each formulation sample in the standardized experimental dataset through intervention experiment simulation. The changes in performance output are collected, and the nonlinear mapping relationship between component ratio, environmental variables and performance output is extracted to construct a performance correlation path including direct influence path, indirect influence path and interactive influence path.

4. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 3, characterized in that, The prior logic constraints include triplets of influence relationships between components and performance, triplets of antagonistic or synergistic relationships between components, and triplets of constraint relationships between environmental factors and performance indicators stored in the oil product R&D knowledge graph. Each triplet is associated with a confidence score and edge weight. The interaction influence path includes the obstacle path of high olefin components and environmental humidity on combustion efficiency through the emulsification tendency variable.

5. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 4, characterized in that, By applying hypothesized environmental variable perturbations to each formulation sample in the standardized experimental dataset through interventional experimental simulations, the changes in performance outputs are collected, including: Obtain the component proportion vector, environmental variable vector, and performance output vector corresponding to each formulation sample in the standardized experimental dataset. The component proportion vector includes the mass fraction of high olefins, the mass fraction of aromatics, the mass fraction of saturated hydrocarbons, and the mass fraction of each additive. The environmental variable vector includes the temperature value, humidity value, and intake pressure value. The performance output vector includes the calorific value of combustion, the standard deviation of combustion pressure oscillation, and the concentration of exhaust gas emissions. For each formulation sample, based on the edge weights of the triplet constraint relationship between olefin entities and moisture content entities stored in the oil product R&D knowledge graph and the critical conditions, the environmental humidity perturbation response function of the formulation sample is constructed. Based on the environmental humidity disturbance response function, a disturbance is applied to each formulation sample, and the change in the standard deviation of combustion pressure oscillation under each disturbance is calculated. The degree to which the change in the standard deviation of combustion pressure oscillation exceeds the threshold of the bifurcation parameter corresponding to the critical condition of the nonlinear transition of the combustion mode is taken as the change in performance output.

6. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 5, characterized in that, The change in the standard deviation of combustion pressure oscillation under each disturbance is calculated, and the degree to which the change in the standard deviation of combustion pressure oscillation exceeds the bifurcation parameter threshold corresponding to the critical condition for the nonlinear transition of the combustion mode is used as the change in performance output, including: Obtain the baseline combustion pressure oscillation standard deviation of the formulation sample under undisturbed state, and the perturbed combustion pressure oscillation standard deviation of the formulation sample under a positive humidity perturbation state, and calculate the deviation of the perturbed combustion pressure oscillation standard deviation relative to the baseline combustion pressure oscillation standard deviation. Obtain the critical stability change corresponding to the critical condition for the nonlinear transition of the combustion mode of the formula sample; The deviation of the standard deviation of combustion pressure oscillation is compared with the change in critical stability. The excess of the deviation exceeding the change in critical stability is calculated, and the excess is mapped to the dimensionless change in performance output within the interval [0,1].

7. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 6, characterized in that, Based on the conflict deviation degree combined with discrete boundary enumeration and backtracking search, the optimal component ratio range that satisfies the performance constraints is found, and an optimized formulation path is generated, including the following steps: The list of incompatible formulations is traversed, and formulations with a conflict deviation greater than the conflict deviation threshold but less than the second conflict deviation threshold, and which simultaneously meet the preset performance and quality standards, are selected as formulations to be improved. For each formulation to be improved, obtain the names of each component in the component proportion vector; In the knowledge graph of oil product research and development, search for the components that have antagonistic or synergistic relationships with the formulation to be improved, based on the names of each component. A set of feasible interval constraints for component proportions is formed based on all influencing components; For each component in the feasible interval constraint set of component proportions, a candidate proportion list for that component is generated based on the upper and lower boundaries of each component and the coupling adjustment step rule. The initial grid point set is constructed based on all candidate ratio lists; For each candidate ratio list in the initial grid point set, the predicted value is calculated by importing a thermodynamic analysis model from a standardized experimental dataset; Based on the predicted values, a feasible grid subset consisting of multiple feasible grid points is obtained by comparing and filtering them through preset performance and quality standards. Multiple starting grid points are selected based on the Euclidean distance between each feasible grid point in the feasible grid subset and the component proportion vector of the formulation to be improved; By traversing each starting grid point and its corresponding neighboring grid points, the candidate optimal component ratio range of each starting grid point and the corresponding movement trajectory of each starting grid point are obtained through analysis. The target optimal component ratio range is generated by perturbing the environmental humidity based on the candidate optimal component ratio range. By connecting the target optimal component ratio range with the movement trajectory, an optimized formulation path is obtained.

8. The method for small-sample screening and invalid formula elimination in the research and development process according to claim 1, characterized in that, Based on all influencing components, a set of feasible interval constraints for component proportions is formed, including the following operational steps: Extract the original component names of each component from the component ratio vector of the formulation to be improved, and obtain the set of original component names; Using each component name in the original component name set as the query key, retrieve all triples with the original component name as the head or tail entity in the oil product R&D knowledge graph; take the other component in the retrieved triple whose name is different from the original component name as the influencing component, record the relationship type and confidence score of the triple, and determine the adjustment direction information according to the order of the head and tail entities in the triple. Constraint parameters are extracted from the attribute constraint layer of the knowledge graph of oil product R&D by the relationship type of the triples corresponding to each influencing component. Calculate the upper and lower boundaries of the feasible proportion interval for each influencing component based on the constraint parameters. Remove all names of influencing components from the original set of component names to obtain a set of non-influencing component names; for each component in the set of non-influencing component names, calculate the upper boundary and lower boundary of the feasible non-influencing ratio interval by combining the upper and lower floating ratios of the default adjustment range in the attribute constraint layer with the current quality score of each component. By summarizing the upper and lower boundaries of the feasible proportion intervals for all influencing components, along with the upper and lower boundaries of the non-influencing feasible proportion intervals, we obtain the set of component proportion feasible interval constraints for the formulation to be improved.

9. The method for small-sample screening and elimination of ineffective formulations in the research and development process according to claim 8, characterized in that, The constraint parameters include antagonistic constraint parameters and cooperative constraint parameters. Specifically, antagonistic constraint parameters include a preset upper limit scaling factor and a minimum effective addition amount, while cooperative constraint parameters include a synchronous adjustment step value, a maximum allowable offset, a minimum effective addition amount, and a maximum allowable addition amount.

10. A sample screening and invalid formula elimination system in the research and development process, comprising a knowledge graph construction module, a data acquisition module, a deviation analysis module, a correlation analysis module, and an optimization module; in, The knowledge graph construction module is used to extract entities from basic data of oil research and development using natural language processing, establish the relationship between the entities based on a preset chemical knowledge database, form a structured set of triples and store them in a graph database, and construct an oil research and development knowledge graph. The data acquisition module is used to collect sample test data in real time and preprocess the sample test data to form a standardized experimental dataset. The deviation analysis module is used to import the standardized experimental dataset into the thermodynamic analysis model. By importing the standardized experimental dataset into the thermodynamic analysis model, energy conservation equations and mass conservation equations are constructed to simulate the oxidation reaction rate of high olefin components under different oxygen content environments. The contribution rate of high olefin components to the overall combustion heat value improvement within a specific proportion range is calculated. At the same time, variance analysis is used to identify the amount of combustion stability decay caused by the high olefin components under the conditions of fluctuating ambient temperature and humidity, and the conflict deviation between the improvement contribution rate and the combustion stability decay is calculated. The association analysis module is used to analyze the standardized experimental dataset based on the prior logic provided by the oil product R&D knowledge graph, and to extract the nonlinear mapping relationship between component ratio, environmental variables and performance output by applying causal reasoning algorithms, and to construct a performance association path from the input end to the output end. The system also includes an optimization module, which compares the performance-related path with preset performance quality standards, identifies out-of-limit logic nodes, and divides the formula into a suitable formula list, a formula list to be improved, and an unsuitable formula list based on the out-of-limit logic nodes. At the same time, for the formula list to be improved, the system uses conflict deviation degree combined with discrete boundary enumeration and backtracking search to find the optimal component ratio range that meets the performance constraints, and generates an optimized formula path.