Method, device, equipment, medium and product for generating molecular composition of complex material based on macroscopic properties
By constructing mechanistic and diffusion models and generating virtual datasets using measured data, the problem of low efficiency in generating molecular composition information for complex materials was solved, and efficient generation of molecular composition data was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNIV OF PETROLEUM (BEIJING)
- Filing Date
- 2026-03-23
- Publication Date
- 2026-07-03
AI Technical Summary
In existing technologies, the generation efficiency of molecular composition information for complex materials is low, and the inefficiency of the model back-reasoning process is caused by limitations in physical equipment and high-dimensional nonlinear problems.
By acquiring measured macroscopic property data and molecular composition data of complex materials, a mechanism model is constructed, a virtual dataset is generated, and a diffusion model is trained, thereby realizing the generation of molecular composition data based on macroscopic property data.
It improves the efficiency of generating molecular composition information for complex materials, avoids the problem of scarce real data, and achieves efficient generation of molecular composition data.
Smart Images

Figure CN122337375A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer-aided chemistry or molecular modeling technology, and in particular to a method, apparatus, equipment, medium and product for generating complex material molecular compositions based on macroscopic properties. Background Technology
[0002] In the energy and chemical industry, accurate characterization of the molecular composition of complex materials is fundamental to optimizing processing and improving product quality. Therefore, to enhance the processing efficiency of complex materials, precise information collection at the molecular level is necessary to ensure accurate molecular composition details.
[0003] In existing technologies, the main method for obtaining information on the molecular composition of complex materials is to use molecular characterization techniques to detect the molecular composition of complex materials, thereby obtaining information on the molecular composition of complex materials.
[0004] Because existing technologies rely on existing detection equipment to obtain molecular composition information of complex materials, the methods for obtaining this information in large-scale industrial scenarios are limited by physical equipment, resulting in the technical problem of low efficiency in generating molecular composition information of complex materials. Summary of the Invention
[0005] This application provides methods, apparatus, equipment, media, and products for generating complex material molecular compositions based on macroscopic properties, thereby achieving the technical effect of improving the efficiency of generating complex material molecular composition information.
[0006] In a first aspect, embodiments of this application provide a method for generating complex material molecular compositions based on macroscopic properties, including:
[0007] To obtain measured macroscopic property data and measured molecular composition data of complex materials;
[0008] Mechanistic models of complex materials are constructed based on measured macroscopic property data and measured molecular composition data;
[0009] A virtual dataset containing multiple sets of macroscopic property data and molecular composition data is generated based on the mechanism model;
[0010] The initial diffusion model is trained using a virtual dataset to obtain the target diffusion model;
[0011] The macroscopic property data to be processed is input into the target diffusion model to obtain the target molecular composition data corresponding to the macroscopic property data to be processed.
[0012] In one possible implementation, a mechanistic model of complex materials is constructed based on measured macroscopic property data and measured molecular composition data, including:
[0013] A molecular library of complex materials was constructed based on measured molecular composition data.
[0014] Based on molecular libraries, measured molecular composition data, and pre-defined multi-dimensional molecular distribution constraints for complex materials, the target probability density function composition scheme for complex materials is determined.
[0015] Based on measured macroscopic property data and preset mixing rules, the composition scheme of the target probability density function is verified. When the verification result indicates that the verification is successful, a mechanism model is generated based on the composition scheme of the target probability density function and the preset mixing rules.
[0016] In one possible implementation, based on a molecular library, measured molecular composition data, and preset multi-dimensional molecular distribution constraints corresponding to complex materials, the target probability density function composition scheme for the complex material is determined, including:
[0017] Based on the molecular library and measured molecular composition data, the fitting result corresponding to each candidate probability density function is determined, and the target probability density function corresponding to each molecular composition feature dimension in complex materials is determined based on the fitting result.
[0018] Based on the pre-defined multi-dimensional molecular distribution constraints corresponding to complex materials, and the target probability density function corresponding to each molecular composition feature dimension, a combination scheme of probability density functions for complex materials is constructed.
[0019] In one possible implementation, the target probability density function composition scheme is verified based on measured macroscopic property data and preset mixing rules. When the verification result indicates that the verification is successful, a mechanistic model is generated based on the target probability density function composition scheme and preset mixing rules, including:
[0020] Based on the preset combination of probability density parameters, determine the parameter values of key parameters in the composition scheme of the target probability density function;
[0021] Based on the parameter values and the target probability density function composition scheme, the molecular composition matrix corresponding to complex materials is generated.
[0022] By inputting the molecular composition matrix into the preset mixing rules, theoretical macroscopic property data corresponding to complex materials can be obtained;
[0023] The error value is obtained by comparing theoretical macroscopic property data with measured macroscopic property data of complex materials;
[0024] When the error value is lower than the preset threshold, a mechanism model is generated based on the target probability density function composition scheme and the preset mixing rules.
[0025] In one possible implementation, a virtual dataset containing multiple sets of macroscopic property data and molecular composition data is generated based on a mechanistic model, including:
[0026] Based on measured molecular composition data, the value range of key parameters corresponding to the combination scheme of target probability density function in the mechanism model is determined, and a list of value ranges consisting of multiple key parameter value ranges is obtained.
[0027] Based on the combination scheme of the target probability density function and the list of value ranges in the mechanistic model, large-scale molecular composition data is generated by sampling; the large-scale molecular composition data includes multiple sample molecular composition data.
[0028] The molecular composition matrix corresponding to each sample molecular composition data is input into the preset mixing rules to obtain the sample macroscopic property data corresponding to the sample molecular composition data;
[0029] A virtual dataset is generated based on the molecular composition data of each sample and the corresponding macroscopic property data of each sample.
[0030] In one possible implementation, the initial diffusion model is trained based on a virtual dataset to obtain the target diffusion model, including:
[0031] The training dataset and the test dataset are obtained by standardizing and partitioning the virtual dataset.
[0032] Noise is added to the training dataset to obtain the complete training dataset.
[0033] The initial diffusion model is iteratively trained based on the complete training dataset to obtain the diffusion model to be determined for the current iteration round.
[0034] The diffusion model to be determined is validated based on the test dataset. When the validation result indicates that the validation is successful, the diffusion model to be determined is determined as the target diffusion model.
[0035] Secondly, embodiments of this application provide an apparatus for generating complex material molecular compositions based on macroscopic properties, comprising:
[0036] The acquisition module is used to acquire measured macroscopic property data and measured molecular composition data of complex materials;
[0037] The first processing module is used to construct mechanistic models of complex materials based on measured macroscopic property data and measured molecular composition data;
[0038] The second processing module is used to generate a virtual dataset containing multiple sets of macroscopic property data and molecular composition data based on the mechanism model;
[0039] The third processing module is used to train the initial diffusion model based on the virtual dataset to obtain the target diffusion model;
[0040] The fourth processing module is used to input the macroscopic property data to be processed into the target diffusion model to obtain the target molecular composition data corresponding to the macroscopic property data to be processed.
[0041] In one possible implementation, the first processing module is further configured to:
[0042] A molecular library of complex materials was constructed based on measured molecular composition data.
[0043] Based on molecular libraries, measured molecular composition data, and pre-defined multi-dimensional molecular distribution constraints for complex materials, the target probability density function composition scheme for complex materials is determined.
[0044] Based on measured macroscopic property data and preset mixing rules, the composition scheme of the target probability density function is verified. When the verification result indicates that the verification is successful, a mechanism model is generated based on the composition scheme of the target probability density function and the preset mixing rules.
[0045] In one possible implementation, the first processing module is further configured to:
[0046] Based on the molecular library and measured molecular composition data, the fitting result corresponding to each candidate probability density function is determined, and the target probability density function corresponding to each molecular composition feature dimension in complex materials is determined based on the fitting result.
[0047] Based on the pre-defined multi-dimensional molecular distribution constraints corresponding to complex materials, and the target probability density function corresponding to each molecular composition feature dimension, a combination scheme of probability density functions for complex materials is constructed.
[0048] In one possible implementation, the first processing module is further configured to:
[0049] Based on the preset combination of probability density parameters, determine the parameter values of key parameters in the composition scheme of the target probability density function;
[0050] Based on the parameter values and the target probability density function composition scheme, the molecular composition matrix corresponding to complex materials is generated.
[0051] By inputting the molecular composition matrix into the preset mixing rules, theoretical macroscopic property data corresponding to complex materials can be obtained;
[0052] The error value is obtained by comparing theoretical macroscopic property data with measured macroscopic property data of complex materials;
[0053] When the error value is lower than the preset threshold, a mechanism model is generated based on the target probability density function composition scheme and the preset mixing rules.
[0054] In one possible implementation, the second processing module is further configured to:
[0055] Based on measured molecular composition data, the value range of key parameters corresponding to the combination scheme of target probability density function in the mechanism model is determined, and a list of value ranges consisting of multiple key parameter value ranges is obtained.
[0056] Based on the combination scheme of the target probability density function and the list of value ranges in the mechanistic model, large-scale molecular composition data is generated by sampling; the large-scale molecular composition data includes multiple sample molecular composition data.
[0057] The molecular composition matrix corresponding to each sample molecular composition data is input into the preset mixing rules to obtain the sample macroscopic property data corresponding to the sample molecular composition data;
[0058] A virtual dataset is generated based on the molecular composition data of each sample and the corresponding macroscopic property data of each sample.
[0059] In one possible implementation, the third processing module is further configured to:
[0060] The training dataset and the test dataset are obtained by standardizing and partitioning the virtual dataset.
[0061] Noise is added to the training dataset to obtain the complete training dataset.
[0062] The initial diffusion model is iteratively trained based on the complete training dataset to obtain the diffusion model to be determined for the current iteration round.
[0063] The diffusion model to be determined is validated based on the test dataset. When the validation result indicates that the validation is successful, the diffusion model to be determined is determined as the target diffusion model.
[0064] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0065] The memory stores instructions that the computer executes;
[0066] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect above and various possible implementations of the first aspect.
[0067] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and various possible implementations thereof.
[0068] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and various possible implementations thereof.
[0069] This application provides a method, apparatus, device, medium, and product for generating molecular composition of complex materials based on macroscopic properties. The method involves collecting actual data on complex materials to obtain measured macroscopic property data and measured molecular composition data. A mechanistic model of the complex material is constructed using the measured data, clarifying the correlation between the macroscopic property data and the molecular composition data. A virtual dataset is constructed for the complex material using the mechanistic model, resulting in a dataset containing multiple macroscopic property data and molecular composition data, avoiding the problem of scarce real data during model training and obtaining sufficient effective training data. An initial diffusion model is trained using the virtual dataset, driving the diffusion model to learn the distribution characteristics of the virtual dataset, thereby obtaining a model that can generate corresponding molecular composition data based on the input macroscopic property data, and realizing the generation of molecular composition data based on this model. This achieves the technical effect of improving the efficiency of generating molecular composition information for complex materials. Attached Figure Description
[0070] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0071] Figure 1 A flowchart illustrating the method for generating complex material molecular compositions based on macroscopic properties provided in this application. Figure 1 ;
[0072] Figure 2 A flowchart illustrating the method for generating complex material molecular compositions based on macroscopic properties provided in this application. Figure 2 ;
[0073] Figure 3 This is a schematic diagram of a combination of probability density functions provided in an embodiment of this application;
[0074] Figure 4 A flowchart illustrating the method for generating complex material molecular compositions based on macroscopic properties provided in this application. Figure 3 ;
[0075] Figure 5A schematic diagram illustrating the macroscopic properties of a complex material provided in an embodiment of this application;
[0076] Figure 6 This is a schematic diagram of the molecular distribution of the training set provided in an embodiment of this application;
[0077] Figure 7 This is a schematic diagram of the molecular distribution of the test set provided in an embodiment of this application;
[0078] Figure 8 A comparative diagram of the hydrocarbon composition of straight-run gasoline provided in this application;
[0079] Figure 9 A comparative diagram of the distillation curves of straight-run gasoline provided in this application;
[0080] Figure 10 A schematic diagram of the molecular weight distribution of straight-run gasoline provided in this application;
[0081] Figure 11 A comparative schematic diagram of the hydrocarbon composition of the blended gasoline provided in this application;
[0082] Figure 12 A comparative diagram of the distillation curves of the blended gasoline provided in this application;
[0083] Figure 13 A schematic diagram of the molecular weight distribution of the blended gasoline provided in this application;
[0084] Figure 14 A comparative schematic diagram of the hydrocarbon composition of the coking diesel provided in this application;
[0085] Figure 15 A comparative diagram of the distillation curves of the coking diesel provided in this application;
[0086] Figure 16 A schematic diagram of the molecular weight distribution of coking diesel provided in this application;
[0087] Figure 17 A comparative schematic diagram of the hydrocarbon composition of the catalytic cracking diesel provided in this application;
[0088] Figure 18 A comparative schematic diagram of distillation curves for catalytic cracked diesel provided in this application;
[0089] Figure 19 A schematic diagram of the molecular weight distribution of the catalytic cracked diesel fuel provided in this application;
[0090] Figure 20 A diagram showing a comparison of the computational speeds of the generation method provided in this application and the reconstruction method based on a global optimization algorithm;
[0091] Figure 21A schematic diagram of the device for generating complex material molecular compositions based on macroscopic properties, provided in this application;
[0092] Figure 22 A schematic diagram of the structure of the electronic device provided in this application.
[0093] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0094] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0095] First, let me explain the terms used in this application:
[0096] U-shaped Network (U-Net): refers to a deep convolutional neural network architecture used for image segmentation tasks. This network adopts a symmetrical U-shaped structure and includes an encoder for downsampling feature extraction and a decoder for upsampling to restore spatial resolution.
[0097] In existing technologies, the methods for generating molecular composition information for complex materials are either direct detection using molecular characterization techniques, or inverse deduction using mechanistic models of complex materials, specifically, inverse deduction of the molecular composition of complex materials through global optimization algorithms.
[0098] However, molecular characterization techniques rely on specific physical equipment and practical operations, thus incurring significant time and equipment costs. Mechanistic model inversion relies on initially set guesses, and in high-dimensional nonlinear problems, the inversion process can become bogged down in slow convergence of the optimal solution, reducing efficiency. Therefore, existing technologies suffer from low efficiency in generating information about the molecular composition of complex materials.
[0099] To address the aforementioned technical problems, this application proposes the following technical concept: A mechanistic model of complex materials is constructed by acquiring measured macroscopic property data and measured molecular composition data; multiple virtual datasets of macroscopic property data and molecular composition data are constructed using the mechanistic model; an initial diffusion model is trained using the virtual datasets to obtain a target diffusion model capable of generating molecular composition data based on macroscopic property data. The target diffusion model is then used to process the macroscopic property data to be processed, thereby generating the corresponding target molecular composition data. Compared with existing technologies, this application utilizes the construction of mechanistic and diffusion models to establish a high-dimensional nonlinear mapping relationship between macroscopic properties and molecular composition, thereby achieving efficient generation of molecular compositions of complex materials and improving the technical efficiency of complex material molecular composition generation.
[0100] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0101] Figure 1 A flowchart illustrating the method for generating complex material molecular compositions based on macroscopic properties provided in this application. Figure 1 ,like Figure 1 As shown, the method includes:
[0102] S101. Obtain measured macroscopic property data and measured molecular composition data of complex materials.
[0103] In this step, the types of complex materials include, but are not limited to, complex molecular mixtures such as petroleum fractions, coal liquefaction oil, waste plastic pyrolysis oil, and biomass pyrolysis oil. The measured molecular composition data and measured macroscopic property data are obtained by analyzing the complex materials using analytical instruments. For molecular composition analysis, the analytical instruments used include, but are not limited to, gas chromatography, liquid chromatography, and mass spectrometry, which can detect molecular composition. For macroscopic property analysis of complex materials, the analytical instruments that can be used include, but are not limited to, instruments capable of detecting the macroscopic properties of materials such as simulated distillation, density detection, elemental composition detection, and hydrocarbon group composition analysis.
[0104] For example, the complex material is a blended gasoline. The measured macroscopic properties and molecular composition data of this complex material are obtained as follows: A gas chromatography-flame ionization detector (GC-FID) is used to analyze the blended gasoline sample, separating and identifying 86 molecules in the C5-C12 range, including n-alkanes, isoalkanes, alkenes, and monocyclic aromatics. The mass fraction of each molecule is recorded to obtain molecular composition data. A simulated distillation apparatus is used to measure the distillation curve, a densitometer is used to measure the density at 20°C, an elemental analyzer is used to measure the C and H content, and a hydrocarbon analyzer is used to measure the proportions of alkanes, alkenes, and aromatics to obtain macroscopic property data.
[0105] S102. Construct a mechanism model for complex materials based on measured macroscopic property data and measured molecular composition data.
[0106] In this step, the mechanism model refers to the quantitative correlation model between molecular composition and macroscopic properties. Its core is to calculate the relationship between molecular composition and macroscopic properties by constraining the distribution through probability density function and combining it with mixing rules.
[0107] Alternatively, one possible approach to constructing a mechanistic model of complex materials is as follows:
[0108] S1021. Construct a molecular library of complex materials based on measured molecular composition data.
[0109] In this step, constructing a molecular library refers to using measured molecular composition data as the core basis to screen the core molecular types in the target complex material, expanding molecular diversity by combining the physicochemical properties of the material, and finally forming a structured molecular library that covers the actual molecular distribution of the material, providing a basic molecular list for subsequent probability density function fitting and molecular composition calculation.
[0110] Alternatively, the molecular library can be constructed in the following ways:
[0111] a1. Extract core molecular information from measured molecular composition data to identify key characteristics such as molecular type, carbon chain length, and functional groups.
[0112] a2. Based on the boiling point range of complex materials and industrial application scenarios, determine the carbon chain length or structural extension boundary of molecules.
[0113] a3. Based on the core molecule, expand the types of molecules by adding side chains of different lengths and adjusting the positions of functional groups to ensure coverage of possible derivative molecules in the material.
[0114] a4. All molecules are classified and organized according to the logic of molecular type-carbon chain length to form a structured molecular library.
[0115] For example, using catalytic cracked diesel as the target material, the three sets of measured molecular composition data of catalytic cracked diesel were analyzed, and five core molecules were extracted: n-alkanes (C12, C14, C16, C18), isoalkanes (C12, C14, C16, C18), monocyclic cycloalkanes (C12, C14, C16, C18), monocyclic aromatics (C12, C14, C16, C18), and bicyclic aromatics (C14, C16, C18), and the mass fraction range of each type of molecule was determined.
[0116] Based on the boiling point range of catalytic cracked diesel (220~340℃), the carbon chain length extension boundary is determined to be C10~C20, where C10 has a boiling point of approximately 174℃ and C20 has a boiling point of approximately 343℃, covering the target boiling point range.
[0117] Molecular expansion based on carbon chain length extension boundaries:
[0118] n-Alkanes: expanded from C10 to C20, with the addition of C10, C11, C13, C15, C17, C19, and C20, for a total of 11 types.
[0119] Isomerized alkanes: Based on the core molecule, 1 to 2 methyl side chains are added to each carbon chain length, resulting in 3 types of C10 to C20 after expansion, for a total of 33 types.
[0120] Cycloalkanes and aromatics: The carbon chain length is extended to C10~C20 according to the same logic, and one typical structure is retained for each carbon chain length, with 11 structures in each category.
[0121] A molecular library table was constructed using five categories of molecules—n-alkanes, isoalkanes, monocyclic cycloalkanes, monocyclic aromatics, and bicyclic aromatics—as rows and C10–C20 carbon chain lengths as columns. The table recorded the name, carbon chain length, density, boiling point, and other basic properties of each molecule. The final molecular library contained 11+33+11+11+11=77 molecules.
[0122] S1022. Based on the molecular library, measured molecular composition data, and preset multi-dimensional molecular distribution constraints corresponding to complex materials, determine the composition scheme of the target probability density function corresponding to complex materials.
[0123] In this step, determining the target probability density function composition scheme means combining the molecular type characteristics of the molecular library, the distribution law of measured molecular composition, and the preset multi-dimensional molecular distribution constraints to screen suitable probability density functions and construct a combination scheme, thereby achieving synergistic constraints on molecular distribution and ensuring that the generated molecular composition conforms to the actual material characteristics.
[0124] S1023. Based on measured macroscopic property data and preset mixing rules, verify the composition scheme of the target probability density function. When the verification result indicates that the verification is successful, generate a mechanism model based on the composition scheme of the target probability density function and preset mixing rules.
[0125] In this step, generating the mechanism model refers to using a preset mixing rule applicable to complex materials, inputting the molecular composition matrix generated by the target probability density function composition scheme into the mixing rule, calculating the theoretical macroscopic properties; comparing the theoretical values with the measured macroscopic property data, if the error is lower than the preset threshold, then the target function combination scheme and the mixing rule are solidified to generate a mechanism model that can link the relationship between molecular composition and macroscopic properties.
[0126] It should be noted that the generation of the target probability density function composition scheme and the construction of the mechanism model in this step are described below. Figure 2 Further descriptions are provided in the embodiments shown, but will not be repeated here.
[0127] S103. Generate a virtual dataset containing multiple sets of macroscopic property data and molecular composition data based on the mechanism model.
[0128] In this step, the virtual dataset is generated by using the parameter constraints of the mechanistic model to generate large-scale molecular composition data through sampling. Then, the macroscopic properties corresponding to each molecular composition data are calculated to obtain a virtual dataset with paired molecular composition and macroscopic properties.
[0129] It should be noted that the generation of the virtual dataset is as follows: Figure 4 Further descriptions are provided in the embodiments shown, but will not be repeated here.
[0130] S104. Train the initial diffusion model based on the virtual dataset to obtain the target diffusion model.
[0131] In this step, the process of training the diffusion model involves using a virtual dataset as a foundation. Data preprocessing of the virtual dataset yields a standardized sample dataset suitable for model training. This standardized sample dataset is then divided into training and testing sets. The initial diffusion model is iteratively trained using the training set, and validated using the testing set, resulting in a target diffusion model that can be used for subsequent molecular composition data generation. The purpose of using the diffusion model is to predict and progressively remove noise from a noisy state based on given macroscopic properties of the material, ultimately generating the target molecular composition matrix. This model takes the noisy molecular composition matrix and the corresponding diffusion time step as input, and uses the macroscopic property vector as a conditional embedding in the network to guide the denoising process at each step.
[0132] Alternatively, the diffusion model can be structured as an encoder-decoder structure.
[0133] For example, the detailed structure of the diffusion model is as follows: The encoder consists of four downsampling modules stacked sequentially, each containing two residual blocks. The number of feature channels increases progressively, at 64, 128, 256, and 512 respectively. The first three downsampling modules are standard convolutional blocks, while the fourth module integrates a cross-attention layer to fuse the macroscopic property vector with the feature map. After four 2x downsampling iterations, the spatial size of the input feature map decreases from 64×64 to 4×4. The decoder contains four symmetrical upsampling modules, the first of which also contains a cross-attention layer, while the subsequent three are standard upsampling residual convolutional blocks, with the number of feature channels decreasing sequentially from 512 to 64. Feature concatenation is performed between corresponding layers of the encoder and decoder via skip connections to preserve high-resolution details of molecular composition. The network finally maps the 64-channel features to a single-channel output through a 1×1 convolutional layer, predicting the noise added at the current time step. After 50 iterations of denoising, a molecular composition matrix matching the given macroscopic properties is obtained.
[0134] It should be noted that the diffusion model in this application can also be replaced by other models that can generate matrices from input vectors, such as variational autoencoder models, generative adversarial network models, etc.
[0135] Alternatively, one possible implementation for training the diffusion model is as follows:
[0136] S1041. Based on the virtual dataset, standardize and divide it to obtain the training dataset and the test dataset.
[0137] In this step, the standardization and partitioning can be performed as follows: Macroscopic property data are normalized, mapping them to the range of 0-1. The molecular composition data format remains unchanged, ensuring that the elements in the molecular composition data represent their content percentages. The dataset is randomly divided into training and test sets according to a preset ratio.
[0138] For example, a virtual dataset of 500,000 macroscopic properties and molecular compositions generated by a mechanistic model was used for model pre-training. First, the input macroscopic property data was standardized to the range of 0-1, serving as the conditional embedding vector; while the molecular composition data was transformed into a 64×64 matrix of molecular homologues, serving as the model's target output. The training data was randomly divided into training and test sets in an 8:2 ratio.
[0139] S1042. Add noise to the training dataset to obtain the complete training dataset.
[0140] In this step, the complete training dataset is obtained by setting diffusion parameters, including the total number of noise addition steps and scheduling parameters. The noise intensity at each time step is calculated based on the cosine noise scheduling function. A random time step is selected, and a matrix of noise-adding molecules is generated according to the formula. The noise-adding matrix, time steps, macroscopic property vectors, and real noise are paired to form the complete training dataset.
[0141] For example, if the total number of noise addition steps is 1000 and the number of samples in the training dataset is 400,000, the complete training dataset can be obtained by adding noise in the following ways:
[0142] b1. Calculate the noise intensity and noise removal accumulation coefficient for each time step, specifically:
[0143]
[0144]
[0145] in, The noise intensity at each time step, This is the cumulative coefficient for noise removal, where t is the current time step and s is the scheduling parameter with a value of 0.008. Using formulas 1 and 2, noise can be gradually increased until a highly noisy sample is finally generated.
[0146] b2. For each group of samples in the training dataset, calculate the noise matrix for the non-molecular composition data. Specifically, randomly select a time step t from 1-1000, generate standard Gaussian noise of the same size as the molecular composition matrix of the molecular composition data, and construct the noisy sample x based on this time step and the standard Gaussian noise. t The construction method is shown in Formula 3:
[0147]
[0148] Where x0 is the molecular composition matrix corresponding to the molecular composition data in the sample training data. It is standard Gaussian noise, and the other parameters are explained as shown in Formula 2 above.
[0149] b3. Match the noisy samples, standard Gaussian noise, and original sample training data corresponding to each group of training data to obtain the complete training dataset.
[0150] S1043. Based on the complete training dataset, perform iterative training on the initial diffusion model to obtain the diffusion model to be determined for the current iteration round.
[0151] In this step, complete training samples are input into the initial diffusion model, the prediction error is calculated using the loss function, the model parameters are adjusted using the optimizer, and after iterative training to the preset number of training steps, a diffusion model to be determined with noise prediction capability is obtained.
[0152] Alternatively, the diffusion model to be determined can be obtained as follows:
[0153] c1. Build the initial diffusion model. This diffusion model uses the U-Net framework and includes an encoder, decoder, and cross-attention layer.
[0154] c2. Set the training parameters of the model, including the optimizer, learning rate, and loss function weights.
[0155] c3. Input complete training samples in batches, and the model outputs the prediction noise corresponding to the complete training samples.
[0156] c4. Calculate the loss value of the model based on the predicted noise, and use the loss value to backpropagate and update the model parameters.
[0157] For example, the loss value is calculated as shown in Formula 4:
[0158]
[0159] Where L is the loss value, y represents the noise predicted by the model, and y represents the actual noise. and These are the weighted hyperparameters for mean squared error and cosine similarity, respectively.
[0160] c5. Repeat the iterative training until the preset number of steps are reached to obtain the diffusion model to be determined.
[0161] For example, during training, the macroscopic properties of the input can be masked with a random probability of 30% to 70%. This helps the model generate accurate molecular compositions even when some macroscopic properties are missing. An optimizer is used during model training, with an initial learning rate set to 5 × 10⁻⁶. -4 The weight decay was 0.01. A cosine annealing strategy with a 500-step warm-up period was used for learning rate scheduling to ensure stability in the early stages of training and effectively achieve convergence in the later stages. To further accelerate training and save GPU memory, automatic mixed precision technology was used throughout the training process. Finally, a pre-trained diffusion model to be determined was obtained after 500,000 training steps.
[0162] S1044. Validate the diffusion model to be determined based on the test dataset. When the validation result indicates that the validation is successful, determine the diffusion model to be determined as the target diffusion model.
[0163] In this step, the generalization ability and generation accuracy of the diffusion model to be determined are evaluated using a test dataset. The model is verified by generating molecular composition data, calculating macroscopic property data, and comparing it with the measured macroscopic property data. When the calculated error meets the standard, the current model is solidified, and the target diffusion model is obtained.
[0164] Alternatively, the target diffusion model can be determined as follows:
[0165] The macroscopic property vectors in the test set are standardized; the standardized macroscopic properties are input into the model to be determined, and the molecular composition matrix is generated through 50 steps of denoising iteration; the molecular composition matrix is input into the preset mixing rules of the value mechanism model to calculate the theoretical macroscopic properties; the original macroscopic properties of the theory and the test set are compared, the average error is calculated, and if it is lower than the preset verification threshold, the verification is passed.
[0166] S105. Input the macroscopic property data to be processed into the target diffusion model to obtain the target molecular composition data corresponding to the macroscopic property data to be processed.
[0167] In this step, the macroscopic property data of the complex material to be processed, obtained from the industrial scenario, is input into the target diffusion model to directly generate the molecular composition matrix corresponding to the macroscopic property data. Then, the molecular composition data corresponding to the current macroscopic property data to be processed is obtained by parsing the matrix elements in the molecular composition matrix.
[0168] For example, the material to be processed is a certain coking diesel. The macroscopic properties obtained by online detection are: density 0.88 g / cm³, carbon content 88 wt%, initial boiling point 220℃, and final boiling point 340℃. After standardization, the material is input into the target diffusion model to generate a molecular composition matrix. The molecular composition matrix is analyzed to obtain the corresponding molecular composition data: n-alkanes (20 wt%), isoalkanes (30 wt%), cycloalkanes (15 wt%), aromatics (35 wt%), and the proportion of molecules with different carbon chain lengths within each hydrocarbon group.
[0169] This application provides a method for generating the molecular composition of complex materials based on macroscopic properties. This method involves collecting actual data on the complex materials to obtain measured macroscopic property data and measured molecular composition data. A mechanistic model of the complex material is constructed using the measured data, clarifying the correlation between the macroscopic property data and the molecular composition data. A virtual dataset is constructed for the complex material using the mechanistic model, resulting in a dataset containing multiple macroscopic property data and molecular composition data. This avoids the problem of scarce real data during model training and obtains sufficient and effective training data. An initial diffusion model is trained using the virtual dataset, driving the diffusion model to learn the distribution characteristics of the virtual dataset. This yields a model that can generate corresponding molecular composition data based on the input macroscopic property data, and the generation of molecular composition data is achieved based on this model. This achieves the technical effect of improving the efficiency of generating molecular composition information for complex materials.
[0170] Figure 2 A flowchart illustrating the method for generating complex material molecular compositions based on macroscopic properties provided in this application. Figure 2 ,like Figure 2 As shown, the method includes:
[0171] S201. Based on the molecular library and measured molecular composition data, determine the fitting result corresponding to each candidate probability density function, and based on the fitting result, determine the target probability density function corresponding to each molecular composition feature dimension in the complex material.
[0172] In this step, the molecular composition feature dimension refers to the molecular distribution characteristics in the molecular composition data of complex materials. The method for constructing the target density function corresponding to each molecular composition feature dimension is as follows:
[0173] d1. Based on measured molecular composition data, core molecular types are screened and expanded to generate a molecular library.
[0174] d2. Based on the molecular library splitting, multiple molecular composition feature dimensions are obtained, such as hydrocarbon content dimension and carbon chain length dimension.
[0175] d3. Select the corresponding candidate probability density function for each molecular composition feature dimension, such as gamma distribution, histogram distribution, or Poisson distribution function.
[0176] d4. Use the least squares method to fit the measured data in each dimension and calculate the goodness of fit of each candidate probability density function in each molecular composition feature dimension.
[0177] d5. For each molecular composition feature dimension, select the function with the highest goodness of fit as the target probability density function corresponding to that molecular composition feature dimension.
[0178] For example, Figure 3 This is a schematic diagram of a combination of probability density functions provided in an embodiment of this application, such as... Figure 3 As shown, the expanded molecular library includes 32 molecules: n-alkanes (C5-C12), isoalkanes (C5-C12), alkenes (C5-C12), and monocyclic aromatics (C5-C12). The candidate probability density functions include histogram distribution and gamma distribution. The 32 molecules are divided into two feature dimensions: hydrocarbon group content, corresponding to the percentage of n-alkanes, isoalkanes, alkenes, and monocyclic aromatics; and carbon chain length, corresponding to the C5-C12 distribution within each hydrocarbon group.
[0179] The histogram distribution function is shown in Formula 5:
[0180]
[0181] in, It is the probability density function (PDF) value of the x-th sample point, that is, the PDF value of each molecule; This refers to the measured PDF value corresponding to the i-th sample point, i.e., the measured PDF value of the i-th molecule, which is also the decimal form of the measured mass percentage of that molecule; x refers to the discrete sample point in the feature dimension, which here refers to discrete objects such as n-alkanes, isoalkanes, alkenes, and monocyclic aromatics in the hydrocarbon content dimension, or C5, C6...C12 in the carbon chain length dimension. The purpose of this formula is to directly use the measured content percentage of each discrete sample point in the feature dimension as its probability density value, so as to achieve rapid fitting of the distribution law of discrete, categorical molecules.
[0182] The gamma distribution function is shown in Equation 6:
[0183]
[0184] Where x refers to a continuous independent variable, and in this case, it refers to the length of the carbon chain; , , This refers to the fitting parameters, where For shape factor, As a scale factor, For position parameters, These are parameters The purpose of this formula is to obtain the optimal fitting parameters for molecular characteristic dimensions with continuous changing trends, so that the gamma distribution curve is as close as possible to the measured molecular proportion data, thereby achieving accurate fitting of the continuous molecular distribution law.
[0185] For each feature dimension, a goodness-of-fit calculation is performed. This involves determining the degree of fit between the fit of each candidate probability density function for each feature dimension and the measured value for that feature dimension. The goodness-of-fit value ranges from 0 to 1. A value closer to 1 indicates a stronger explanatory power of the candidate probability density function for the measured data, and a higher degree of fit between the fitted value and the measured value. Conversely, a value closer to 0 indicates that the candidate probability density function cannot explain the fluctuations in the measured data, and the deviation between the fitted value and the measured value is extremely large.
[0186] Hydrocarbon content: The measured data were 18 wt% n-alkanes, 42 wt% isoalkanes, 10 wt% alkenes, and 30 wt% monocyclic aromatics. When fitted with a histogram distribution, the PDF values were 0.18, 0.42, 0.10, and 0.30, respectively, with a goodness-of-fit R-value. 2 =0.93; when fitting with a gamma distribution, =2.1、 =0.2, goodness of fit R 2 =0.76.
[0187] Carbon chain length dimension: Taking isoalkanes as an example, the measured proportions of C5-C12 are 5%, 8%, 12%, 15%, 20%, 18%, 12%, and 10%, respectively. When fitting with a gamma distribution, =3.5、 =2.8, goodness-of-fit R 2 =0.90; When fitting with a histogram distribution, the goodness of fit R0.90 2 =0.62.
[0188] Based on the selection method that yields the best fit, the target probability density function for the hydrocarbon content dimension is determined to be a histogram distribution, and the target probability density function for the carbon chain length dimension is determined to be a gamma distribution.
[0189] S202. Based on the preset multi-dimensional molecular distribution constraints corresponding to complex materials, and the target probability density function corresponding to each molecular composition feature dimension, construct a combination scheme of probability density functions for complex materials.
[0190] In this step, the preset dimensional molecular distribution constraints refer to the combination of constraints for each molecular composition feature dimension, which is determined based on the characteristics of complex materials and industrial needs. The probability density function combination scheme refers to the scheme of the target probability density functions corresponding to each molecular composition feature dimension. The scheme includes the combination logic between functions, priority, and constraints for each molecular composition feature dimension.
[0191] For example, the preset multi-dimensional molecular distribution constraints are: the total content of hydrocarbon groups in straight-run gasoline is 100wt%; the carbon chain length range is C5-C12; the olefin content is ≤15wt%; and the monocyclic aromatic hydrocarbon content is ≤35wt%. The priority and constraint dimensions of the target probability density function are: histogram distribution, priority 1, used to constrain the hydrocarbon content, satisfying a total of 100%, olefins ≤15%, and aromatic hydrocarbons ≤35%. Gamma distribution, priority 2, used to constrain the carbon chain length distribution within each hydrocarbon group, satisfying C5-C12. Based on the priority and constraint dimensions, the combination logic between functions is defined, and the combination scheme is solidified to obtain the probability density function combination scheme corresponding to complex materials.
[0192] S203. Based on the preset combination of probability density parameters, determine the parameter values of key parameters in the composition scheme of the target probability density function.
[0193] In this step, the parameter values of key parameters are determined as follows: based on the measured molecular composition data, a candidate range for key parameters is determined, and multiple sets of preset probability density parameter combinations corresponding to candidate parameters are generated based on this candidate range. The parameter values corresponding to each candidate parameter in the preset probability density parameter combinations are then substituted into the above probability density function combination scheme, and it is verified whether the probability density function combination scheme satisfies the preset multi-dimensional molecular distribution constraints. The parameter values corresponding to the candidate parameters that satisfy the preset multi-dimensional molecular distribution constraints and are closest to the measured molecular composition data are determined as the parameter values of the key parameters in the target probability density function composition scheme.
[0194] S204. Based on the parameter values and the target probability density function composition scheme, generate the molecular composition matrix corresponding to the complex material.
[0195] In this step, the molecular composition matrix is generated as follows: based on the logic in the combination scheme, the total content of each hydrocarbon group is calculated using parameter values; for each hydrocarbon group, the proportion of molecules with each carbon chain length is calculated using gamma distribution parameters; the absolute content of each molecule in each molecular combination data is calculated. The measured molecular composition data are organized into a matrix by molecular type and listed by carbon chain length to obtain the molecular composition matrix.
[0196] S205. Input the molecular composition matrix into the preset mixing rules to obtain the theoretical macroscopic property data corresponding to the complex material.
[0197] In this step, the theoretical macroscopic property data is generated by selecting a preset mixing rule that is compatible with the molecular composition data, based on the basic properties of each molecule in the molecular composition matrix, and then incorporating the content and basic properties of the molecules in the molecular composition matrix into the preset mixing rule to calculate the theoretical macroscopic property data corresponding to the molecular composition matrix.
[0198] For example, the basic properties of each molecule can be obtained by querying the density and boiling point at 20°C of 32 molecules in the molecular composition matrix from a hydrocarbon property database.
[0199] Preset mixing rules are determined based on fundamental properties. When the fundamental property is density, a weighted average mixing rule is selected; when the fundamental property is carbon content, a weighted average rule is selected; and when the fundamental property is distillation curve, an association rule based on boiling point distribution is selected.
[0200] The specific values corresponding to different basic properties are calculated based on preset mixing rules. For example, the density is calculated using the weighted average mixing rule: density = (n-alkanes C5: 0.9% × 0.626) + (n-alkanes C6: 1.44% × 0.659) + ... + (monocyclic aromatics C12: 3.0% × 0.891) = 0.732 g / cm³. The carbon content is calculated using the weighted average rule: carbon content = (0.9% × 83.3%) + (1.44% × 84.0%) + ... + (3.0% × 90.0%) = 86.5 wt%. The key points of the distillation curve are determined by calculating the cumulative mass fraction. The points are sorted from low to high molecular boiling point, and the temperatures corresponding to 10%, 50%, and 90% cumulative mass fractions are calculated, yielding an initial boiling point of 34℃, a 50% distillation temperature of 95℃, and a final boiling point of 178℃. The theoretical macroscopic properties obtained are: density 0.732 g / cm³, carbon content 86.5 wt%, and distillation curves (34℃, 95℃, 178℃).
[0201] S206. Compare the theoretical macroscopic property data with the measured macroscopic property data of complex materials to obtain the error value.
[0202] In this step, calculating the error value refers to determining the deviation between the theoretical macroscopic properties and the measured macroscopic properties using either relative or absolute error calculation methods. First, it's necessary to identify the macroscopic property indicators used for error value calculation, such as density, carbon content, and the critical temperature of the distillation curve. The error for each indicator is calculated, and the average of the errors for multiple indicators is then calculated to obtain the final error value.
[0203] S207. When the error value is lower than the preset threshold, generate a mechanism model based on the target probability density function composition scheme and the preset mixing rules.
[0204] In this step, the mechanism model is generated as follows: a preset threshold is determined based on the type of complex material and the application scenario; the error value is compared with the preset threshold, and when the error value is lower than the preset threshold, the target probability objective function combination scheme and preset mixing rules are saved, and the mechanism model is generated. At the same time, the applicable scope, input and output format and specific usage of the model can also be clarified.
[0205] Figure 4A flowchart illustrating the method for generating complex material molecular compositions based on macroscopic properties provided in this application. Figure 3 ,like Figure 4 As shown, the method includes:
[0206] S401. Based on the measured molecular composition data, determine the value range of the key parameters corresponding to the combination scheme of the target probability density function in the mechanism model, and obtain a list of value ranges composed of multiple key parameter value ranges.
[0207] In this step, the list of value ranges is obtained by: deducing the actual values of key parameters from the measured molecular composition data, and statistically analyzing the distribution range of the actual values of key parameters, such as minimum, maximum, and mean values; expanding the value range by ±30% of the mean, or by multiplying the minimum value by 80% and the maximum value by 120%, to obtain the value range of each key parameter; and integrating the value ranges of multiple key parameters to obtain the list of value ranges.
[0208] For example, the key parameters corresponding to the target probability density function combination scheme are the PDF value of each molecule in the histogram distribution function, and the shape factor and scale factor in the gamma distribution function. The value range of the key parameters is determined based on the measured molecular composition data of three sets of straight-run gasoline. In the histogram distribution, there are three molecules with the following contents: molecule 1, maximum content 19wt%, minimum content 17wt%, mean 18wt%; molecule 2, maximum content 11wt%, minimum content 9wt%, mean 10wt%; molecule 3, maximum content 31wt%, minimum content 29wt%, mean 30wt%. In the gamma distribution, the maximum value of the shape factor is 3.7, and the minimum value is 3.3; the maximum value of the scale factor is 3.0, and the minimum value is 2.6.
[0209] The value ranges of each key parameter, calculated by multiplying the minimum value by 80% to the maximum value by 120%, are as follows: PDF value X1 of molecule 1: 13.6~22.8wt%, PDF value X3 of molecule 2: 7.2~13.2wt%, and PDF value X3 of molecule 3: 23.2~37.2wt%; the value range of the shape factor is α: 2.64~4.44, and the value range of the scale factor is β: 2.08~3.60.
[0210] S402. Based on the combination scheme of the target probability density function in the mechanism model and the list of value ranges, large-scale molecular composition data are generated by sampling.
[0211] In this step, the large-scale molecular composition data includes multiple sample molecular composition data. The method for generating large-scale molecular composition data is to select a corresponding sampling method based on the parameter distribution characteristics, and to perform random sampling according to a list of value ranges to obtain a large number of parameter combinations. Each parameter combination is then substituted into a probability density function combination scheme to calculate the molecular composition matrix corresponding to each parameter combination. Based on the molecular composition data of each molecular composition matrix, large-scale molecular composition data is generated. The sampling methods include, but are not limited to, Monte Carlo sampling, normal distribution sampling, and uniform sampling.
[0212] For example, the sampling method selected is normal distribution sampling, and the key parameter values of the current target probability density function combination scheme are as follows: PDF value X1 of molecule 1: 13.6~22.8wt%, PDF value X2 of molecule 2: 7.2~13.2wt%, PDF value X3 of molecule 3: 23.2~37.2wt%; the value range of the shape factor is α: 2.64~4.44, and the value range of the scale factor is β: 2.08~3.60.
[0213] The normal distribution sampling method involves calculating the mean and standard deviation of each key parameter. The mean of X1 is 18.2, and the standard deviation is 1.53; the mean of X2 is 10.2, and the standard deviation is 1; the mean of X3 is 30.2, and the standard deviation is 2.3; the mean of α is 3.54, and the standard deviation is 0.3; and the mean of β is 2.84, and the standard deviation is 0.25. Based on the normal distribution sampling, 500,000 parameter combinations are generated. Each combination includes X1, X2, X3, α, and β, and the specific values of the parameters in each combination are within their respective ranges.
[0214] Each set of parameters generated by sampling is substituted into the target probability density function combination scheme to calculate the molecular composition matrix corresponding to each parameter combination. Based on the molecular composition matrix, corresponding molecular composition data is generated, thereby generating large-scale molecular composition data.
[0215] S403. Input the molecular composition matrix corresponding to each sample molecular composition data into the preset mixing rule to obtain the sample macroscopic property data corresponding to the sample molecular composition data.
[0216] In this step, the preset mixing rule refers to the preset mixing rule used in the mechanism model. It traverses each data point in the large-scale molecular composition data, substitutes the molecular composition matrix of each data point into the preset mixing rule to calculate the macroscopic properties, and records the macroscopic property data corresponding to each data point.
[0217] For example, Figure 5 This is a schematic diagram illustrating the macroscopic properties of a complex material provided in an embodiment of this application, such as... Figure 5As shown, random sampling using a normal distribution generated 500,000 sets of molecular composition data. These data were then input into the mixing rules to calculate the corresponding macroscopic properties. The molecular compositions are stored as a 64×64 matrix, organized by molecular type homologues. The calculated macroscopic properties are as follows: Figure 5 As shown in the figure. The results indicate that the density is mainly distributed between 0.7 and 0.95 g / cm³. 3 The concentration range exhibits a bimodal characteristic; the carbon content is concentrated between 85 and 90 wt%, with a peak at 86.5 wt%; the content of alkanes, cycloalkanes, and aromatics covers almost the entire range from 0 to 100 wt%. Overall, the macroscopic properties corresponding to the molecular composition of the generated product can well cover the actual property range of common diesel fuels.
[0218] S404. Based on the molecular composition data of each sample and the corresponding macroscopic property data of each sample, generate a virtual dataset.
[0219] In this step, the virtual dataset includes multiple samples, each containing macroscopic property data and molecular composition data. The molecular composition data is stored in the virtual dataset in the form of a molecular composition matrix, which is convenient to use directly when training the model based on the virtual dataset.
[0220] For example, straight-run gasoline includes 500,000 sample molecular composition data. The 64×64 molecular composition matrix corresponding to each sample is input into a preset mixing rule. The basic properties are calculated by density-weighted average, carbon content-weighted average, and boiling point correlation of distillation curve, resulting in 500,000 corresponding sample macroscopic property data. Each macroscopic property data, the corresponding molecular composition matrix of the macroscopic property data, and the molecular composition data corresponding to the molecular composition matrix are determined as a sample. The combination of multiple samples is determined as a virtual dataset.
[0221] Based on the above embodiments, this application also verifies the trained target diffusion model. Figure 6 This is a schematic diagram of the molecular distribution of the training set provided in an embodiment of this application. Figure 7 This is a schematic diagram of the molecular distribution of the test set provided in the embodiments of this application; as shown below. Figure 6 and Figure 7 As shown, in both the training and testing sets, the molecular composition distributions generated by the target diffusion model and the mechanistic model are consistent, indicating that the trained target diffusion model can accurately capture the complex relationship between macroscopic properties and molecular composition. Figure 6 and Figure 7 The vertical axis represents molecular composition information generated by the mechanistic model; the horizontal axis represents molecular composition information generated by the diffusion model. Figure 6The molecular step-by-step results for the training set. Figure 7 This shows the molecular distribution results for the test set.
[0222] Figure 8-10 This is a schematic diagram illustrating the formation result of a straight-run gasoline provided in this application. Figure 8 A comparative schematic diagram of the hydrocarbon composition of straight-run gasoline provided in this application is shown below. Figure 8 As shown, the hydrocarbon group composition corresponding to the molecular composition data generated based on the diffusion model is consistent with the hydrocarbon group composition corresponding to the molecular composition data in the experimental data. Figure 9 A comparative diagram of the distillation curves of straight-run gasoline provided in this application is shown below. Figure 9 As shown, the boiling point distribution in the distillation curves corresponding to the molecular composition data generated based on the diffusion model is consistent with the boiling point distribution in the experimental data distillation curves. Furthermore, according to... Figure 8 and Figure 9 It can be determined that the gasoline sample used for testing of this straight-run gasoline has a high content of n-alkanes and isoalkanes, does not contain olefins, and its boiling point is mainly concentrated in the range of 300K~500K. Figure 10 This is a schematic diagram of the molecular weight distribution of straight-run gasoline provided in this application, as shown below. Figure 10 As shown, the molecular weight of this straight-run gasoline is mostly distributed between 50 and 200.
[0223] Figure 11-13 This is a schematic diagram illustrating the generation result of blended gasoline provided in this application. Figure 11 A comparative schematic diagram of the hydrocarbon composition of the blended gasoline provided in this application is shown below. Figure 11 As shown, the hydrocarbon group composition corresponding to the molecular composition data generated based on the diffusion model is consistent with the hydrocarbon group composition corresponding to the molecular composition data in the experimental data. Figure 12 A comparative diagram of the distillation curves of the blended gasoline provided in this application, such as... Figure 12 As shown, the boiling point distribution in the distillation curves corresponding to the molecular composition data generated based on the diffusion model is consistent with the boiling point distribution in the experimental data distillation curves. Furthermore, according to... Figure 11 and Figure 12 It can be determined that the gasoline sample used for testing of this blended gasoline has a high content of isoalkanes, olefins and aromatics, and a low content of n-alkanes, with its boiling point mainly concentrated in the range of 300K to 500K. Figure 13 This is a schematic diagram of the molecular weight distribution of the blended gasoline provided in this application, as shown below. Figure 13 As shown, the molecular weight of this blended gasoline is mostly distributed between 50 and 200.
[0224] Figure 14-16 This is a schematic diagram illustrating the formation result of coking diesel fuel provided in this application. Figure 14 A comparative schematic diagram of the hydrocarbon composition of the coking diesel provided in this application is shown below. Figure 14As shown, the hydrocarbon group composition corresponding to the molecular composition data generated based on the diffusion model is consistent with the hydrocarbon group composition corresponding to the molecular composition data in the experimental data. Figure 15 This is a comparative diagram of the distillation curves of the coking diesel provided in this application, as shown in the figure. Figure 15 As shown, the boiling point distribution in the distillation curves corresponding to the molecular composition data generated based on the diffusion model is consistent with the boiling point distribution in the experimental data distillation curves. Furthermore, according to... Figure 14 and Figure 15 It can be determined that the gasoline sample used for testing this coking diesel has a high content of alkanes, monocyclic aromatics and bicyclic aromatics, and its boiling point is mainly concentrated in the range of 500K~650K. Figure 16 This is a schematic diagram of the molecular weight distribution of coking diesel provided in this application, as shown below. Figure 16 As shown, the molecular weight of this coking diesel oil is mostly distributed between 100 and 400.
[0225] Figure 17-19 This is a schematic diagram illustrating the formation result of catalytic cracking diesel fuel provided in this application. Figure 17 A comparative schematic diagram of the hydrocarbon composition of the catalytic cracking diesel provided in this application is shown below. Figure 17 As shown, the hydrocarbon group composition corresponding to the molecular composition data generated based on the diffusion model is consistent with the hydrocarbon group composition corresponding to the molecular composition data in the experimental data. Figure 18 A comparative schematic diagram of the distillation curves of the catalytic cracked diesel provided in this application, as shown below. Figure 18 As shown, the boiling point distribution in the distillation curves corresponding to the molecular composition data generated based on the diffusion model is consistent with the boiling point distribution in the experimental data distillation curves. Figure 19 This is a schematic diagram of the molecular weight distribution of the catalytic cracked diesel fuel provided in this application. Figure 17 , Figure 18 and Figure 19 It can be determined that this catalytic cracked diesel oil is similar to [the original] in terms of group composition, distillation profile, and molecular weight distribution. Figures 14-16 The coking diesel shown is similar.
[0226] It should be noted that, in order to verify the computational efficiency of the diffusion model trained in this application, this embodiment compares the computational speed of the two methods, the diffusion model-based generation method and the global optimization algorithm-based reconstruction method, to verify the conclusion that the method in this application can improve the efficiency of molecular composition generation. Figure 20 A comparative diagram of the computational speeds of the generation method and the reconstruction method based on the global optimization algorithm provided in this application is shown below. Figure 20 As shown, the two methods perform similarly in predicting macroscopic properties, but the computation time of the molecular reconstruction algorithm is much longer than that of the diffusion model, about 110 times longer.
[0227] Figure 21A schematic diagram of the device for generating complex material molecular compositions based on macroscopic properties provided in this application is shown below. Figure 21 As shown, the apparatus for generating complex material molecular compositions based on macroscopic properties provided in this embodiment includes:
[0228] The acquisition module 2101 is used to acquire measured macroscopic property data and measured molecular composition data of complex materials.
[0229] The first processing module 2102 is used to construct a mechanism model of complex materials based on measured macroscopic property data and measured molecular composition data.
[0230] The second processing module 2103 is used to generate a virtual dataset containing multiple sets of macroscopic property data and molecular composition data based on the mechanism model.
[0231] The third processing module 2104 is used to train the initial diffusion model based on the virtual dataset to obtain the target diffusion model.
[0232] The fourth processing module 2105 is used to input the macroscopic property data to be processed into the target diffusion model to obtain the target molecular composition data corresponding to the macroscopic property data to be processed.
[0233] Optionally, in one possible implementation, the first processing module 2102 is further configured to:
[0234] A molecular library of complex materials is constructed based on measured molecular composition data.
[0235] Based on molecular libraries, measured molecular composition data, and pre-defined multi-dimensional molecular distribution constraints for complex materials, the composition scheme of the target probability density function for complex materials is determined.
[0236] Based on measured macroscopic property data and preset mixing rules, the composition scheme of the target probability density function is verified. When the verification result indicates that the verification is successful, a mechanism model is generated based on the composition scheme of the target probability density function and the preset mixing rules.
[0237] Optionally, in one possible implementation, the first processing module 2102 is further configured to:
[0238] Based on the molecular library and measured molecular composition data, the fitting result corresponding to each candidate probability density function is determined, and the target probability density function corresponding to each molecular composition feature dimension in complex materials is determined based on the fitting result.
[0239] Based on the pre-defined multi-dimensional molecular distribution constraints corresponding to complex materials, and the target probability density function corresponding to each molecular composition feature dimension, a combination scheme of probability density functions for complex materials is constructed.
[0240] Optionally, in one possible implementation, the first processing module 2102 is further configured to:
[0241] Based on the preset combination of probability density parameters, determine the parameter values of key parameters in the composition scheme of the target probability density function.
[0242] Based on the parameter values and the target probability density function composition scheme, the molecular composition matrix corresponding to the complex material is generated.
[0243] By inputting the molecular composition matrix into the preset mixing rules, theoretical macroscopic property data corresponding to complex materials are obtained.
[0244] The error value is obtained by comparing theoretical macroscopic property data with measured macroscopic property data of complex materials.
[0245] When the error value is lower than the preset threshold, a mechanism model is generated based on the target probability density function composition scheme and the preset mixing rules.
[0246] Optionally, in one possible implementation, the second processing module 2103 is further configured to:
[0247] Based on measured molecular composition data, the value ranges of key parameters corresponding to the target probability density function combination scheme in the mechanism model are determined, resulting in a list of value ranges composed of multiple key parameter value ranges.
[0248] Based on the combination scheme of the target probability density function and the list of value ranges in the mechanistic model, large-scale molecular composition data is generated by sampling; the large-scale molecular composition data includes multiple sample molecular composition data.
[0249] The molecular composition matrix corresponding to each sample molecular composition data is input into the preset mixing rules to obtain the sample macroscopic property data corresponding to the sample molecular composition data.
[0250] A virtual dataset is generated based on the molecular composition data of each sample and the corresponding macroscopic property data of each sample.
[0251] Alternatively, in one possible implementation, the third processing module 2104 is further configured to:
[0252] The training dataset and the test dataset are obtained by standardizing and partitioning the virtual dataset.
[0253] Noise is added to the training dataset to obtain the complete training dataset.
[0254] The initial diffusion model is iteratively trained based on the complete training dataset to obtain the diffusion model to be determined for the current iteration round.
[0255] The diffusion model to be determined is validated based on the test dataset. When the validation result indicates that the validation is successful, the diffusion model to be determined is determined as the target diffusion model.
[0256] The apparatus provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0257] Figure 22 A schematic diagram of the structure of the electronic device provided in this application. Figure 22 As shown, the electronic device provided in this embodiment includes at least one processor 2201 and a memory 2202. Optionally, the device further includes a communication component 2203. The processor 2201, memory 2202, and communication component 2203 are connected via a bus 2204.
[0258] In the specific implementation process, at least one processor 2201 executes computer execution instructions stored in memory 2202, causing at least one processor 2201 to execute the above-mentioned method or approach for generating complex material molecular composition based on macroscopic properties.
[0259] The specific implementation process of processor 2201 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0260] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0261] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0262] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0263] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0264] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0265] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0266] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0267] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0268] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0269] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0270] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0271] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0272] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for generating complex material molecular compositions based on macroscopic properties, characterized in that, include: To obtain measured macroscopic property data and measured molecular composition data of complex materials; A mechanistic model of the complex material is constructed based on the measured macroscopic property data and the measured molecular composition data. Based on the aforementioned mechanism model, a virtual dataset containing multiple sets of macroscopic property data and molecular composition data is generated; The initial diffusion model is trained based on the virtual dataset to obtain the target diffusion model; The macroscopic property data to be processed is input into the target diffusion model to obtain the target molecular composition data corresponding to the macroscopic property data to be processed.
2. The method according to claim 1, characterized in that, The construction of the mechanism model for the complex material based on the measured macroscopic property data and the measured molecular composition data includes: A molecular library of the complex material is constructed based on the measured molecular composition data; Based on the molecular library, the measured molecular composition data, and the preset multi-dimensional molecular distribution constraints corresponding to the complex material, the target probability density function composition scheme corresponding to the complex material is determined. Based on the measured macroscopic property data and the preset mixing rules, the composition scheme of the target probability density function is verified. When the verification result indicates that the verification is successful, the mechanism model is generated based on the composition scheme of the target probability density function and the preset mixing rules.
3. The method according to claim 2, characterized in that, The step of determining the target probability density function composition scheme for the complex material based on the molecular library, the measured molecular composition data, and the preset multi-dimensional molecular distribution constraints corresponding to the complex material includes: Based on the molecular library and the measured molecular composition data, the fitting result corresponding to each candidate probability density function is determined, and the target probability density function corresponding to each molecular composition feature dimension in the complex material is determined based on the fitting result. Based on the preset multi-dimensional molecular distribution constraints corresponding to the complex material, and the target probability density function corresponding to each molecular composition feature dimension, a combination scheme of probability density functions corresponding to the complex material is constructed.
4. The method according to claim 2, characterized in that, The step involves verifying the target probability density function composition scheme based on the measured macroscopic property data and preset mixing rules. When the verification result indicates that the verification is successful, the mechanistic model is generated based on the target probability density function composition scheme and the preset mixing rules, including: Based on the preset combination of probability density parameters, the parameter values of key parameters in the composition scheme of the target probability density function are determined; Based on the parameter values and the target probability density function composition scheme, the molecular composition matrix corresponding to the complex material is generated; The molecular composition matrix is input into a preset mixing rule to obtain the theoretical macroscopic property data corresponding to the complex material; The error value is obtained by comparing the theoretical macroscopic property data with the measured macroscopic property data of the complex material. When the error value is lower than a preset threshold, the mechanism model is generated based on the target probability density function composition scheme and the preset mixing rule.
5. The method according to claim 1, characterized in that, The virtual dataset generated based on the aforementioned mechanism model, containing multiple sets of macroscopic property data and molecular composition data, includes: Based on the measured molecular composition data, the value range of key parameters corresponding to the target probability density function combination scheme in the mechanism model is determined, and a list of value ranges consisting of multiple key parameter value ranges is obtained. Based on the target probability density function combination scheme in the mechanistic model and the list of value ranges, large-scale molecular composition data is generated by sampling; wherein, the large-scale molecular composition data includes multiple sample molecular composition data; The molecular composition matrix corresponding to each sample molecular composition data is input into a preset mixing rule to obtain the sample macroscopic property data corresponding to the sample molecular composition data; The virtual dataset is generated based on the molecular composition data of each sample and the macroscopic property data of the sample corresponding to each molecular composition data.
6. The method according to claim 1, characterized in that, The step of training the initial diffusion model based on the virtual dataset to obtain the target diffusion model includes: The virtual dataset is standardized and divided to obtain training and test datasets. Noise is added to the training dataset to obtain the complete training dataset. Based on the complete training dataset, the initial diffusion model is iteratively trained to obtain the diffusion model to be determined for the current iteration round. The diffusion model to be determined is validated based on the test dataset. When the validation result indicates that the validation is successful, the diffusion model to be determined is determined as the target diffusion model.
7. A device for generating complex material molecular compositions based on macroscopic properties, characterized in that, include: The acquisition module is used to acquire measured macroscopic property data and measured molecular composition data of complex materials; The first processing module is used to construct a mechanism model of the complex material based on the measured macroscopic property data and the measured molecular composition data; The second processing module is used to generate a virtual dataset containing multiple sets of macroscopic property data and molecular composition data based on the mechanism model; The third processing module is used to train the initial diffusion model based on the virtual dataset to obtain the target diffusion model; The fourth processing module is used to input the macroscopic property data to be processed into the target diffusion model to obtain the target molecular composition data corresponding to the macroscopic property data to be processed.
8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.