Atlas-based molecular composition recognition and physical property prediction method and system
By combining infrared spectroscopy and gas chromatography detection, and utilizing a dual spectral database and an integrated learning neural network, the problems of low accuracy in identifying the molecular composition of gasoline and inaccurate prediction of its physical properties were solved, achieving high-precision molecular composition identification and physical property prediction.
Patent Information
- Application Number
- CN202511621249.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-11-07
AI Technical Summary
Existing technologies lack precision in identifying the molecular composition of gasoline, and single detection methods have significant limitations and inaccurate property predictions, making it difficult to meet the practical needs for accurate prediction.
By combining infrared spectroscopy and gas chromatography, and utilizing a dual spectral database for hierarchical similarity comparison, a molecular property prediction plugin is constructed through integrated learning and neural networks to achieve molecular composition identification and property prediction.
It improves the accuracy of gasoline molecular composition identification and the reliability of property prediction. By obtaining complementary information through dual detection methods, it achieves high-precision molecular composition identification and property prediction.
Smart Images

Figure CN121068532B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of materials analysis and testing technology, and in particular to a method and system for molecular composition identification and property prediction based on spectra. Background Technology
[0002] As an important petroleum product, the accurate identification of gasoline's molecular composition and the prediction of its physical properties are of great significance for oil quality control, process optimization, and product development. Gasoline contains a variety of complex components such as saturated hydrocarbons, aromatics, and olefins, and its molecular composition directly affects key performance indicators such as octane number, volatility, and stability.
[0003] Currently, gasoline molecular composition identification mainly employs single detection methods such as infrared spectroscopy or gas chromatography. Infrared spectroscopy can rapidly acquire molecular vibrational information, but its distinguishing ability for structurally similar compounds is limited, and spectral overlap is prone to occur, leading to low identification accuracy. While gas chromatography offers good separation, it has limitations in detecting trace components in complex matrices, and when used alone, it cannot comprehensively reflect the overall molecular composition characteristics of gasoline. Regarding property prediction, existing technologies typically rely on empirical formulas or simple linear models, lacking in-depth analysis of the relationship between molecular composition and properties. Due to the inaccurate identification of gasoline molecular composition, the reliability of subsequent property predictions, such as boiling point, critical parameters, and thermodynamic properties, is poor, failing to meet the practical needs for accurate prediction. Summary of the Invention
[0004] This invention addresses the technical problems of low accuracy in identifying the molecular composition of gasoline, the limitations of single detection methods, and inaccurate prediction of physical properties in existing technologies by providing a method and system for molecular composition identification and property prediction based on spectra.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0006] In a first aspect, the present invention provides a method for molecular composition identification and property prediction based on spectra, comprising: performing infrared spectroscopy and gas chromatography detection on the gasoline to be tested to obtain infrared spectral data and gas chromatography data; using a dual spectral database, performing hierarchical similarity comparison on the infrared spectral data, and identifying a first predicted molecular composition based on the similar sample spectral dataset; identifying a second predicted molecular composition based on the gas chromatography data, and linearly summing the first and second predicted molecular compositions to obtain an overall predicted molecular composition; constructing a molecular property prediction plugin based on ensemble learning and neural networks, analyzing the overall predicted molecular composition to obtain predicted molecular attribute information, and using the overall predicted molecular composition and predicted molecular attribute information as the molecular detection result of the gasoline to be tested.
[0007] Secondly, the present invention provides a spectrum-based molecular composition identification and property prediction system, comprising: a spectral data acquisition module for performing infrared spectral detection and gas chromatography detection on the gasoline to be tested to acquire infrared spectral data and gas chromatography data; an infrared spectral identification module for performing hierarchical similarity comparison on the infrared spectral data using a dual spectral database, and identifying a first predicted molecular composition based on a similar sample spectral dataset; a chromatographic composition fusion module for identifying a second predicted molecular composition based on the gas chromatography data, and linearly summing the first and second predicted molecular compositions to obtain an overall predicted molecular composition; and a property prediction analysis module for constructing a molecular property prediction plugin based on ensemble learning and neural networks, analyzing the overall predicted molecular composition to obtain predicted molecular attribute information, and using the overall predicted molecular composition and predicted molecular attribute information as the molecular detection results of the gasoline to be tested.
[0008] The beneficial effects of this invention are:
[0009] The gasoline to be tested is subjected to infrared spectroscopy and gas chromatography detection to obtain infrared spectral data and gas chromatography data. The complementary molecular information obtained by the dual detection method provides a data foundation for subsequent accurate identification. The infrared spectral data is compared hierarchically using a dual spectral database. The first predicted molecular composition is identified based on the similar sample spectral dataset, realizing molecular composition identification based on infrared spectroscopy. The second predicted molecular composition is identified based on gas chromatography data. The first and second predicted molecular compositions are linearly summed to obtain the overall predicted molecular composition. The accuracy of molecular composition identification is improved by fusing the results of the two detection methods. A molecular property prediction plugin is constructed based on ensemble learning and neural networks. The predicted molecular attribute information is obtained by analyzing the overall predicted molecular composition. The overall predicted molecular composition and the predicted molecular attribute information are used as the molecular detection results of the gasoline to be tested, realizing the prediction of physical properties based on accurate molecular composition.
[0010] By combining infrared spectroscopy and gas chromatography, the above technical solution improves the comparison accuracy by utilizing a dual spectral database, fuses the dual detection results through linear addition, and employs integrated learning and neural networks to achieve intelligent property prediction. This achieves the technical effect of improving the accuracy of molecular composition identification through dual spectral data fusion and realizing reliable property prediction based on accurate molecular composition. Attached Figure Description
[0011] Figure 1 This is a schematic flowchart of the spectrum-based molecular composition identification and property prediction method provided by the present invention.
[0012] Figure 2This is a schematic diagram of the structure of the spectrum-based molecular composition identification and property prediction system provided by the present invention.
[0013] In the attached diagram, the components represented by each number are as follows:
[0014] 11. Spectral data acquisition module, 12. Infrared spectral identification module, 13. Chromatographic composition fusion module, 14. Physical property prediction and analysis module. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0017] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0018] Example 1, as Figure 1 As shown, embodiments of the present invention provide a method for molecular composition identification and property prediction based on spectra, including:
[0019] S1. Perform infrared spectroscopy and gas chromatography on the gasoline to be tested to obtain infrared spectral data and gas chromatography data.
[0020] Specifically, the gasoline to be tested refers to gasoline samples that require molecular composition identification and property prediction. These can be intermediate products from the refinery production process, finished gasoline, or gasoline samples that require quality testing and performance evaluation, including but not limited to various types of gasoline products such as commercial gasoline, blended gasoline, refined gasoline, reformed gasoline, catalytic cracking gasoline, and hydrocracking gasoline.
[0021] First, the gasoline to be tested is subjected to infrared spectroscopy, preferably using a Fourier transform infrared spectrometer (FTIR). Infrared spectroscopy is based on the principle of molecular vibrational absorption. When infrared light shines on the gasoline, the organic molecules in it will vibrate and absorb light at specific wavenumbers. Different molecular groups (such as CH bonds, C=C bonds, CO bonds, etc.) have characteristic infrared absorption peaks. The sample is detected within the infrared spectral range (typically 4000-4000 cm⁻¹). -1 The changes in absorbance of gasoline were analyzed to obtain infrared spectral data reflecting the molecular structure of gasoline. Infrared spectral data contains molecular vibrational characteristics of various hydrocarbon compounds in gasoline, providing important structural data for subsequent molecular composition identification.
[0022] Simultaneously, gas chromatography (GC) analysis was performed on the same gasoline sample, preferably using gas chromatography-flame ionization detector (GC-FID). GC detection is based on the separation principle of different compounds in a chromatographic column. The components in gasoline are separated by time in the column according to differences in their physicochemical properties such as boiling point and polarity. Gas chromatographic data are obtained by detecting the signal intensity at different retention times. The GC data contains quantitative information such as retention time, peak area, and peak height of each component in the gasoline, accurately reflecting the content distribution of different types of compounds in the gasoline, including saturated hydrocarbons, aromatics, oxygen-containing compounds, and sulfur-containing compounds.
[0023] The dual-spectral detection method described above can simultaneously acquire structural feature information (infrared spectral data) and component content information (gas chromatographic data) of the gasoline to be tested, providing complementary detection data for molecular composition identification in subsequent steps. This overcomes the limitation of incomplete information in single detection methods and improves the accuracy and reliability of molecular composition identification.
[0024] S2. Using a dual spectral database, perform hierarchical similarity comparison on the infrared spectral data, and identify the first predicted molecular composition based on the similar sample spectral dataset.
[0025] Specifically, a pre-constructed dual-spectral database is used to perform hierarchical similarity comparison analysis on the obtained infrared spectral data, and the molecular composition is initially identified through the principle of similarity matching.
[0026] The dual-spectral database employs a hierarchical architecture, comprising a first spectral alignment layer and a second spectral alignment layer. The first spectral alignment layer stores representative infrared spectral data from multiple samples, serving as a preliminary screening layer; the second spectral alignment layer stores detailed sample spectral data based on a cluster structure, serving as a precise matching layer. This hierarchical structure significantly improves retrieval efficiency while ensuring alignment accuracy.
[0027] First, the infrared spectral data of the gasoline to be tested is compared with the representative infrared spectral data of the samples in the first spectral comparison layer. The Dynamic Time Warping (DTW) similarity comparison algorithm is used to calculate the similarity between the spectra, and candidate spectral datasets that meet the first similarity threshold are selected. Then, based on the comparison results of the first layer, the corresponding mapping clusters in the second spectral comparison layer are called for a more refined similarity comparison. A more stringent second similarity threshold is used for selection, and finally, a similar sample spectral dataset is obtained.
[0028] Subsequently, based on the obtained similar sample spectral dataset, the corresponding sample molecular composition information is extracted to form a sample molecular composition set. Considering the complex nonlinear relationship between infrared spectra and molecular composition, a nonlinear fitting method is used to perform data fusion processing on the sample molecular composition set. A weighted calculation is then performed by combining the confidence weights of each similar sample to output the first predicted molecular composition. The first predicted molecular composition reflects the gasoline molecular composition information obtained based on infrared spectral feature identification.
[0029] Through the aforementioned hierarchical similarity comparison mechanism, efficient and accurate infrared spectral matching was achieved, avoiding the inefficiency of full-database traversal search. At the same time, nonlinear fitting processing improved the accuracy of molecular composition identification, laying the foundation for subsequent composition fusion.
[0030] S3. Based on the gas chromatography data, a second predicted molecular composition is identified, and the first and second predicted molecular compositions are linearly summed to obtain the overall predicted molecular composition.
[0031] Specifically, independent molecular composition identification is performed based on the obtained gas chromatography data, and the results are fused with the obtained first predicted molecular composition to obtain a more accurate and reliable overall predicted molecular composition.
[0032] First, molecular composition identification is performed based on gas chromatography (GC) data, leveraging the high separation capability and quantitative accuracy of GC. By analyzing the retention time, peak area, and peak shape characteristics of each component in the GC data, and combining this with a database of standard substance chromatographic behavior, various compound components in gasoline are identified. This identification process includes qualitative peak identification and quantitative analysis. Qualitative identification determines the compound type based on retention time matching, while quantitative analysis determines the relative content of each component based on peak area calculation. Through the above analytical process, a second predicted molecular composition based on GC detection is obtained, which exhibits good quantitative accuracy and component resolution.
[0033] Subsequently, the first and second predicted molecular compositions were fused using a linear summation method. Specifically, based on the characteristics and reliability of infrared spectroscopy and gas chromatography, fusion weight coefficients were assigned to the first and second predicted molecular compositions respectively. For different component types in gasoline (saturated hydrocarbons, aromatics, gums, asphaltenes, sulfur-containing compounds, oxygen-containing compounds, etc.), the weight allocation was dynamically adjusted based on the detection sensitivity and accuracy of each detection method for that type of component. Through weighted linear combination, the two prediction results were fused according to the optimal weight ratio to obtain the overall predicted molecular composition. The linear summation method fully leverages the advantages of infrared spectroscopy in structural identification and gas chromatography in quantitative analysis, effectively compensating for the limitations of a single detection method through complementary information fusion. The combination of molecular structural fingerprint information provided by infrared spectroscopy data and component content information provided by gas chromatography data significantly improves the comprehensiveness and accuracy of molecular composition identification.
[0034] Through the above dual prediction fusion process, the overall predicted molecular composition combines the detection advantages of two spectroscopic techniques. Compared with the identification results of a single method, it has higher reliability and completeness, providing a high-quality molecular composition data foundation for subsequent property prediction.
[0035] S4. Construct a molecular property prediction plugin based on ensemble learning and neural networks, obtain predicted molecular attribute information based on the overall predicted molecular composition analysis, and use the overall predicted molecular composition and predicted molecular attribute information as the molecular detection result of the gasoline to be detected.
[0036] Specifically, an intelligent molecular property prediction plugin is constructed based on ensemble learning theory and neural networks. It uses the obtained overall predicted molecular composition data to predict molecular attribute information, and integrates the overall predicted molecular composition and predicted molecular attribute information as the molecular detection result of the gasoline to be tested.
[0037] The molecular property prediction plugin employs an ensemble learning framework, constructing and optimizing multiple backpropagation (BP) neural network prediction units. First, a training dataset is built based on historical gasoline testing records, including a set of sample molecular compositions and corresponding sets of sample molecular attribute information. This molecular attribute information includes fundamental thermodynamic properties such as boiling point, critical temperature, critical pressure, critical volume, molar volume, eccentricity factor, melting point, Gibbs free energy, standard enthalpy of formation, enthalpy of melting, enthalpy of vaporization, and solubility parameters. The training data is divided into P equal parts (P being an integer greater than 10 and less than or equal to 50). P different training subsets are obtained through sampling with replacement, and each subset is used to train one of the P BP neural network prediction units until the loss function converges. Finally, these subsets are integrated to obtain a molecular property prediction plugin with strong generalization capabilities.
[0038] To improve prediction accuracy and computational efficiency, an adaptive unit selection mechanism based on prediction complexity was designed. Prediction complexity is evaluated based on the overall predicted molecular composition, where it is positively correlated with the number of molecules in the composition and negatively correlated with the historical frequency of molecule occurrence. The ratio of prediction complexity to standard prediction complexity is set as the unit selection coefficient, dynamically determining the number K of suitable units to select. K units are randomly selected from the P prediction units in the molecular property prediction plugin for prediction calculation. The K prediction results are then averaged to obtain the final predicted molecular attribute information.
[0039] Subsequently, the obtained overall predicted molecular composition and predicted molecular attribute information are integrated to form the molecular detection results of the gasoline to be tested. These results include the types, content distribution, and corresponding physicochemical properties of various molecular components in the gasoline, providing comprehensive data support for gasoline quality assessment, performance prediction, and process optimization.
[0040] Through the combined application of integrated learning and neural networks, the constructed molecular property prediction plugin has strong prediction accuracy and robustness. It can reliably predict various physical property parameters of gasoline based on accurate molecular composition information, realizing the transformation from molecular composition identification to physical property prediction. This solves the technical problems of low accuracy in gasoline molecular composition identification, large limitations of single detection methods, and inaccurate physical property prediction in existing technologies.
[0041] Furthermore, infrared spectral data and gas chromatographic data are acquired, including:
[0042] S11. Perform Fourier transform infrared spectroscopy and gas chromatography on the gasoline to be tested to obtain initial infrared spectral data and initial gas chromatography data.
[0043] S12. Perform first-order and second-order derivative processing on the initial infrared spectral data to obtain infrared spectral data;
[0044] S13. Perform data preprocessing on the initial gas chromatographic data to obtain gas chromatographic data, wherein the data preprocessing includes baseline correction, noise reduction, smoothing, and peak identification and alignment.
[0045] In a preferred embodiment, the gasoline to be tested is first subjected to Fourier transform infrared spectroscopy and gas chromatography to obtain initial infrared spectral data and initial gas chromatographic data, respectively. Specifically, in the infrared spectroscopy detection, a Fourier transform infrared spectrometer (FTIR) is used to scan and detect the gasoline to be tested, and the preferred detection wavenumber range is... The resolution is set to Or higher, with 32 or more scans, to obtain initial infrared spectral data containing complete molecular vibrational information. In gas chromatography detection, a gas chromatograph is used to separate and detect the gasoline to be tested. The gas chromatograph can be a single-column gas chromatograph, a multidimensional gas chromatograph, or a fully two-dimensional gas chromatograph. Among them, single-column gas chromatography is suitable for routine component analysis, multidimensional gas chromatography is suitable for further separation of complex components, and fully two-dimensional gas chromatography is suitable for ultra-high resolution full component analysis. Through the above detection process, initial gas chromatographic data containing information such as the retention time and peak intensity of each component are obtained.
[0046] Subsequently, the initial infrared spectral data underwent mathematical processing to enhance spectral features and eliminate interference signals. First, first-order derivative processing was performed, calculating the difference between adjacent data points to enhance spectral resolution, highlight the separation of overlapping peaks, and reduce the impact of baseline drift. Second-order derivative processing was then performed to further enhance the sharpness and separation of spectral peaks, improving the ability to identify weak peaks. This derivative processing can be implemented using the Savitzky-Golay filtering algorithm or other appropriate numerical differentiation algorithms, resulting in spectral data with clearer characteristic peak information and a better signal-to-noise ratio. Through these processes, infrared spectral data suitable for subsequent similarity comparison analysis was obtained.
[0047] Subsequently, the initial gas chromatographic data underwent data preprocessing to improve data quality and analytical accuracy. Specifically, firstly, baseline correction was performed using polynomial fitting, asymptotic least squares, or other appropriate baseline correction algorithms to eliminate baseline drift and tilt, ensuring baseline stability and consistency, and providing a foundation for accurate peak area integral calculation. Next, denoising was performed using wavelet transform, moving average filtering, or other digital filtering techniques to remove high-frequency noise and random interference signals from the chromatographic data, effectively improving the signal-to-noise ratio and preserving the true shape characteristics of the chromatographic peaks. Based on denoising, further smoothing was performed using Savitzky-Golay smoothing, Gaussian smoothing, or other data smoothing techniques to optimize the chromatographic data, reducing random fluctuations while maintaining the integrity and resolution of the chromatographic peaks. Following this, peak identification and alignment were performed, using peak detection algorithms to automatically identify the start point, peak apex position, and end point of each chromatographic peak in the chromatogram, establishing a complete peak information database. Simultaneously, retention time correction and peak alignment were performed to eliminate potential time shifts between different batches of analysis, ensuring data consistency and comparability. Through the above preprocessing, gas chromatographic data with excellent data quality are obtained, providing reliable data support for subsequent molecular composition identification and analysis.
[0048] The above data acquisition and preprocessing process ensures the high quality and standardization of infrared spectroscopy and gas chromatography data, laying a data foundation for subsequent layered similarity comparison and molecular composition identification, thereby improving the accuracy and reliability of the entire detection.
[0049] Furthermore, the method for constructing the dual spectral database includes:
[0050] S211. A Fourier transform infrared spectrometer is used to detect several gasoline samples to obtain several initial sample infrared spectral data, and molecular detection is performed to obtain the molecular composition of several samples.
[0051] S212. Effective wavelength selection and preprocessing are performed on the infrared spectral data of the several initial samples respectively to obtain infrared spectral data of several samples;
[0052] S213. Use the K-means algorithm to cluster the infrared spectral data of the several samples to obtain multiple clusters and representative infrared spectral data of multiple samples;
[0053] S214. Construct a first spectral comparison layer based on the representative infrared spectral data of the multiple samples;
[0054] S215. Establish the mapping relationship between the representative data of the sample infrared spectrum and the clusters, and construct a second spectral comparison layer based on the multiple clusters;
[0055] S216. A dual spectral database is constructed based on a hierarchical alignment mechanism, consisting of a first spectral alignment layer, a second spectral alignment layer, and several sample molecules.
[0056] In a preferred embodiment, firstly, a Fourier transform infrared spectrometer is used to detect several different types of gasoline samples, obtaining initial sample infrared spectral data covering different compositional characteristics, resulting in several initial sample infrared spectral data. The gasoline samples include, but are not limited to, representative gasoline products with different refining processes, different crude oil sources, and different octane ratings, ensuring sample diversity and representativeness. Simultaneously, molecular detection techniques (such as gas chromatography-mass spectrometry, nuclear magnetic resonance, etc.) are used to perform detailed molecular composition analysis on the above gasoline samples, obtaining the corresponding sample molecular composition, resulting in several sample molecular compositions, and establishing the correlation between spectral data and molecular composition.
[0057] Subsequently, the acquired initial sample infrared spectral data underwent standardization preprocessing to improve data quality. Specifically, effective wavelengths were first selected based on the characteristic absorption bands of gasoline molecules, choosing effective wavelength ranges with rich information content and low interference, while removing redundant information and bands with significant noise interference. Data preprocessing was then performed, including baseline correction, normalization, and smoothing filtering, to eliminate the influence of instrument differences and environmental factors, obtaining standardized sample infrared spectral data. This resulted in several sample infrared spectral data sets, ensuring data consistency and comparability.
[0058] Then, the standardized sample infrared spectral data were clustered using the K-means clustering algorithm. Based on the similarity characteristics of the spectral data, samples with similar spectral features were grouped into the same cluster, forming multiple clusters. Each cluster represents a class of gasoline samples with similar molecular composition characteristics. By calculating the centroid of the infrared spectral data of all samples within each cluster, representative infrared spectral data of multiple samples were obtained. The representative infrared spectral data of the samples is the average spectral feature of the infrared spectral data of all samples in the corresponding cluster, which can effectively describe the typical spectral fingerprint of this type of gasoline sample, providing a high-quality reference standard for subsequent rapid comparison.
[0059] Next, a first spectral comparison layer is constructed based on the obtained representative infrared spectral data of multiple samples, serving as the initial screening structure of the database. This first spectral comparison layer employs a simplified design, storing only the representative infrared spectral data of each cluster, thus achieving efficient indexing of a large amount of sample data. This first spectral comparison layer is primarily used for rapid initial screening; by comparing the similarity with the representative infrared spectral data of the samples, it quickly locates potentially matching cluster regions, improving retrieval efficiency and avoiding the computational burden of a full database traversal search.
[0060] Subsequently, a precise mapping relationship was established between representative infrared spectral data of the samples and their corresponding clusters, and a second spectral alignment layer was constructed based on each cluster. This second spectral alignment layer stores detailed sample infrared spectral data and corresponding molecular composition information within each cluster, providing data support for precise matching and detailed analysis. By establishing this mapping relationship, rapid location and retrieval from representative data to detailed data was achieved, providing the technical foundation for the hierarchical alignment mechanism.
[0061] Subsequently, based on the constructed hierarchical comparison mechanism, the first spectral comparison layer, the second spectral comparison layer, and the corresponding sample molecular composition data, a complete dual-spectral database was integrated and constructed. This dual-spectral database employs a hierarchical retrieval strategy: first, a single-level comparison is performed through the first spectral comparison layer; then, based on the results of the single-level comparison, the second spectral comparison layer retrieves the dual data for precise comparison. While ensuring similarity comparison accuracy, the hierarchical architecture significantly improves retrieval efficiency. The dual-spectral database supports various similarity measurement algorithms, including Euclidean distance for routine comparisons and Dynamic Time Warping (DTW) for handling complex oil products with significant variations in spectral characteristics, ensuring adaptability and accuracy for different types of gasoline samples.
[0062] The dual-spectral database established using the above database construction method has high-efficiency retrieval performance and accurate matching capability, providing technical support for the rapid and accurate identification of gasoline molecular composition and solving the technical problems of low retrieval efficiency and insufficient accuracy of traditional single-layer databases.
[0063] Furthermore, using a dual spectral database, the infrared spectral data is subjected to hierarchical similarity comparison, and the first predicted molecular composition is identified based on the similar sample spectral dataset, including:
[0064] S22. Using the first spectral comparison layer in the dual spectral database, the DWT similarity comparison algorithm is used to perform similarity comparison on the infrared spectral data, and outputs a first similar spectral dataset that meets the first similarity threshold.
[0065] S23. Based on the first similar spectral dataset, call multiple mapping clusters of the second spectral comparison layer to perform similarity comparison, and output a similar sample spectral dataset that meets the second similarity threshold, wherein the second similarity threshold is greater than the first similarity threshold;
[0066] S24. Obtain the sample molecular composition set of the similar sample spectral dataset, and perform nonlinear fitting on the sample molecular composition set to output the first predicted molecular composition.
[0067] In a preferred embodiment, preliminary similarity screening is first performed using a first spectral alignment layer in the dual spectral database. For example, a dynamic time-warping similarity alignment algorithm is used to calculate the similarity between the infrared spectral data of the gasoline to be tested and the representative infrared spectral data of the samples stored in the first spectral alignment layer. Compared with the traditional Euclidean distance method, the dynamic time-warping similarity alignment algorithm has stronger adaptability and can effectively handle complex oil samples with large variations in spectral characteristics. By allowing flexible alignment of spectral data points, it calculates the optimal matching path and obtains more accurate similarity assessment results. A first similarity threshold is set as the initial screening criterion, and representative spectral data that meet the first similarity threshold are selected to form a first similar spectral dataset, enabling rapid candidate region localization.
[0068] Then, based on the screening results of the first similar spectral dataset, multiple mapped clusters in the corresponding second spectral comparison layer are invoked for more refined similarity comparison analysis. The mapping relationship established through the first layer of screening accurately locates the relevant clusters, avoiding computational redundancy from a full-database search. In the second layer of comparison, the spectral data to be detected is compared one by one with the detailed sample spectral data within each cluster, using a more stringent second similarity threshold for screening. The second similarity threshold is greater than the first similarity threshold, ensuring that the finally selected samples have higher similarity and reliability. Through this two-layer progressive comparison mechanism, a similar sample spectral dataset that meets high-precision matching requirements is output.
[0069] Subsequently, the molecular composition information of samples corresponding to similar sample spectral datasets is obtained, forming a sample molecular composition set. Considering the complex nonlinear mapping relationship between infrared spectra and molecular composition, especially in complex gasoline systems, simple linear fitting methods are insufficient to accurately describe this complex relationship. Therefore, a nonlinear fitting approach is used to process the sample molecular composition set. Nonlinear fitting can be implemented using support vector machine (SVM) regression, neural network regression, or other advanced machine learning algorithms, which can effectively capture the nonlinear correlation patterns between spectral features and molecular composition. By weighted fusion and nonlinear modeling of the molecular composition data of multiple similar samples, the first predicted molecular composition based on infrared spectral feature identification is output.
[0070] Through the aforementioned hierarchical similarity comparison and nonlinear fitting processing, efficient and accurate infrared spectral matching and molecular composition prediction were achieved. The hierarchical comparison mechanism significantly improved retrieval efficiency while ensuring matching accuracy, while nonlinear fitting effectively solved the technical challenge of modeling complex spectral-composition relationships, providing high-quality basic data for subsequent data fusion.
[0071] Furthermore, a nonlinear fitting is performed on the sample molecular composition set to output a first predicted molecular composition, including:
[0072] S241. Based on historical gasoline detection records, collect multiple sample training data, including historical spectral datasets and historical molecular composition sets, and the historical spectral data are labeled with reliable weights.
[0073] S242. Train the support vector machine using the multiple sample training data until convergence to obtain the molecular composition predictor;
[0074] S243. Obtain multiple second alignment similarities of multiple similar sample spectral data in the similar sample spectral dataset, and set multiple data confidence weights based on the multiple second alignment similarities to obtain a data confidence weight set, wherein the data confidence weights are positively correlated with the second alignment similarities;
[0075] S244. Using the molecular composition predictor, the first predicted molecular composition is obtained by analyzing the sample molecular composition set and the data confidence weight set.
[0076] In a preferred embodiment, firstly, a training dataset is constructed based on historical gasoline detection records, and multiple sample training data are collected for training and optimization of the support vector machine model. The sample training data includes historical spectral datasets and corresponding historical molecular composition datasets, forming complete input-output data pairs. To improve training effectiveness and model reliability, the historical spectral data undergoes quality assessment, and each set of historical spectral data is assigned a corresponding credibility weight based on factors such as the reliability of the data source, detection accuracy, and reproducibility. The setting of credibility weights considers multiple dimensions, including data detection conditions, instrument accuracy, and operational standardization, ensuring that high-quality data plays a more significant role in model training and improving the overall performance of the molecular composition predictor.
[0077] Subsequently, the Support Vector Machine (SVM) was trained using multiple constructed sample training data until the model parameters converged and reached optimal performance. SVM is a nonlinear regression algorithm that effectively handles the complex nonlinear relationship between spectral data and molecular composition through kernel function mapping. During training, cross-validation techniques were used to optimize model parameters, including the penalty parameter C and kernel function parameters, to ensure good generalization ability. Through iterative training until the loss function converged, a stable molecular composition predictor was obtained, capable of accurately predicting corresponding molecular composition information based on spectral features.
[0078] Next, the second alignment similarity, calculated during the second-layer alignment process, is obtained from the spectral data of each similar sample in the similar sample spectral dataset, resulting in multiple second alignment similarities. Based on these multiple second alignment similarities, corresponding data confidence weights are assigned, establishing a positive correlation between the second alignment similarities and the data confidence weights. Specifically, similar sample spectral data with higher second alignment similarities are assigned higher data confidence weights, reflecting their importance and reliability in the prediction process. For example, firstly, all obtained second alignment similarities are normalized to a standard range of zero to one, eliminating the influence of numerical range differences. Subsequently, exponential or power function transformations are used to nonlinearly amplify the normalized second alignment similarities, enhancing the weight difference between high-similarity and low-similarity samples, resulting in samples with higher second alignment similarities receiving significantly larger weight coefficients. Finally, all transformed weight values are standardized to ensure that the sum of all data confidence weights is one, forming multiple data confidence weight sets. This similarity-based weighting mechanism ensures that historical data that is more similar to the gasoline being tested plays a more important role in the prediction, thus improving prediction accuracy.
[0079] Subsequently, the trained molecular composition predictor is used to perform weighted prediction analysis by combining the sample molecular composition set and the corresponding data credibility weight set. During the prediction process, the prediction results are weighted and fused according to the data credibility weights of the spectral data of each similar sample, allowing highly similar and highly credible sample data to have a greater impact on the final prediction result. Through weighted nonlinear regression calculation, the molecular composition information and credibility of multiple similar sample spectral data are comprehensively considered to output the first predicted molecular composition based on infrared spectral features. This first predicted molecular composition fully utilizes the empirical information of historical data and the similarity characteristics of the current sample, exhibiting high prediction accuracy and reliability.
[0080] High-precision molecular composition prediction was achieved through weighted nonlinear fitting. The introduction of the weighting mechanism effectively improved the accuracy and robustness of the prediction model, enabling the prediction results to better reflect the true molecular composition characteristics of the samples and providing reliable basic data for subsequent data fusion and property prediction.
[0081] Furthermore, molecular property prediction plugins are constructed based on ensemble learning and neural networks, including:
[0082] S41. Based on the historical test records of gasoline, collect the sample molecular composition set and obtain the historical molecular attribute information of different sample molecular compositions to obtain the sample molecular attribute information set. The molecular attribute information includes at least boiling point, critical temperature, critical pressure, critical volume, molar volume, eccentricity factor, melting point, Gibbs free energy, standard enthalpy of formation, enthalpy of melting, enthalpy of vaporization and solubility parameters.
[0083] S42. The sample molecular composition set and sample molecular attribute information set are used as training data and divided into P equal parts. The first training set is obtained by selecting P times with replacement. The first training set is obtained by iteratively selecting P times.
[0084] S43. Using the P training sets, supervise the training of the BP neural network until the loss function converges, to obtain P molecular property prediction units, and integrate them to obtain a molecular property prediction plugin.
[0085] In a preferred embodiment, firstly, data is collected based on historical gasoline testing records to obtain a set of sample molecular compositions covering different gasoline types and compositional characteristics, forming the input part of the training data. Simultaneously, historical molecular attribute information corresponding to each sample molecular composition is obtained through experimental measurement or theoretical calculation, constructing a complete set of sample molecular attribute information as the output part of the training data. This molecular attribute information covers key thermodynamic and physicochemical properties of molecules, including at least fundamental thermodynamic parameters such as boiling point, critical temperature, critical pressure, critical volume, molar volume, eccentricity factor, melting point, Gibbs free energy, standard enthalpy of formation, enthalpy of melting, enthalpy of vaporization, and solubility parameters. These attribute parameters comprehensively describe the physical properties of gasoline molecules, providing a data foundation for constructing a high-precision molecular property prediction plugin.
[0086] Subsequently, the sample molecular composition set and sample molecular attribute information set were used as the original training data, and the training sets were diversified through sampling methods. First, the original training data was divided into P equal parts, where P is an integer greater than 10 and less than or equal to 50, ensuring sufficient data segmentation granularity. Then, sampling with replacement was used to select P times from the P parts of data to form the first training set. This process allowed the same data sample to appear repeatedly in the training set, increasing the diversity of the training data. By iteratively repeating the above sampling process P times, P different training sets were finally obtained, each with different data combination characteristics. This sampling strategy effectively enhanced the randomness and representativeness of the training data, providing a diversified data foundation for ensemble learning and helping to improve the generalization ability and robustness of the final molecular property prediction plugin.
[0087] Subsequently, P training sets were constructed to independently supervise the training of the BP neural network. Each training set corresponded to training an independent neural network prediction unit, serving as the molecular property prediction unit. The BP neural network employed the backpropagation algorithm for parameter optimization, calculating the predicted output through forward propagation and updating the network weights and bias parameters through backpropagation. During training, changes in the loss function were monitored, and the model performance was evaluated using mean squared error or other appropriate loss functions. Training was stopped when the loss function converged to a preset threshold or the training epochs reached their maximum value. Through this training process, P stable molecular property prediction units were obtained, each capable of predicting corresponding molecular attribute parameters based on molecular composition information. Finally, the P molecular property prediction units were integrated to form a molecular property prediction plugin. This plugin possesses the advantages of multi-model collaborative prediction, effectively reducing the prediction error of a single model and improving the reliability and accuracy of the prediction results.
[0088] By constructing a neural network based on ensemble learning, the obtained molecular property prediction plugin has strong predictive power and good generalization performance. It can accurately predict a variety of thermodynamic properties of gasoline molecules, providing a reliable technical means for gasoline quality assessment and performance prediction.
[0089] Furthermore, based on the overall predicted molecular composition analysis, predicted molecular attribute information is obtained, including:
[0090] S44. Based on the overall predicted molecular composition, evaluate the prediction complexity and output the prediction complexity. The prediction complexity is positively correlated with the number of molecules in the molecular composition and negatively correlated with the historical frequency of the molecules.
[0091] S45. Set the ratio of the prediction complexity to the standard prediction complexity as the unit selection coefficient, and take the integer part of the product of the unit selection coefficient and the initial unit selection number to obtain the adaptive unit selection number K, wherein the initial unit selection number is 5, and K is greater than or equal to 1 and less than or equal to P.
[0092] S46. Randomly select K molecular property prediction units from the P molecular property prediction units of the molecular property prediction plugin, predict molecular properties based on the overall predicted molecular composition, and calculate the mean of the K prediction results to obtain the predicted molecular property information.
[0093] In a preferred embodiment, firstly, the prediction complexity is assessed based on the obtained overall predicted molecular composition by analyzing the characteristic parameters of the molecular composition. The calculation of prediction complexity considers two key factors: firstly, prediction complexity is positively correlated with the number of molecules in the molecular composition; that is, the more types of molecules there are, the more complex the intermolecular interactions, and the greater the difficulty of predicting physical properties. Secondly, prediction complexity is negatively correlated with the frequency of each molecule's occurrence in the historical database; that is, the lower the historical frequency of a molecule, the scarcer its physical property data, and the higher the uncertainty of the prediction. By comprehensively considering the molecule number factor and the historical frequency factor, the prediction complexity is quantitatively assessed, and a prediction complexity reflecting the difficulty of the current prediction task is output, providing a decision-making basis for subsequent prediction unit selection. For example, firstly, the total number of molecular types included in the overall predicted molecular composition is counted and denoted as the molecule number factor; then, the occurrence frequency of each molecule is queried from the historical database, the occurrence frequency of each molecule is calculated, and the reciprocals of the occurrence frequencies of all molecules are summed to obtain the frequency factor. The molecular quantity factor and frequency factor are weighted and summed, with the weighting coefficient for the molecular quantity factor set to 0.6 and the weighting coefficient for the frequency factor set to 0.4. The weighted sum is the prediction complexity. For ease of subsequent processing, the calculated prediction complexity can be mapped to the [0,1] range through logarithmic transformation or normalization. Through the above quantitative evaluation, the prediction complexity reflecting the difficulty of the current prediction task is output, providing a decision-making basis for the subsequent selection of molecular property prediction units.
[0094] Subsequently, the calculated prediction complexity is compared with the preset standard prediction complexity to obtain the unit selection coefficient, which reflects the complexity of the current prediction task relative to the standard task. The unit selection coefficient is then multiplied by the initial unit selection number and rounded to obtain the appropriate unit selection number K. The initial unit selection number is set to 5, serving as the baseline selection number for the standard prediction task. Through the above calculations, when the prediction task complexity is high, more molecular property prediction units are selected to improve accuracy; when the prediction task is relatively simple, fewer molecular property prediction units are selected to improve computational efficiency. To ensure the feasibility and effectiveness of the prediction, the value of K is set to be greater than or equal to 1 and less than or equal to P, meaning at least one molecular property prediction unit is selected, and at most all P molecular property prediction units are selected.
[0095] Subsequently, from the P molecular property prediction units included in the molecular property prediction plugin, K molecular property prediction units are randomly selected to participate in the prediction calculation. This random selection strategy ensures the diversity of prediction units and avoids the bias problems that might arise from fixed selection. The overall predicted molecular composition is used as input data and fed into the selected K molecular property prediction units for parallel prediction calculations. Each molecular property prediction unit outputs the corresponding molecular property prediction result based on its trained model parameters. Finally, the prediction results of the K molecular property prediction units are calculated by arithmetic mean, and the final predicted molecular property information is obtained through ensemble fusion. This multi-model ensemble prediction method effectively reduces the prediction error and uncertainty of a single model, and improves the reliability and stability of the prediction results.
[0096] Through the aforementioned adaptive prediction unit selection and integrated prediction mechanism, an intelligent prediction strategy that dynamically adjusts computing resources according to the complexity of the prediction task is realized. While ensuring prediction accuracy, the computational efficiency is optimized, providing personalized physical property prediction services for gasoline under test with different levels of complexity.
[0097] Example 2, as Figure 2 As shown, based on the same inventive concept as the spectrum-based molecular composition identification and property prediction method provided in Embodiment 1, this embodiment of the invention also provides a spectrum-based molecular composition identification and property prediction system, including:
[0098] The spectral data acquisition module 11 is used to perform infrared spectral detection and gas chromatography detection on the gasoline to be tested to acquire infrared spectral data and gas chromatography data.
[0099] The infrared spectral recognition module 12 is used to perform hierarchical similarity comparison of the infrared spectral data using a dual spectral database, and to identify the first predicted molecular composition based on the similar sample spectral dataset.
[0100] The chromatographic composition fusion module 13 is used to identify a second predicted molecular composition based on the gas chromatographic data, and to linearly sum the first and second predicted molecular compositions to obtain the overall predicted molecular composition.
[0101] The property prediction and analysis module 14 is used to construct a molecular property prediction plugin based on ensemble learning and neural networks, obtain predicted molecular attribute information based on the overall predicted molecular composition analysis, and use the overall predicted molecular composition and predicted molecular attribute information as the molecular detection result of the gasoline to be tested.
[0102] Furthermore, the execution steps of the spectral data acquisition module 11 include:
[0103] The gasoline to be tested was subjected to Fourier transform infrared spectroscopy and gas chromatography to obtain initial infrared spectral data and initial gas chromatography data.
[0104] The initial infrared spectral data is processed by first-order and second-order derivatives to obtain infrared spectral data.
[0105] The initial gas chromatographic data is preprocessed to obtain gas chromatographic data. The preprocessing includes baseline correction, noise reduction, smoothing, and peak identification and alignment.
[0106] Furthermore, the method for constructing the dual spectral database includes:
[0107] Fourier transform infrared spectroscopy was used to detect several gasoline samples, obtain several initial sample infrared spectral data, and molecular detection was performed to obtain the molecular composition of several samples.
[0108] Effective wavelengths are selected and preprocessed on the infrared spectral data of the initial samples to obtain infrared spectral data of the samples.
[0109] The K-means algorithm is used to cluster the infrared spectral data of the samples to obtain multiple clusters and representative infrared spectral data of the samples.
[0110] A first spectral comparison layer is constructed based on the representative infrared spectral data of the multiple samples;
[0111] Establish a mapping relationship between representative infrared spectral data of samples and clusters, and construct a second spectral comparison layer based on the multiple clusters;
[0112] A dual-spectral database is constructed based on a hierarchical alignment mechanism, a first spectral alignment layer, a second spectral alignment layer, and several sample molecules.
[0113] Furthermore, the execution steps of the infrared spectral recognition module 12 include:
[0114] Using the first spectral comparison layer of the dual spectral database, the DWT similarity comparison algorithm is used to perform similarity comparison on the infrared spectral data, and outputs a first similar spectral dataset that meets the first similarity threshold.
[0115] Based on the first similar spectral dataset, multiple mapping clusters of the second spectral comparison layer are called to perform similarity comparison, and a similar sample spectral dataset that meets the second similarity threshold is output, wherein the second similarity threshold is greater than the first similarity threshold;
[0116] Obtain the sample molecular composition set of the similar sample spectral dataset, and perform nonlinear fitting on the sample molecular composition set to output the first predicted molecular composition.
[0117] Furthermore, the execution steps of the infrared spectral recognition module 12 also include:
[0118] Based on historical gasoline testing records, multiple sample training data were collected. The sample training data included historical spectral datasets and historical molecular composition datasets, and the historical spectral data were labeled with reliable weights.
[0119] The support vector machine is trained until convergence using the multiple sample training data to obtain the molecular composition predictor;
[0120] Multiple second alignment similarities are obtained from the spectral data of multiple similar samples in the similar sample spectral dataset, and multiple data confidence weights are set based on the multiple second alignment similarities to obtain a data confidence weight set, wherein the data confidence weights are positively correlated with the second alignment similarities;
[0121] Using the molecular composition predictor, the first predicted molecular composition is obtained by analyzing the sample molecular composition set and the data confidence weight set.
[0122] Furthermore, the execution steps of the property prediction and analysis module 14 include:
[0123] Based on historical test records of gasoline, a set of sample molecular composition is collected, and historical molecular attribute information of different sample molecular compositions is obtained to obtain a set of sample molecular attribute information. Among them, the molecular attribute information includes at least boiling point, critical temperature, critical pressure, critical volume, molar volume, eccentricity factor, melting point, Gibbs free energy, standard enthalpy of formation, enthalpy of melting, enthalpy of vaporization, and solubility parameters.
[0124] The sample molecular composition set and sample molecular attribute information set are used as training data and divided into P equal parts. The first training set is obtained by selecting P times with replacement. The first training set is obtained by iteratively selecting P times.
[0125] Using the P training sets, supervised training is performed on the BP neural network until the loss function converges, resulting in P molecular property prediction units, which are then integrated to obtain a molecular property prediction plugin.
[0126] Furthermore, the execution steps of the property prediction and analysis module 14 also include:
[0127] The prediction complexity is evaluated based on the overall predicted molecular composition, and the prediction complexity is output. The prediction complexity is positively correlated with the number of molecules in the molecular composition and negatively correlated with the historical frequency of the molecules.
[0128] The ratio of the prediction complexity to the standard prediction complexity is set as the unit selection coefficient. The product of the unit selection coefficient and the initial unit selection number is rounded to obtain the number of adapted units selected, K, where the initial unit selection number is 5, and K is greater than or equal to 1 and less than or equal to P.
[0129] K molecular property prediction units are randomly selected from the P molecular property prediction units of the molecular property prediction plugin. Molecular property prediction is performed based on the overall predicted molecular composition. The mean of the K prediction results is calculated to obtain the predicted molecular property information.
[0130] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0131] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0132] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0133] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0134] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0135] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.
[0136] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for molecular composition identification and property prediction based on spectra, characterized in that, The methods include: The gasoline to be tested was subjected to infrared spectroscopy and gas chromatography to obtain infrared spectral data and gas chromatography data, respectively. Using a dual spectral database, the infrared spectral data are subjected to hierarchical similarity comparison. Based on the similar sample spectral dataset, the first predicted molecular composition is identified, including: Using the first spectral comparison layer of the dual spectral database, the DWT similarity comparison algorithm is used to perform similarity comparison on the infrared spectral data, and outputs a first similar spectral dataset that meets the first similarity threshold. Based on the first similar spectral dataset, multiple mapping clusters of the second spectral alignment layer are called to perform similarity comparison, and a similar sample spectral dataset that meets the second similarity threshold is output, wherein the second similarity threshold is greater than the first similarity threshold; Obtain the sample molecular composition set of the similar sample spectral dataset, and perform nonlinear fitting on the sample molecular composition set to output the first predicted molecular composition; The method for constructing the dual spectral database includes: Fourier transform infrared spectroscopy was used to detect several gasoline samples, obtain several initial sample infrared spectral data, and molecular detection was performed to obtain the molecular composition of several samples. Effective wavelengths are selected and preprocessed on the infrared spectral data of the initial samples to obtain infrared spectral data of the samples. The K-means algorithm is used to cluster the infrared spectral data of the samples to obtain multiple clusters and representative infrared spectral data of the samples. A first spectral comparison layer is constructed based on the representative infrared spectral data of the multiple samples; Establish a mapping relationship between representative infrared spectral data of samples and clusters, and construct a second spectral comparison layer based on the multiple clusters; A dual-spectral database is constructed based on a hierarchical alignment mechanism, a first spectral alignment layer, a second spectral alignment layer, and several sample molecules. The second predicted molecular composition is identified based on the gas chromatography data. The overall predicted molecular composition is obtained by linearly summing the first and second predicted molecular compositions. A molecular property prediction plugin is constructed based on ensemble learning and neural networks. Predicted molecular attribute information is obtained from the overall predicted molecular composition analysis. The overall predicted molecular composition and predicted molecular attribute information are used as the molecular detection results of the gasoline to be tested.
2. The method for molecular composition identification and property prediction based on spectra according to claim 1, characterized in that, Acquire infrared spectroscopy data and gas chromatography data, including: The gasoline to be tested was subjected to Fourier transform infrared spectroscopy and gas chromatography to obtain initial infrared spectral data and initial gas chromatography data. The initial infrared spectral data is processed by first-order and second-order derivatives to obtain infrared spectral data. The initial gas chromatographic data is preprocessed to obtain gas chromatographic data. The preprocessing includes baseline correction, noise reduction, smoothing, and peak identification and alignment.
3. The method for molecular composition identification and property prediction based on spectra according to claim 1, characterized in that, The sample molecular composition set is subjected to nonlinear fitting to output the first predicted molecular composition, including: Based on historical gasoline testing records, multiple sample training data were collected. The sample training data included historical spectral datasets and historical molecular composition datasets, and the historical spectral data were labeled with reliable weights. The support vector machine is trained until convergence using the multiple sample training data to obtain the molecular composition predictor; Multiple second alignment similarities are obtained from the spectral data of multiple similar samples in the similar sample spectral dataset, and multiple data confidence weights are set based on the multiple second alignment similarities to obtain a data confidence weight set, wherein the data confidence weights are positively correlated with the second alignment similarities; Using the molecular composition predictor, the first predicted molecular composition is obtained by analyzing the sample molecular composition set and the data confidence weight set.
4. The method for molecular composition identification and property prediction based on spectra according to claim 1, characterized in that, Molecular property prediction plugins built based on ensemble learning and neural networks include: Based on historical test records of gasoline, a set of sample molecular composition is collected, and historical molecular attribute information of different sample molecular compositions is obtained to obtain a set of sample molecular attribute information. Among them, the molecular attribute information includes at least boiling point, critical temperature, critical pressure, critical volume, molar volume, eccentricity factor, melting point, Gibbs free energy, standard enthalpy of formation, enthalpy of melting, enthalpy of vaporization, and solubility parameters. The sample molecular composition set and sample molecular attribute information set are used as training data and divided into P equal parts. The first training set is obtained by selecting P times with replacement. The first training set is obtained by iteratively selecting P times. Using the P training sets, supervised training is performed on the BP neural network until the loss function converges, resulting in P molecular property prediction units, which are then integrated to obtain a molecular property prediction plugin.
5. The method for molecular composition identification and property prediction based on spectra according to claim 4, characterized in that, Based on the overall predicted molecular composition analysis, predicted molecular attribute information is obtained, including: The prediction complexity is evaluated based on the overall predicted molecular composition, and the prediction complexity is output. The prediction complexity is positively correlated with the number of molecules in the molecular composition and negatively correlated with the historical frequency of the molecules. The ratio of the prediction complexity to the standard prediction complexity is set as the unit selection coefficient. The product of the unit selection coefficient and the initial unit selection number is rounded to obtain the number of adapted units selected, K, where the initial unit selection number is 5, and K is greater than or equal to 1 and less than or equal to P. K molecular property prediction units are randomly selected from the P molecular property prediction units of the molecular property prediction plugin. Molecular property prediction is performed based on the overall predicted molecular composition. The mean of the K prediction results is calculated to obtain the predicted molecular property information.
6. A system for molecular composition identification and property prediction based on spectra, characterized in that, The system for implementing the spectrum-based molecular composition identification and property prediction method as described in any one of claims 1 to 5 includes: The spectral data acquisition module is used to perform infrared spectral detection and gas chromatography detection on the gasoline to be tested, and to acquire infrared spectral data and gas chromatography data. The infrared spectroscopy recognition module is used to perform hierarchical similarity comparison of the infrared spectral data using a dual spectral database, and to identify the first predicted molecular composition based on the similar sample spectral dataset; The chromatographic composition fusion module is used to identify a second predicted molecular composition based on the gas chromatographic data, and to linearly sum the first and second predicted molecular compositions to obtain the overall predicted molecular composition. The property prediction and analysis module is used to construct a molecular property prediction plugin based on ensemble learning and neural networks. It obtains predicted molecular attribute information based on the overall predicted molecular composition analysis and uses the overall predicted molecular composition and predicted molecular attribute information as the molecular detection results of the gasoline to be tested.
Citation Information
Patent Citations
Method and apparatus for rapid analysis of molecular components of naphtha
WO2025076965A1