Shale oil chromatography fingerprint identification model, model construction method and identification method

The shale oil chromatography fingerprint recognition model constructed by gas chromatography and mass spectrometry combined with K-Shape algorithm solves the problem of inaccurate fingerprint recognition of shale oil saturated hydrocarbon chromatography in the existing technology, and achieves high-precision recognition results.

CN120028473APending Publication Date: 2025-05-23CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311578917.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, the fingerprint recognition method for shale oil saturated hydrocarbon chromatography has a great influence on multiple solutions, mutually exclusive and human subjective factors, resulting in inaccurate identification results.

Method used

The initial detection data was obtained by using gas chromatography mass spectrometry combined technology, and the key attribute field data were extracted. The target chromatography fingerprint time series signal was obtained through baseline correction, filtering and noise reduction and delay correction processing. The shale oil chromatography fingerprint recognition model was constructed based on the K-Shape algorithm.

Benefits of technology

Accurate identification of shale oil saturated hydrocarbon chromatography fingerprints is achieved, reducing the influence of human subjective factors and improving the accuracy and reliability of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120028473A_ABST
    Figure CN120028473A_ABST
Patent Text Reader

Abstract

The invention relates to the field of oil and gas exploration and development, and discloses a shale oil chromatography fingerprint identification model, a model construction method and an identification method, and the method comprises the following steps: detecting a plurality of shale oil saturated hydrocarbon samples by using a gas chromatography-mass spectrometry technology to obtain initial detection data; extracting retention time and chromatographic response data from the initial detection data; based on the retention time and the chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, obtaining each initial chromatographic fingerprint time sequence signal, and optimizing each initial chromatographic fingerprint time sequence signal to obtain each target chromatographic fingerprint time sequence signal; and based on a K-Shape algorithm and each target chromatographic fingerprint time sequence signal, constructing a shale oil chromatographic fingerprint identification model. The shale oil chromatographic fingerprint identification model is constructed based on a K-Shape algorithm with a better clustering effect, so that the shale oil saturated hydrocarbon chromatographic fingerprint can be accurately identified by using the shale oil chromatographic fingerprint identification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oil and gas exploration and development, and in particular to a shale oil chromatographic fingerprint recognition model, a shale oil chromatographic fingerprint recognition model construction method, and a shale oil chromatographic fingerprint recognition method. Background Art

[0002] Shale oil is a natural product with a relatively complex composition formed after a long geological process. Saturated hydrocarbons are one of the main components of shale oil, usually accounting for 50% to 80%; therefore, accurate identification of shale oil saturated hydrocarbon chromatographic fingerprints is the key to achieving fine exploration and development of shale oil.

[0003] At present, the commonly used chromatographic fingerprint identification methods for shale oil saturated hydrocarbons generally include intuitive analysis and comparison method, characteristic parameter construction method, data statistical analysis method, chemometric method, etc.

[0004] However, it has been found in practice that the results obtained by the above-mentioned shale oil saturated hydrocarbon chromatographic fingerprint identification method are multi-solution and mutually exclusive, and the identification process is greatly influenced by human subjective factors, so the identification results are not accurate. Summary of the invention

[0005] The purpose of the present invention is to overcome the problem of inaccurate chromatographic fingerprint identification of shale oil saturated hydrocarbons in the prior art, and to provide a shale oil chromatographic fingerprint identification model, a model construction method and an identification method.

[0006] In order to achieve the above-mentioned object, the first aspect of the present invention provides a method for constructing a shale oil chromatographic fingerprint recognition model, comprising:

[0007] Obtain multiple shale oil saturated hydrocarbon samples;

[0008] The saturated hydrocarbon samples of shale oil are tested by using gas chromatography-mass spectrometry technology to obtain initial test data corresponding to each shale oil saturated hydrocarbon sample;

[0009] Extract key attribute field data from each initial detection data, wherein the key attribute field data at least includes retention time and chromatographic response data;

[0010] Based on the retention time and chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, the initial chromatographic fingerprint time series signals corresponding to each shale oil saturated hydrocarbon sample are obtained, and each initial chromatographic fingerprint time series signal is optimized to obtain the target chromatographic fingerprint time series signals corresponding to each shale oil saturated hydrocarbon sample;

[0011] Based on the K-Shape algorithm and the time series signals of each target chromatographic fingerprint, a shale oil chromatographic fingerprint recognition model was constructed.

[0012] In the embodiment of the present application, the step of obtaining a plurality of shale oil saturated hydrocarbon samples comprises:

[0013] Based on the wellhead produced fluid and / or shale rock samples, multiple shale oil samples are obtained, and multiple shale oil saturated hydrocarbon samples are obtained from the multiple shale oil samples, and one shale oil sample corresponds to one shale oil saturated hydrocarbon sample.

[0014] In an embodiment of the present application, the key attribute field data also includes peak ID, peak name, retention index, distribution coefficient, number of plates, scan number ID, peak area, peak height, signal-to-noise ratio and half-peak height symmetry.

[0015] In the embodiment of the present application, the optimization processing of each initial chromatographic fingerprint time series signal to obtain the target chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample includes:

[0016] Performing baseline correction processing on the initial chromatographic fingerprint time series signal to obtain a first chromatographic fingerprint time series signal;

[0017] Performing filtering and noise reduction processing on the first chromatographic fingerprint time series signal to obtain a second chromatographic fingerprint time series signal;

[0018] The second chromatographic fingerprint time series signal is subjected to delay correction processing to obtain a target chromatographic fingerprint time series signal.

[0019] In the embodiment of the present application, the baseline correction process is performed on the initial chromatographic fingerprint time series signal to obtain the first chromatographic fingerprint time series signal, including:

[0020] Using fast wavelet transform, the initial chromatographic fingerprint time series signal is decomposed to obtain multiple sub-signals corresponding to the initial chromatographic fingerprint time series signal;

[0021] Threshold processing is performed on the multiple sub-signals, and the multiple sub-signals after the threshold processing are reconstructed to obtain a first chromatographic fingerprint time series signal.

[0022] In the embodiment of the present application, the filtering and noise reduction process is performed on the first chromatographic fingerprint time series signal to obtain the second chromatographic fingerprint time series signal, including:

[0023] Using a local polynomial function to fit the first chromatographic fingerprint time series signal within the filtering window to obtain a fitting value corresponding to the first chromatographic fingerprint time series signal;

[0024] The data corresponding to the first chromatographic fingerprint time series signal is replaced with the fitting value to obtain the second chromatographic fingerprint time series signal.

[0025] In the embodiment of the present application, the shale oil chromatographic fingerprint recognition model is constructed based on the K-Shape algorithm and each target chromatographic fingerprint time series signal, including:

[0026] The distances between the chromatographic fingerprints corresponding to the target chromatographic fingerprint time series signals are calculated using a dynamic time warping algorithm;

[0027] Based on the distance between each chromatographic fingerprint, the K-Means algorithm is used to cluster the time series of each chromatographic fingerprint, and a shale oil chromatographic fingerprint recognition model is constructed according to the clustering results.

[0028] In the embodiment of the present application, after the shale oil chromatographic fingerprint recognition model is constructed, the method further includes:

[0029] The shale oil chromatographic fingerprint recognition model is verified, and the shale oil chromatographic fingerprint recognition model is optimized according to the verification result.

[0030] The second aspect of the present application provides a shale oil chromatographic fingerprint recognition model, which is constructed by the shale oil chromatographic fingerprint recognition model construction method provided by the first aspect of the present application.

[0031] The third aspect of the present application provides a shale oil chromatographic fingerprint identification method, wherein the shale oil chromatographic fingerprint identification method uses the shale oil chromatographic fingerprint identification model provided in the second aspect of the present application to perform shale oil chromatographic fingerprint identification, and the shale oil chromatographic fingerprint identification method comprises:

[0032] Obtaining a saturated hydrocarbon sample of shale oil to be tested;

[0033] Detecting the shale oil saturated hydrocarbon sample to be tested by using gas chromatography-mass spectrometry to obtain initial detection data corresponding to the shale oil saturated hydrocarbon sample to be tested;

[0034] Based on the retention time and chromatographic response data corresponding to the initial detection data, an initial chromatographic fingerprint time series signal corresponding to the initial detection data is obtained, and the initial chromatographic fingerprint time series signal is optimized to obtain a chromatographic fingerprint time series signal to be identified corresponding to the shale oil saturated hydrocarbon sample to be tested;

[0035] The chromatographic fingerprint time series signal to be identified is identified based on the shale oil chromatographic fingerprint identification model.

[0036] Through the above technical scheme, the technical scheme includes: obtaining multiple shale oil saturated hydrocarbon samples; using gas chromatography-mass spectrometry to detect each shale oil saturated hydrocarbon sample, and obtaining initial detection data corresponding to each shale oil saturated hydrocarbon sample; extracting key attribute field data from each initial detection data, and the key attribute field data at least includes retention time and chromatographic response data; based on the retention time and chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, obtaining the initial chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample, optimizing each initial chromatographic fingerprint time series signal, and obtaining the target chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample; based on the K-Shape algorithm and each target chromatographic fingerprint time series signal, constructing a shale oil chromatographic fingerprint recognition model. Based on the scheme provided in the embodiment of the present application, a shale oil chromatographic fingerprint recognition model can be constructed. Since the shale oil chromatographic fingerprint recognition model is constructed based on the K-Shape algorithm with better clustering effect, the shale oil chromatographic fingerprint can be accurately identified based on the shale oil chromatographic fingerprint recognition model.

[0037] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings:

[0039] Figure 1 A schematic diagram of a process for constructing a shale oil chromatographic fingerprint recognition model according to an embodiment of the present application is shown;

[0040] Figure 2 A schematic diagram of a shale oil saturated hydrocarbon chromatographic fingerprint similarity measurement and optimal alignment path according to an embodiment of the present application is schematically shown;

[0041] Figure 3 A schematic diagram of an unsupervised clustering identification process of shale oil saturated hydrocarbon chromatographic fingerprint according to an embodiment of the present application is shown schematically;

[0042] Figure 4-1 A schematic diagram shows a two-cluster identification result of a shale oil saturated hydrocarbon chromatographic fingerprint according to an embodiment of the present application;

[0043] Figure 4-2 A schematic diagram shows a three-cluster identification result of a shale oil saturated hydrocarbon chromatographic fingerprint according to an embodiment of the present application;

[0044] Figure 4-3Schematically showing a four-cluster identification result of a shale oil saturated hydrocarbon chromatographic fingerprint according to an embodiment of the present application;

[0045] Figure 4-4 A schematic diagram shows the five-cluster identification result of a shale oil saturated hydrocarbon chromatographic fingerprint according to an embodiment of the present application;

[0046] Figure 5 A schematic diagram of similarity inheritance relationship of shale oil saturated hydrocarbon chromatographic fingerprint identification according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0048] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back...), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0049] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0050] As described in the background technology, shale oil is a natural product with a relatively complex composition formed after a long geological process. According to the chemical composition, the components in shale oil can be divided into four categories: saturated hydrocarbons, aromatic hydrocarbons, non-hydrocarbons and asphaltene. Among them, the proportion of saturated hydrocarbons in shale oil can usually reach 50% to 80%; accurate identification of the chromatographic fingerprint of shale oil saturated hydrocarbons is the key to achieving fine exploration and development of shale oil. Accurate identification of the chromatographic fingerprint of shale oil saturated hydrocarbons enables accurate analysis of shale oil, which in turn helps in the subsequent fracturing network impact description, formation capacity allocation, well-to-well interference analysis, well spacing optimization plan, reservoir sweet spot identification, residual oil potential evaluation, fracturing construction plan optimization, dynamic capacity prediction, etc.

[0051] In practical applications, the process of analyzing samples based on gas chromatography technology usually includes: using a gas chromatograph to separate the components in a sample containing complex components; then, based on the "uniqueness" of the chromatographic fingerprint, the components in the sample are identified by comparing the shapes of the chromatographic fingerprint. At present, the commonly used shale oil saturated hydrocarbon chromatographic fingerprint identification methods generally include intuitive analysis and comparison methods, characteristic parameter construction methods, data statistical analysis methods, chemometric methods, etc.

[0052] Among them, the intuitive analysis and comparison method includes: using chromatographic methods to separate compounds in samples and obtain chromatograms; overlapping and drawing chromatograms of different samples on the same graph to intuitively compare and evaluate similarities and differences. This method is suitable for situations where the number of samples is small, the number of compound peaks contained in the samples is small, the chromatographic fingerprint structure is simple and clear, and it relies on manual qualitative identification.

[0053] Since common peaks refer to compounds with relative retention values ​​within ±1% as the same compounds, the number of common peaks can be used to indicate the similarity of samples, and the overlap rate characteristic parameters can be constructed to characterize the similarity between samples. This is the principle of the characteristic parameter construction method. Specifically, all chromatographic peaks that appear can be sorted from large to small according to their normalized peak area or relative area value, and the top 70% of compound peaks can be selected as the main components of the samples to identify the differences between samples. In specific implementation, the similarity and difference rates of characteristic fingerprints can be further used as a measure of the similarity between samples. The characteristic fingerprint is a peak group composed of a series of chromatographic peaks, which reflects the characteristics of the biological source of shale oil. The characteristic fingerprint is not limited to completely common chromatographic peaks, and the peaks common to 80% to 90% of the samples can be used as characteristic fingerprints.

[0054] Data statistical analysis combines chemical pattern recognition technology with mathematical statistics. Data statistical analysis can generally be divided into fuzzy information analysis, artificial neural network, grey correlation clustering, principal component analysis, etc. Among them, fuzzy information analysis is to find the characteristics of fuzzy information in the comparison between certain information and random information, extract sufficient characteristics to reflect the quality and quantity of the sample from the chromatographic fingerprint, discard the information that has no effect on the classification decision, and then reflect the fuzziness through "subordinate degree", and use the maximum subordinate principle or threshold principle to realize the identification of the sample. Artificial neural network method is based on certain learning rules, with the help of back propagation to correct the connection weights between each node, based on a wide range of samples and a large amount of training, to obtain a recognition model, and then identify the samples to be identified based on the recognition model. Grey correlation clustering method uses relative correlation as a measurement indicator and constructs a pattern recognition model, and then uses the sequence relative correlation to judge the similarity between samples to realize the identification of shale oil samples. Principal component analysis is to decompose the singular value of the chromatographic data matrix into the product of two orthogonal matrices and a diagonal matrix, and identify the subtle differences between samples from a low-dimensional space.

[0055] The chemometric method uses internal and external standard injection to solve the absolute concentration of a small number of specified compounds between samples. Then, hierarchical cluster analysis, principal component analysis, multidimensional scaling algorithm, alternating least squares method and other methods are used to achieve the purpose of crude oil type classification, differential feature analysis, representative end member selection and mixing ratio contribution distribution. Chemometric methods are usually suitable for oil source comparison when there is a significant gradient in compound concentration.

[0056] However, it has been found in practice that the chromatographic fingerprint identification methods of shale oil saturated hydrocarbons listed above are mostly targeted, pertinent, objective and biased identification methods. Due to the insufficient complexity of the sample data itself, the sample cannot be uniquely identified. In addition, there are often obvious contradictions and conflicts between the identification results of different indicators of different identification methods, so the results obtained by identification may be multi-solution and mutually exclusive. In addition, the chromatographic fingerprint identification methods of shale oil saturated hydrocarbons listed above are mostly artificially selected indicators for absolute qualitative and quantitative analysis, so the range of compounds that can be covered is very limited, and shale oil saturated hydrocarbons cannot be fully identified. Moreover, the indicators are mostly selected based on mechanism knowledge and expert experience, and the selection of indicators may not be appropriate and accurate in fact. It can be seen that the current identification methods are not accurate in identifying the chromatographic fingerprints of shale oil saturated hydrocarbons, making it difficult to meet the needs of refined comparison of shale oil in unconventional tight oil reservoirs.

[0057] Example 1

[0058] In view of the above problems, an embodiment of the present application provides a method for constructing a shale oil chromatographic fingerprint recognition model. The recognition model can specifically recognize the shale oil saturated hydrocarbon chromatographic fingerprint, so the shale oil chromatographic fingerprint recognition model can also be called a shale oil saturated hydrocarbon chromatographic fingerprint recognition model. Accordingly, the method for constructing a shale oil chromatographic fingerprint recognition model provided in the embodiment of the present application can specifically be a method for constructing a shale oil saturated hydrocarbon chromatographic fingerprint recognition model. Figure 1 As shown, the method may include the following steps:

[0059] Step 101: Obtain multiple shale oil saturated hydrocarbon samples.

[0060] In an embodiment of the present application, step 101 of obtaining multiple shale oil saturated hydrocarbon samples may include: obtaining multiple shale oil samples based on wellhead produced fluid and / or shale rock samples, and obtaining multiple shale oil saturated hydrocarbon samples from the multiple shale oil samples, one shale oil sample corresponding to one shale oil saturated hydrocarbon sample.

[0061] In the case where shale oil samples are obtained based on wellhead produced fluid, multiple shale oil samples are obtained based on the wellhead produced fluid, and multiple shale oil saturated hydrocarbon samples are obtained from the multiple shale oil samples, which can include: collecting wellhead produced fluid from different wells as the multiple shale oil samples, or collecting wellhead produced fluid from different production time stages of the same well as the above-mentioned shale oil samples, performing oil-water separation on the multiple shale oil samples to obtain respectively corresponding oil components, performing column chromatography separation treatment on each oil component using a chromatographic column to obtain shale oil saturated hydrocarbon components corresponding to each shale oil sample, and taking at least a portion of each shale oil saturated hydrocarbon component as the corresponding shale oil saturated hydrocarbon sample, that is, obtaining multiple shale oil saturated hydrocarbon samples corresponding to the multiple shale oil samples.

[0062] In specific implementation, the wellhead produced fluid can be collected at the on-site wellhead or obtained from a wellhead separator. When performing oil-water separation on shale oil samples, solvents such as dichloromethane can be used as extraction solvents to achieve oil-water separation. The chromatographic column can specifically be a silica gel-alumina chromatographic column, so that the saturated hydrocarbon components of shale oil can be better separated from the oil components.

[0063] In the case where the shale oil sample is obtained based on the shale rock sample, obtaining multiple shale oil samples based on the shale rock sample, and obtaining multiple shale oil saturated hydrocarbon samples from the multiple shale oil samples can include: grinding the multiple shale rock samples respectively to obtain ground particles corresponding to each shale rock sample, extracting shale oil saturated hydrocarbon components from the soluble organic matter (which can be regarded as the multiple shale oil samples) of the ground particles corresponding to each shale rock sample by Soxhlet extraction and group component separation, and taking at least a part of each shale oil saturated hydrocarbon component as the corresponding shale oil saturated hydrocarbon sample, that is, obtaining multiple shale oil saturated hydrocarbon samples corresponding to the multiple shale oil samples.

[0064] In specific implementation, shale core or rock cutting samples collected from different well locations, different layers, and different development stages can be collected as the multiple shale rock samples. When grinding the shale rock samples, they can be ground to 120 meshes to improve the dissolution effect of soluble organic matter.

[0065] Step 102, using gas chromatography-mass spectrometry to detect each shale oil saturated hydrocarbon sample, and obtaining initial detection data corresponding to each shale oil saturated hydrocarbon sample.

[0066] In practical applications, the test platform of gas chromatography-mass spectrometry (GC-MS) can be used in combination with an HP6890N gas chromatograph and an HP5973N mass spectrometer in series. Among them, the chromatographic column in the gas chromatograph can be an HP-5 elastic quartz capillary column, and the size can be 30m×0.25mm×0.25μm. The column box temperature can be heated from 80°C to 290°C, with a heating rate of 4°C / minute, and the temperature can be maintained at 80°C for 2min and at 290°C for 30min. The carrier gas can be high-purity helium with a purity of 99.999%. In a single quadrupole mass spectrometer, the EI ion source temperature can be 280°C, the ionization energy can be 70eV, dichloromethane (chromatographic grade) can be used as a solvent, and the manual injector can be set to a splitless injection volume of 2μL.

[0067] Step 103: extract key attribute field data from each initial detection data, wherein the key attribute field data at least includes retention time and chromatographic response data.

[0068] In specific implementation, the laboratory information platform can be used to interactively decode each initial detection data to extract key attribute field data. Specifically, the MassHunter workstation software can be used to decode each initial detection data obtained by the detection, and the Nist17 spectral library can be used to qualitatively and quantitatively identify characteristic compounds. Then, the identification results are submitted to the laboratory big data platform, and the key attribute field data are extracted using the platform interactive tools.

[0069] In the embodiments of the present application, retention time refers to the time required for compounds in a sample to be separated by a chromatographic column and enter a mass spectrometer detector in gas chromatography technology; it is one of the characteristics of a peak in a chromatogram and can be used to identify and distinguish compounds.

[0070] In chromatography, the chromatographic response data can be direct chromatographic detector intensity or one of the peak heights and peak areas of different compounds in the chromatographic column, which has a linear relationship with the concentration and mole fraction of the compound.

[0071] In a specific implementation, the key attribute field data may also include peak ID, peak name, retention index, distribution coefficient, plate number, scan ID, peak area, peak height, signal-to-noise ratio, and half-peak height symmetry.

[0072] The peak ID refers to a unique identifier for identifying and naming a specific peak in a chromatogram and is associated with a specific chromatographic peak.

[0073] Peak name refers to the name or symbol used to identify and name the peaks that appear in the chromatogram, and is usually used to identify and describe characteristic peaks or representative peaks in the chromatogram. Peak names can generally be identified based on the location of the peak on the chromatogram or the special shape of the peak, or can be determined based on the distribution of mass spectrum fragment characteristics. Common naming methods include M+ peak, M-1 peak, base peak, fragment peak, etc. It should be pointed out that the peak name is different from the peak ID. The peak ID is a unique identifier for identifying and naming a specific peak, while the peak name is a description and classification of the peak in chemical language.

[0074] The retention index is the measured value of the retention time of a compound relative to a reference substance under specific chromatographic conditions. Since compounds with similar boiling points and chemical properties have similar retention times on the chromatographic column, the retention time of a compound relative to a reference substance (i.e., the retention index) can be used to distinguish compounds and perform qualitative analysis.

[0075] The distribution coefficient refers to the distribution ratio of a compound between the non-polar stationary phase and the gas phase mobile phase in GC-MS analysis. Generally, the higher the distribution coefficient, the higher the distribution ratio of the compound in the non-polar stationary phase, the longer its retention time in the gas phase, and the corresponding signal intensity in the chromatographic detector will be stronger.

[0076] The plate number refers to the number and characteristics of peaks detected in the sample chromatogram under gas chromatography-mass spectrometry conditions. The plate number is an indicator of the complexity of the sample and the uniformity of the component distribution. The higher the plate number, the higher the separation efficiency, the narrower the peak base between adjacent sample peaks and the easier to identify, and the more obvious the separation of signal peaks.

[0077] The scan ID is a unique number generated by the analytical equipment when scanning a sample. This number generally consists of two parts, one is the instrument type and parameter setting information, and the other is the specific scan number. The scan ID can facilitate the quick search of spectrum-related information, and is also conducive to data management and data processing, making it easy to find and compare data between different samples.

[0078] The signal-to-noise ratio refers to the ratio between the chromatographic signal intensity and the background noise intensity. A higher signal-to-noise ratio can improve the stability and reliability of the mass spectrometry signal, thereby reducing misjudgment and errors.

[0079] Half-peak height symmetry refers to the symmetry between the left and right sides of the chromatographic peak. In chromatographic analysis, a symmetrical chromatographic peak shape usually means that the distribution of sample components is relatively uniform. A symmetrical peak shape makes it easier to accurately measure and calculate the peak area or peak height.

[0080] In practical applications, the extracted key attribute field data can be further used to construct a shale oil saturated hydrocarbon fingerprint database.

[0081] Step 104, based on the retention time and chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, obtain the initial chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample, optimize each initial chromatographic fingerprint time series signal, and obtain the target chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample.

[0082] Among them, based on the retention time and the chromatographic response data, a chromatographic fingerprint response signal, i.e., the initial chromatographic fingerprint time series signal, can be obtained. During the chromatographic separation process, as the sample components are successively distilled out on the chromatographic column, the fixed frequency full scan (Full Scan) detector can output a curve of signal intensity changing with time, so the chromatographic fingerprint response signal is a time series signal. At the same time, the area or height of the chromatographic peak is related to the concentration or mass of the component, so the chromatographic fingerprint response signal meets the basic characteristics of time series data analysis.

[0083] In an embodiment of the present application, step 104 optimizes the processing of each initial chromatographic fingerprint time series signal to obtain a target chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample, which may include: for each initial chromatographic fingerprint time series signal corresponding to a shale oil saturated hydrocarbon sample, performing baseline correction processing on the initial chromatographic fingerprint time series signal to obtain a first chromatographic fingerprint time series signal; performing filtering and noise reduction processing on the first chromatographic fingerprint time series signal to obtain a second chromatographic fingerprint time series signal; performing delay correction processing on the second chromatographic fingerprint time series signal to obtain a target chromatographic fingerprint time series signal.

[0084] In practical applications, there are a large number of cyclized components and isomers in shale oil saturated hydrocarbons. The strong intermolecular interaction makes it difficult for one-dimensional chromatographic columns to separate them, resulting in baseline drift. At the same time, column aging or column loss will also cause the chromatographic fingerprint signal to be in a non-stationary state, thereby causing baseline drift. Therefore, baseline correction can be performed on each initial chromatographic fingerprint time series signal to improve the accuracy and reliability of shale oil saturated hydrocarbon chromatographic fingerprint identification results.

[0085] In an embodiment of the present application, baseline correction processing is performed on the initial chromatographic fingerprint time series signal to obtain a first chromatographic fingerprint time series signal, which may include: using fast wavelet transform to decompose the initial chromatographic fingerprint time series signal respectively to obtain multiple sub-signals corresponding to the initial chromatographic fingerprint time series signal; performing threshold processing on the multiple sub-signals, and reconstructing the multiple sub-signals after threshold processing to obtain the first chromatographic fingerprint time series signal after baseline correction.

[0086] In the above embodiment, the initial chromatographic fingerprint time series signal is decomposed by fast wavelet transform to obtain multiple sub-signals corresponding to the initial chromatographic fingerprint time series signal, which may specifically include: using Daubechies wavelet basis function to decompose the initial chromatographic fingerprint time series signal into components of different frequencies, and performing multi-scale signal stripping, thereby splitting the initial chromatographic fingerprint time series signal into multiple sub-signals with different frequencies and scales. Generally speaking, high-frequency sub-signals contain noise and high-frequency components, and low-frequency sub-signals contain baseline drift and low-frequency components.

[0087] In the above embodiment, the multiple sub-signals are subjected to threshold processing to eliminate noise and baseline drift. Furthermore, the multiple sub-signals subjected to threshold processing are reconstructed to obtain a chromatographic fingerprint time series signal after removing noise and baseline drift.

[0088] In practical applications, the signal-to-noise ratio (S / N) is a key control factor for distinguishing two highly similar shale oils, which directly affects the accuracy of the chromatographic fingerprint recognition results. Therefore, after obtaining the baseline-corrected first chromatographic fingerprint time series signals, the first chromatographic fingerprint time series signals can be further filtered and denoised to further improve the accuracy and reliability of the shale oil saturated hydrocarbon chromatographic fingerprint recognition results.

[0089] In an embodiment of the present application, filtering and denoising the first chromatographic fingerprint time series signal to obtain a second chromatographic fingerprint time series signal may include: fitting the first chromatographic fingerprint time series signal using a local polynomial function within a filtering window to obtain fitting values ​​corresponding to the first chromatographic fingerprint time series signal, and replacing the data corresponding to the first chromatographic fingerprint time series signal with the fitting values ​​to obtain the second chromatographic fingerprint time series signal after filtering and denoising.

[0090] In specific implementation, the Savitzky-Golay local polynomial function can be used to fit the first chromatographic fingerprint time series signal. By filtering and denoising the first chromatographic fingerprint time series signal, the noise can be effectively removed while maintaining the signal shape. According to experiments, after the above filtering and denoising, the signal-to-noise ratio is significantly improved (>10 times); and the fingerprint peaks of the isomeric compounds that were initially invisible are restored very accurately.

[0091] Affected by the gasification efficiency of the injection port or manual injection, there will be a certain solvent delay between sample chromatographic fingerprints. Therefore, the second chromatographic fingerprint time series signal can be subjected to delay correction processing.

[0092] In practical applications, some chromatographic instruments can directly measure and correct solvent delays. These instruments have the function of automatically calculating solvent delays, and users only need to set and calibrate before the experiment begins. When the chromatographic instrument cannot correct the solvent delay, in an embodiment of the present application, the second chromatographic fingerprint time series signal is subjected to delay correction processing. Specifically, the instrument parameter method can be used to perform solvent delay correction on the second chromatographic fingerprint time series signal. When performing solvent delay correction, the consistency and stability of the injection port must be ensured. After the delay correction processing, each target chromatographic fingerprint time series signal is obtained.

[0093] Step 105: construct a shale oil chromatographic fingerprint recognition model based on the K-Shape algorithm and the target chromatographic fingerprint time series signals.

[0094] The K-Shape algorithm is an algorithm for time series clustering that combines the concepts of k-means clustering and shape distance. Unlike traditional distance-based clustering algorithms, the K-Shape algorithm can handle time series data with similar shapes but different amplitudes and offsets.

[0095] The embodiment of the present application is to abstract the chromatographic fingerprint of shale oil saturated hydrocarbons into time series data, and establish an unsupervised clustering chromatographic fingerprint recognition model of shale oil saturated hydrocarbons based on unequal time series. The fingerprint recognition model involves an improved time series K-Shape clustering method, which has a better clustering recognition effect. The K-Shape algorithm performs clustering by comparing the shapes of each chromatographic fingerprint (i.e., peak intensity and retention time) with each other.

[0096] In an embodiment of the present application, step 105, based on the K-Shape algorithm and each target chromatographic fingerprint time series signal, constructing a shale oil saturated hydrocarbon chromatographic fingerprint recognition model, can include: using a dynamic time warping (Dynamic TimeWarping, DTW) algorithm to calculate the distance between the chromatographic fingerprints corresponding to each target chromatographic fingerprint time series signal, based on the distance between each chromatographic fingerprint, using the K-Means algorithm to cluster the time series of each chromatographic fingerprint, and constructing a shale oil chromatographic fingerprint recognition model according to the clustering result.

[0097] Among them, based on the distance between each chromatographic fingerprint, using the K-Means algorithm to cluster the time series of each chromatographic fingerprint can include: taking the distance between each chromatographic fingerprint calculated based on the dynamic time warping algorithm as input, running the K-Means clustering algorithm to cluster the time series, and dividing the time series into k clusters.

[0098] Compared with the traditional clustering algorithm based on Euclidean distance, the algorithm adopted in the above embodiment can ensure scaling invariance, translation invariance and conversion invariance; at the same time, it can also efficiently compare the similarities between fingerprint sequences.

[0099] In specific implementation, in each iteration, the K-Shape algorithm can cyclically execute the following steps ① and ②:

[0100] ① In the assignment step, based on the DTW distance (i.e., the distance between each chromatographic fingerprint calculated by DTW), each chromatographic fingerprint is compared with all calculated centroids, and the chromatographic fingerprint of each sample is assigned to the cluster closest to the centroid to update the membership in the cluster;

[0101] ②Further, update the cluster center to reflect the changes in cluster members in the previous step.

[0102] Repeat steps ① and ② until the cluster membership does not change or the maximum number of allowed iterations (e.g., 100) is reached.

[0103] In other words, the specific iterative process can perform the following steps (1) and (2):

[0104] (1) Similarity measurement: The dynamic time warping algorithm is used to process chromatographic fingerprints of unequal lengths and give the optimal alignment path. The DTW algorithm can match these chromatographic fingerprints by adaptively stretching or compressing the retention time so that they are aligned on the time axis, thereby calculating the similarity between them.

[0105] (2) Unsupervised clustering: K-Means algorithm is used to cluster time series. The final clustering method for the two chromatographic fingerprint sequences is implemented through iteration, and each iteration is divided into two steps: the first step is to recalculate the centroid, and the second step is to reallocate each sequence to different clusters based on its distance from the new centroid; the iteration cycle continues until the label no longer changes.

[0106] In practical applications, in order to further improve the accuracy of the shale oil chromatographic fingerprint recognition model, after obtaining each target chromatographic fingerprint time series signal in step 104, before constructing the shale oil chromatographic fingerprint recognition model based on the K-Shape algorithm and each target chromatographic fingerprint time series signal in step 105, the time series data corresponding to each target chromatographic fingerprint time series signal can be further standardized and normalized.

[0107] In addition, the embodiment of the present application regards the centroid calculation as an optimization problem, the goal of which is to find the minimum value of the sum of squared distances between all other time series in the class. And, the centroid calculation of the time series is based on the DTW distance metric: first, an optimal offset is calculated for all time series in the class, the cluster center obtained by the previous calculation is used as a reference, and all sequences are aligned with this reference sequence for iterative clustering.

[0108] It can be understood that the shale oil chromatographic fingerprint recognition model construction method provided by the embodiment of the present application is adopted, by obtaining multiple shale oil saturated hydrocarbon samples; using gas chromatography-mass spectrometry technology to detect each shale oil saturated hydrocarbon sample, and obtaining initial detection data corresponding to each shale oil saturated hydrocarbon sample; extracting key attribute field data from each initial detection data, and the key attribute field data at least includes retention time and chromatographic response data; based on the retention time and chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, the initial chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample is obtained, and each initial chromatographic fingerprint time series signal is optimized to obtain the target chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample; based on the K-Shape algorithm and each target chromatographic fingerprint time series signal, a shale oil chromatographic fingerprint recognition model is constructed. Based on the scheme provided by the embodiment of the present application, a shale oil chromatographic fingerprint recognition model can be constructed. Since the shale oil chromatographic fingerprint recognition model is constructed based on the K-Shape algorithm with better clustering effect, the shale oil chromatographic fingerprint can be accurately identified based on the shale oil chromatographic fingerprint recognition model.

[0109] Considering that unsupervised clustering learning models are often based on data sets that are not given clear labels, they mine and identify special structures, characteristic patterns or causal relationships hidden in the data set by clustering and grouping the original data. Therefore, in order to further improve the accuracy and robustness of the shale oil chromatographic fingerprint recognition model, after the shale oil chromatographic fingerprint recognition model is constructed in step 105, the shale oil chromatographic fingerprint recognition model construction method provided in the embodiment of the present application also includes: verifying the shale oil chromatographic fingerprint recognition model, and optimizing the shale oil chromatographic fingerprint recognition model according to the verification result.

[0110] In practical applications, expert verification, mechanism model and cross-validation can be used to evaluate the clustering results and verify the shale oil chromatographic fingerprint recognition model.

[0111] Among them, when using expert verification to evaluate the clustering results, the following aspects can be evaluated: ① The rationality of clustering: whether each cluster effectively contains data points with similar characteristics. ② The discrimination of clustering: the appropriate clustering algorithm and clustering parameters can effectively ensure that the differences between clusters are significant enough. ③ The number of clusters: according to the specific feature distribution of the training set sample data, combined with the specific task requirements, the optimal number of clusters is analyzed to decide whether it is necessary to merge or refine certain clusters. ④ Interpretation of results: Since the clustering problem itself is highly exploratory, the results of unsupervised clustering learning often exceed the dimensions of existing cognition, and the conclusions obtained are usually difficult to explain using existing technical means, but can be used as a preprocessing step or sub-process of other technologies, so as to be comprehensively reflected in the accuracy of the upper-level supervised learning tasks. Through expert verification, the reliability and effectiveness of clustering results can be improved, and it is helpful to determine the best clustering model and parameter settings.

[0112] Mechanism models usually have clear causal relationships and interpretability. Systematic sampling and a small amount of group experimental analysis can be carried out in the shale oil saturated hydrocarbon chromatographic fingerprint training set to effectively test the accuracy of the algorithm model from a falsifiable perspective. By comparison, the K-Shape algorithm provided in the embodiment of the present application, which has been modified using the dynamic time warping algorithm, is highly consistent with the actual process of the mechanism model in terms of algorithmic principles, which makes the mechanism model test a key verification method. In the embodiment of the present application, two mechanism models, biomarker compound analysis and chemometric analysis, can be used to verify the shale oil chromatographic fingerprint recognition model.

[0113] The use of cross-validation method in unsupervised learning can evaluate the performance of shale oil chromatographic fingerprint recognition model to a certain extent and provide a reference for its parameter setting. In specific implementation, the optimal number of clusters can be determined by calculating the silhouette coefficient method or using the elbow method. Furthermore, the sum of squares of the distance from each chromatographic fingerprint to the center of the cluster can be calculated, designated as the sum of squares of intra-cluster errors, and an elbow line graph can be drawn to assist in checking whether the number of clusters is reasonable; at the same time, the overlay effect diagram of different fingerprints is drawn to explore the optimal number of clusters in a heuristic search manner.

[0114] It can be understood that after verifying and optimizing the shale oil chromatographic fingerprint recognition model in the above manner, a shale oil chromatographic fingerprint recognition model with higher accuracy, higher recognition efficiency, stronger generalization ability, better robustness and faster speed can be obtained.

[0115] The chromatographic fingerprint recognition method in the prior art usually requires that the chromatographic analysis conditions be as consistent as possible while ensuring the precision and accuracy of the chromatograph. In order to ensure good separation reproducibility, it is generally necessary to use the same detector and keep the same injection volume. The inconsistency of retention values ​​caused by slight differences in experimental conditions will make the subsequent peak identification, peak detection and peak alignment extremely difficult, and rely on absolute qualitative and quantitative analysis of internal and external standards. The shale oil chromatographic fingerprint recognition model provided in the embodiment of the present application is used to identify shale oil saturated hydrocarbon samples, without the need for internal and external standard quantification, and is not affected by the instrument type and chromatographic conditions.

[0116] The saturated hydrocarbon chromatographic fingerprint recognition method in the prior art is basically unable to go beyond the five classifications in the classification of crude oil types, and it is difficult to meet the needs of refined classification comparison. The shale oil saturated hydrocarbon chromatographic fingerprint recognition model provided in the embodiment of the present application compares the shapes of each chromatographic fingerprint with each other and clusters them according to the shapes, so that the chromatographic fingerprint can be more finely divided, thereby meeting the needs of refined classification comparison of crude oil types.

[0117] The saturated hydrocarbon chromatographic fingerprint recognition method in the prior art consumes a lot of time and manpower and material resources in the process of analysis, testing, data processing and recognition comparison, and has low efficiency; and it is impossible to compare a large number of samples with each other. However, the shale oil saturated hydrocarbon chromatographic fingerprint recognition model provided in the embodiment of the present application can compare a large number of samples with each other and greatly improve the recognition efficiency.

[0118] In addition, based on the shale oil saturated hydrocarbon chromatographic fingerprint recognition model provided in the embodiment of the present application, it is possible to scientifically, accurately and efficiently identify different types of shale oil through chromatographic fingerprint time series clustering, which has important guiding significance for the current fine exploration and efficient development of unconventional oil and gas. Combined with the shale oil saturated hydrocarbon fingerprint database, the above-established shale oil saturated hydrocarbon chromatographic fingerprint recognition model can also be applied to actual scenarios such as exploration, development and production to improve the recognition efficiency and accuracy in various scenarios. Further, the chromatographic fingerprint recognition results of the shale oil saturated hydrocarbon chromatographic fingerprint recognition model can also be combined with the source rock parent material type, thermal maturity, sedimentary environment and microbial degradation to achieve a detailed description of the characteristics of hydrocarbon fluids in tight reservoirs. In addition, the shale oil saturated hydrocarbon chromatographic fingerprint recognition model provided in the embodiment of the present application can also be promoted and applied to the development of fracturing fields, such as fracturing evaluation, well-to-well interference, sweet spot identification, production capacity prediction, etc. In actual promotion and application, the parameters of the shale oil saturated hydrocarbon chromatographic fingerprint recognition model can be adjusted according to the adjustment and optimization suggestions fed back by experts to match the needs of the actual application scenario.

[0119] Example 2

[0120] Based on the shale oil chromatographic fingerprint recognition model construction method provided in the above embodiments of the present application, the embodiments of the present application also provide a shale oil chromatographic fingerprint recognition model, and the shale oil chromatographic fingerprint recognition model can be constructed by the shale oil chromatographic fingerprint recognition model construction method provided in any of the above embodiments of the present application.

[0121] It can be understood that the shale oil chromatographic fingerprint recognition model provided in the embodiment of the present application is adopted. Since the shale oil chromatographic fingerprint recognition model is constructed based on the K-Shape algorithm with better clustering effect, the shale oil saturated hydrocarbon chromatographic fingerprint can be accurately identified based on the shale oil chromatographic fingerprint recognition model.

[0122] Example 3

[0123] Based on the shale oil chromatographic fingerprint recognition model provided in the above embodiment of the present application, the embodiment of the present application also provides a shale oil chromatographic fingerprint recognition method. The shale oil chromatographic fingerprint recognition method can use the shale oil chromatographic fingerprint recognition model provided in the above embodiment to identify the shale oil saturated hydrocarbon chromatographic fingerprint. The shale oil chromatographic fingerprint recognition method may include the following steps:

[0124] Step 1: Obtain the shale oil saturated hydrocarbon sample to be tested.

[0125] Among them, the method for obtaining the shale oil saturated hydrocarbon sample to be tested can refer to the above content and will not be repeated here.

[0126] Step 2: Use gas chromatography-mass spectrometry to detect the shale oil saturated hydrocarbon sample to obtain initial detection data corresponding to the shale oil saturated hydrocarbon sample to be tested.

[0127] Step three, based on the retention time and chromatographic response data corresponding to the initial detection data, obtain the initial chromatographic fingerprint time series signal corresponding to the initial detection data, optimize the initial chromatographic fingerprint time series signal, and obtain the chromatographic fingerprint time series signal to be identified corresponding to the shale oil saturated hydrocarbon sample to be tested.

[0128] The optimization process for the initial chromatographic fingerprint time series signal may include baseline correction, filtering and noise reduction, and delay correction. For details, please refer to the above content and will not be described in detail here.

[0129] Step 4: Identify the chromatographic fingerprint time series signal to be identified based on the shale oil chromatographic fingerprint identification model.

[0130] It can be understood that the shale oil chromatographic fingerprint recognition method provided in the embodiment of the present application can accurately identify the chromatographic fingerprint time series signal to be identified based on the shale oil chromatographic fingerprint recognition model, and the shale oil chromatographic fingerprint recognition model is constructed based on the K-Shape algorithm with better clustering effect.

[0131] Example 4

[0132] The shale oil chromatographic fingerprint recognition model, model construction method and recognition method provided in the embodiments of the present application are further explained below with reference to specific examples.

[0133] Taking a shale oil block in eastern China as an example, samples of wellhead produced fluid (oil-water mixture) from 22 wells in the same formation were collected at the oil field site and placed in 250mL sample collection bottles. In the laboratory, 50mL of sample was transferred from each sample collection bottle into a pear-shaped separatory funnel, and dichloromethane (analytical grade) solvent was used for extraction to obtain the oil components corresponding to the wellhead produced fluids of the 22 wells, and a silica gel-alumina packed column was used to separate the oil components into groups, and the saturated hydrocarbon components of shale oil corresponding to the wellhead produced fluids of the 22 wells were obtained.

[0134] Each shale oil saturated hydrocarbon sample was detected and analyzed using a gas chromatography-mass spectrometer to obtain each initial detection data corresponding to each shale oil saturated hydrocarbon sample. The platform uses a combination of an HP6890N gas chromatograph and an HP5973N mass spectrometer. The chromatographic column in the gas chromatograph can be an HP-5 elastic quartz capillary column (size 30m×0.25mm×0.25μm); the column box temperature can be heated from 80°C to 290°C, the heating rate is 4°C / min, and the temperature is maintained at 80°C for 2min and at 290°C for 30min; the carrier gas can be high-purity helium (purity 99.999%). The EI ion source temperature in the single quadrupole mass spectrometer can be 280°C, the ionization energy can be 70eV, dichloromethane (chromatographic grade) is used as the solvent, and the manual injector can have a splitless injection volume of 2μL.

[0135] The MassHunter workstation software was used to interpret the initial detection data obtained, and the Nist17 spectral library was used to qualitatively and quantitatively identify the characteristic compounds. The identification results were submitted to the laboratory big data platform, and the platform interactive tools were used to extract the key attribute field data in the data table to obtain the peak name, peak ID, retention time, chromatographic response data, peak area, peak height, etc., to form a shale oil saturated hydrocarbon fingerprint database with a complete data structure.

[0136] Based on the retention time and chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, the initial chromatographic fingerprint time series signals corresponding to each shale oil saturated hydrocarbon sample are obtained, and each initial chromatographic fingerprint time series signal is optimized. The optimization process includes: first, the initial chromatographic fingerprint time series signal is split into multiple sub-signals with different frequencies and scales by fast wavelet analysis; the multiple sub-signals obtained by decomposition are threshold processed to eliminate noise and baseline drift; then, the multiple sub-signals after threshold processing are reconstructed to obtain the chromatographic fingerprint time series signal after baseline correction. Then, the chromatographic fingerprint time series signal after baseline correction is fitted using a local polynomial function, and the fitting value with a symmetrical waveform and a signal-to-noise ratio S / N>2 is set to replace the original value in the chromatographic fingerprint time series signal after baseline correction to remove the background solvent peak interference. Finally, the peak time of the specified compound of 22 samples is normalized and aligned using the solvent delay time. After the above processing, each target chromatographic fingerprint time series signal is obtained.

[0137] like Figure 2 As shown, Figure 2The diagram is a schematic diagram of the similarity measurement and optimal alignment path of the chromatographic fingerprint of shale oil saturated hydrocarbons. Among them, S201 indicates that the collected samples are tested to obtain the total ion current (TIC) of shale oil saturated hydrocarbons, including chromatographic retention time and chromatographic response data; S202 indicates that the baseline correction processing, filtering and noise reduction processing and delay correction processing are performed on each initial chromatographic fingerprint time series signal.

[0138] The dynamic time warping algorithm is used to calculate the distances between the chromatographic fingerprints corresponding to the time series signals of each target chromatographic fingerprint. Based on the distances between the chromatographic fingerprints, the K-Means algorithm is used to cluster the chromatographic fingerprints in time series. According to the clustering results, a shale oil chromatographic fingerprint recognition model is constructed.

[0139] like Figure 2 As shown, S203 indicates using the dynamic time warping method to calculate the distances between the shale oil saturated hydrocarbon chromatographic fingerprints of 22 wells and the optimal alignment path.

[0140] like Figure 3 As shown, Figure 3 The figure is a schematic diagram of the unsupervised clustering identification process of shale oil saturated hydrocarbon chromatographic fingerprint. S301 represents the input of shale oil saturated hydrocarbon chromatographic fingerprint into the model; S302 represents the output of the training k-Means unsupervised clustering model. During training, the shale oil saturated hydrocarbon chromatographic fingerprint is clustered according to k equal to two, three, four, and five clusters, and the chromatographic fingerprint corresponding to the cluster centroid is calculated through loop iteration.

[0141] As shown in Table 1, the chromatographic fingerprint recognition results of shale oil saturated hydrocarbons in 22 wells in a block in eastern China are shown. Obviously, the chromatographic fingerprint recognition method based on time series clustering identified shale oil saturated hydrocarbons according to two clusters, three clusters, four clusters and five clusters respectively.

[0142] The results of the two-cluster identification in Table 1 are: A1 cluster has 17 wells, A2 cluster has 5 wells. The results of the three-cluster identification are: B1 cluster has 5 wells, B2 cluster has 10 wells, and B3 cluster has 7 wells. The results of the four-cluster identification are: C1 cluster has 1 well, C2 cluster has 2 wells, C3 cluster has 3 wells, and C4 cluster has 16 wells. The results of the five-cluster identification are: X1 cluster has 4 wells, X2 cluster has 3 wells, X3 cluster has 1 well, X4 cluster has 12 wells, and X5 cluster has 2 wells.

[0143] At the same time, it can be seen from Table 1 that compared with the hierarchical classification results based on expert experience and chemometrics, the optimized shale oil saturated hydrocarbon chromatographic fingerprint identification method is highly consistent with the actual process of the mechanism model in terms of algorithm principle, and the patterns and structures learned from the data are highly logically consistent with the processes and phenomena in the real world.

[0144] Table 1 Chromatographic fingerprint identification results of saturated hydrocarbons in shale oil from a block in eastern China

[0145]

[0146] In Table 1, in the expert experience chemometric analysis, II represents the binary classification result; III represents the ternary classification result; VI represents the quaternary classification result. In the time series clustering chromatographic fingerprint recognition, A represents the two-cluster recognition result; B represents the three-cluster recognition result; C represents the four-cluster recognition result; X represents the five-cluster and super-five-cluster recognition results.

[0147] As shown in Table 2, after the experts optimized and verified the model, the results are as follows: ① In the two-cluster clustering, the chromatographic fingerprint identification of shale oil saturated hydrocarbons clearly and accurately identified the significant characteristics of shale oil in 22 wells in the region under biodegradation and thermal degradation; ② Based on the two-cluster clustering, the three-cluster clustering maintained the recognition of the chromatographic fingerprint characteristics of shale oil saturated hydrocarbons of the biodegradation type, and further refined and identified the fingerprint characteristics of shale oil after different degrees of thermal degradation; ③ In the four-cluster clustering, the degree of biodegradation was also well distinguished. After the category refinement, the thermal evolution characteristics of shale oil were identified, a typical front peak type with light carbonaceous components (main peak carbon C 14 ) were identified; ④ In the process of five-cluster clustering and super-five-cluster clustering, the information of normal alkanes, isoalkanes and isoprenoid alkanes in the TIC of shale oil saturated hydrocarbons were fully mined, and special shale oils could be identified under the effects of biodegradation and thermal evolution at the same time.

[0148] Table 2 Identification results of saturated hydrocarbons in shale oil from a block in eastern China by chromatographic fingerprint system

[0149]

[0150] The results of the two-cluster identification of shale oil saturated hydrocarbon chromatographic fingerprints can be shown as follows Figure 4-1 As shown. Figure 4-1 Among them, 411 is the chromatographic thermal degradation fingerprint feature of shale oil saturated hydrocarbons; 412 is the chromatographic biodegradation fingerprint feature of shale oil saturated hydrocarbons.

[0151] The three-cluster identification results of shale oil saturated hydrocarbon chromatographic fingerprint can be shown as follows Figure 4-2 As shown. Figure 4-2 Among them, 421 is the fingerprint characteristic of chromatographic biodegradation of shale oil saturated hydrocarbons; 422 is the fingerprint characteristic of high degree of chromatographic thermal degradation of shale oil saturated hydrocarbons; 423 is the fingerprint characteristic of low degree of chromatographic thermal degradation of shale oil saturated hydrocarbons.

[0152] The four-cluster identification results of shale oil saturated hydrocarbon chromatographic fingerprint can be shown as follows Figure 4-3 As shown. Figure 4-3Among them, 431 is the special light shale oil with a front peak type shown in the shale oil saturated hydrocarbon chromatogram after strong thermal evolution; 432 is the weak fingerprint feature of the shale oil saturated hydrocarbon chromatogram biodegradation degree; 433 is the strong fingerprint feature of the shale oil saturated hydrocarbon chromatogram biodegradation degree; 434 is the fingerprint feature of the shale oil saturated hydrocarbon chromatogram affected by the thermal evolution characteristics.

[0153] The five-cluster identification results of shale oil saturated hydrocarbon chromatographic fingerprint can be shown as follows Figure 4-4 As shown. Figure 4-4 Among them, 441 is the fingerprint feature of high degree of thermal evolution of shale oil saturated hydrocarbon chromatography; 442 is the weak fingerprint feature of shale oil saturated hydrocarbon chromatography biodegradation degree; 443 is the special light shale oil with front peak type shown in shale oil saturated hydrocarbon chromatography after strong thermal evolution; 444 is the fingerprint feature of low degree of thermal evolution of shale oil saturated hydrocarbon chromatography; 445 is the strong fingerprint feature of shale oil saturated hydrocarbon chromatography biodegradation degree.

[0154] like Figure 5 As shown, it is a similarity inheritance relationship diagram of shale oil saturated hydrocarbon chromatographic fingerprint identification provided in the embodiment of the present application. Obviously, the shale oil saturated hydrocarbons of the A2 branch and the A1 branch have two completely independent classification systems in terms of fingerprint feature similarity. Among the shale oils of the 22 wells identified in this embodiment, A1 passes through B2, B3, C4 and then X4, which is the main inheritance evolution path. Therefore, under the dominance of thermal factors, further in-depth refinement of the classification of chromatographic fingerprints is sufficient to solve the key technical problems of fine identification and comparison of shale oil in tight reservoirs, thereby achieving higher-precision fluid tracing.

[0155] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0156] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for constructing a shale oil chromatographic fingerprint recognition model. It is characterized in that The method comprises: Obtain multiple shale oil saturated hydrocarbon samples; Utilizing gas chromatography-mass spectrometry to detect each shale oil saturated hydrocarbon sample, and obtaining initial detection data corresponding to each shale oil saturated hydrocarbon sample; Extract key attribute field data from each initial detection data, wherein the key attribute field data at least includes retention time and chromatographic response data; Based on the retention time and chromatographic response data corresponding to each shale oil saturated hydrocarbon sample, the initial chromatographic fingerprint time series signals corresponding to each shale oil saturated hydrocarbon sample are obtained, and each initial chromatographic fingerprint time series signal is optimized to obtain the target chromatographic fingerprint time series signals corresponding to each shale oil saturated hydrocarbon sample; Based on the K-Shape algorithm and the time series signals of each target chromatographic fingerprint, a shale oil chromatographic fingerprint recognition model was constructed.

2. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 1, It is characterized in that The method of obtaining a plurality of shale oil saturated hydrocarbon samples comprises: Based on the wellhead produced fluid and / or shale rock samples, multiple shale oil samples are obtained, and multiple shale oil saturated hydrocarbon samples are obtained from the multiple shale oil samples, and one shale oil sample corresponds to one shale oil saturated hydrocarbon sample.

3. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 1, It is characterized in that The key attribute field data also includes peak ID, peak name, retention index, distribution coefficient, plate number, scan ID, peak area, peak height, signal-to-noise ratio and half-peak height symmetry.

4. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 1, It is characterized in that The optimization process of each initial chromatographic fingerprint time series signal to obtain a target chromatographic fingerprint time series signal corresponding to each shale oil saturated hydrocarbon sample includes: Performing baseline correction processing on the initial chromatographic fingerprint time series signal to obtain a first chromatographic fingerprint time series signal; Performing filtering and noise reduction processing on the first chromatographic fingerprint time series signal to obtain a second chromatographic fingerprint time series signal; The second chromatographic fingerprint time series signal is subjected to delay correction processing to obtain a target chromatographic fingerprint time series signal.

5. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 4, It is characterized in that The method of performing baseline correction processing on the initial chromatographic fingerprint time series signal to obtain a first chromatographic fingerprint time series signal includes: Using fast wavelet transform, the initial chromatographic fingerprint time series signal is decomposed to obtain multiple sub-signals corresponding to the initial chromatographic fingerprint time series signal; Threshold processing is performed on the multiple sub-signals, and the multiple sub-signals after the threshold processing are reconstructed to obtain a first chromatographic fingerprint time series signal.

6. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 4, It is characterized in that The filtering and noise reduction process is performed on the first chromatographic fingerprint time series signal to obtain the second chromatographic fingerprint time series signal, including: Using a local polynomial function to fit the first chromatographic fingerprint time series signal within the filtering window to obtain a fitting value corresponding to the first chromatographic fingerprint time series signal; The data corresponding to the first chromatographic fingerprint time series signal is replaced with the fitting value to obtain the second chromatographic fingerprint time series signal.

7. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 1, It is characterized in that The shale oil chromatographic fingerprint recognition model is constructed based on the K-Shape algorithm and each target chromatographic fingerprint time series signal, including: The distances between the chromatographic fingerprints corresponding to the target chromatographic fingerprint time series signals are calculated using a dynamic time warping algorithm; Based on the distance between each chromatographic fingerprint, the K-Means algorithm is used to cluster the time series of each chromatographic fingerprint, and a shale oil chromatographic fingerprint recognition model is constructed according to the clustering results.

8. The method for constructing a shale oil chromatographic fingerprint recognition model according to claim 1, It is characterized in that After the shale oil chromatographic fingerprint recognition model is constructed, the method further comprises: The shale oil chromatographic fingerprint recognition model is verified, and the shale oil chromatographic fingerprint recognition model is optimized according to the verification result.

9. A shale oil chromatographic fingerprint recognition model, It is characterized in that The shale oil chromatographic fingerprint recognition model is constructed by the shale oil chromatographic fingerprint recognition model construction method described in any one of claims 1-8.

10. A shale oil chromatographic fingerprint identification method, It is characterized in that The shale oil chromatographic fingerprint identification method uses the shale oil chromatographic fingerprint identification model of claim 9 to perform shale oil chromatographic fingerprint identification, and the shale oil chromatographic fingerprint identification method comprises: Obtaining a saturated hydrocarbon sample of shale oil to be tested; Detecting the shale oil saturated hydrocarbon sample to be tested by using gas chromatography-mass spectrometry to obtain initial detection data corresponding to the shale oil saturated hydrocarbon sample to be tested; Based on the retention time and chromatographic response data corresponding to the initial detection data, an initial chromatographic fingerprint time series signal corresponding to the initial detection data is obtained, and the initial chromatographic fingerprint time series signal is optimized to obtain a chromatographic fingerprint time series signal to be identified corresponding to the shale oil saturated hydrocarbon sample to be tested; The chromatographic fingerprint time series signal to be identified is identified based on the shale oil chromatographic fingerprint identification model.

Citation Information

Cited By

  • Method and device for analyzing content of ions in original formation water of shale

    CN121347716A

  • A method and apparatus for analyzing ion content in shale formation water

    CN121347716B