A method for rapid identification of marine oil spill source based on three-dimensional fluorescence spectral data
Patent Information
- Application Number
- CN202610623948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-18
AI Technical Summary
本发明提供一种海上溢油来源的快速识别方法,以解决传统方法无法在满足识别准确率要求的同时,有着较为快速的识别时间的问题
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine environmental monitoring, specifically a method for rapid detection and identification of marine oil spills. Background Technology
[0002] Marine oil spills generally refer to accidents that occur during the offshore extraction, transportation, loading, unloading, and use of oil. As a typical marine disaster characterized by its suddenness, wide-ranging impact, and difficulty in mitigation, marine oil spills pose a serious threat to the marine ecological environment and human society. Furthermore, oil spills also have a significant impact on coastal economic activities. Oil pollution not only severely affects aquaculture and marine fisheries but also causes a chain reaction affecting industries such as tourism, port logistics, and maritime transportation.
[0003] Existing marine oil spill detection and identification technologies generally include infrared spectroscopy, excitation spectroscopy, emission spectroscopy, and gas chromatography-mass spectrometry. However, most of these methods cannot meet the requirements of both high accuracy and fast identification time. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a rapid identification method for the source of marine oil spills, solving the problem that traditional methods cannot meet the requirements of identification accuracy while also having a relatively fast identification time.
[0005] To address the aforementioned technical problems, this invention proposes a rapid identification method for marine oil spill sources based on three-dimensional fluorescence spectroscopy data. First, an oil sample classification model is constructed, and then this model is used to rapidly identify the sampled oil spill sources. The specific steps are as follows:
[0006] Step 1: Construct a petroleum sample classification model, including: measuring the three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of oil samples from known offshore oil fields; establishing a category table representing the source of the oil sample based on the three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of the oil well samples, wherein the category table includes category number, oil field name, platform name, and oil well number; for each oil sample, preprocessing the three-dimensional fluorescence spectrum data using the ultraviolet fluorescence spectrum data, including: removal of Rayleigh and Raman scattering, conformal interpolation of missing data, correction of internal filtering effect, and norm normalization; the ultraviolet fluorescence spectrum data is used to calculate the internal filtering effect correction factor, and the correction factor is used to recover the true fluorescence intensity of the three-dimensional fluorescence spectrum of the petroleum sample, thereby obtaining the standardized three-dimensional fluorescence spectrum data of the oil sample; extracting 59 fluorescence index data from the standardized three-dimensional fluorescence spectrum data of all oil samples and performing Z-score standardization; importing the Z-score standardized fluorescence index data into the random forest algorithm and mapping it one-to-one with the category number, thereby obtaining the trained petroleum sample classification model.
[0007] Step 2: Use the petroleum sample classification model constructed above to identify the category and obtain the source of the petroleum sample according to the category table.
[0008] Furthermore, in the rapid identification method for the source of marine oil spills described in this invention, wherein:
[0009] In step 1, the excitation wavelength Ex of the three-dimensional fluorescence spectrum is set to 200-580 nm with a scan step size of 5 nm; the emission wavelength is set to 250-580 nm with a scan step size of 5 nm; and the ultraviolet fluorescence spectrum wavelength is set to 200-580 nm with a scan step size of 5 nm.
[0010] In step 1, the 59 fluorescence index data include regional fluorescence integral, classic characteristic fluorescence peaks, key optical ratio index, regional texture and statistical features, global main peak location features, and global matrix distribution features.
[0011] Compared with the prior art, the beneficial effects of the present invention are:
[0012] Since existing methods cannot simultaneously achieve both high accuracy and high identification speed, this invention provides a rapid identification method for the source of marine oil spills. Based on three-dimensional fluorescence spectroscopy, it extracts specific fluorescence indicators and combines them with a random forest algorithm to achieve rapid classification and identification of marine oil spills. The identification time is in the minute range and the accuracy is over 80%. Detailed Implementation
[0013] A rapid method for identifying the source of an oil spill at sea, comprising the following steps:
[0014] S1. The three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of 88 oil samples collected from known oil fields were measured. The excitation wavelength Ex of the three-dimensional fluorescence spectrum was set to 200-580nm and the scanning step size was 5nm; the emission wavelength was set to 250-580nm and the scanning step size was 5nm; the ultraviolet fluorescence spectrum wavelength was set to 200-580nm and the scanning step size was 5nm. In this embodiment, a category table representing the source of the oil sample was established based on the three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of the 88 oil samples. The data in the oil sample source category table include category number, oil field name, platform name and well number, as shown in Table 1 (including continued Table 1(1) to continued Table 1(3)).
[0015] Table 1. Oil Sample Source Categories
[0016]
[0017] Continued from Table 1 (1)
[0018]
[0019] Continued from Table 1 (2)
[0020]
[0021] Continued from Table 1 (3)
[0022]
[0023] S2, For each oil sample, the three-dimensional fluorescence spectral data is preprocessed using the ultraviolet fluorescence spectral data, including the removal of Rayleigh and Raman scattering, conformal interpolation of missing data, correction of internal filtration effect, and norm normalization. The ultraviolet fluorescence spectral data is used to calculate the internal filtration effect correction factor to obtain the standardized three-dimensional fluorescence spectral data of the oil sample.
[0024] S3. From the standardized three-dimensional fluorescence spectral data of all oil samples, the standardized three-dimensional fluorescence spectral data were analyzed, and 59 fluorescence index data were extracted as features for classification and identification. These included regional fluorescence integral, classical characteristic fluorescence peak, key optical ratio index, regional texture and statistical features, global main peak location features, and global matrix distribution features.
[0025] Fluorescence indices 1-49 represent physical statistics such as peaks, ratios, and regional integrals; fluorescence indices 50-59 represent mathematical statistics. Details are as follows:
[0026] 3-1. Fluorescence index 1-5: Fluorescence region integral.
[0027] The three-dimensional fluorescence spectral matrix (EEM) was divided into five regions, and the average values within each region were calculated. The excitation wavelength (Ex) and emission wavelength (Em) ranges for the five regions are shown below:
[0028] Region I: Ex 200-250 nm, Em 300-330 nm;
[0029] Region II: Ex 200-250 nm, Em 330-380 nm;
[0030] Region III: Ex 200-250 nm, Em 250-300 nm;
[0031] Region IV: Ex 250-400 nm, Em 250-400 nm;
[0032] Region V: Ex 250-400 nm, Em 400-500 nm.
[0033] Fluorescence indices 1-5 (FI_I to FI_V): the absolute average fluorescence intensity of the above 5 regions.
[0034] 3-2. Fluorescence index 6-10: Fluorescence region integral percentage.
[0035] The EEM matrix was divided into 5 regions, and the average value within each region was calculated. The excitation wavelength (Ex) and emission wavelength (Em) ranges for the 5 regions are shown below:
[0036] Region I: Ex 200-250 nm, Em 300-330 nm
[0037] Region II: Ex 200-250 nm, Em 330-380 nm
[0038] Region III: Ex 200-250 nm, Em 250-300 nm
[0039] Region IV: Ex 250-400 nm, Em 250-400 nm
[0040] Region V: Ex 250-400 nm, Em 400-500 nm
[0041] Fluorescence index 6-10 (f_I to f_V): the proportion of the above 5 regions in the overall fluorescence (i.e., FIi / ∑FI).
[0042] 3-3. Fluorescence index 11-14: Characteristic peak value.
[0043] Fluorescence index 11 (Peak_A): Maximum value in the region Ex=260 nm, Em=400-460 nm. Represents a humic acid-like peak.
[0044] Fluorescence index 12 (Peak_B): Fixed-point values of Ex=275 nm, Em=305 nm. Represents tyrosine-like peaks.
[0045] Fluorescence index 13 (Peak_N): Fixed-point values at Ex=280 nm and Em=370 nm. Represents the peaks of soluble microbial metabolites.
[0046] Fluorescence index 14 (Peak_T): Fixed-point values of Ex=275 nm, Em=340 nm. Represents the tryptophan-like peak.
[0047] 3-4. Fluorescence index 15-24: Fluorescence ratio index.
[0048] Fluorescence index 15 (Peak TC): I(275, 350) / I(330, 420). Humic acid ratio, used to distinguish the source of humic substances.
[0049] Fluorescence index 16 (HIX) syn : I(390, 408) / I(355, 373). Synchronous humification index, used to assess the maturity of organic matter.
[0050] Fluorescence index 17 ( ): I(254-435, 480) / I(254-300, 345). Emission humification index, f is used to assess the degree of humification.
[0051] Fluorescence index 18 (BIX): I(310, 380) / I(310, 430). This index distinguishes between biodegradation and petroleum pyrolysis origins.
[0052] Fluorescence index 19 (fi): I(370, 450) / I(370, 500). Fluorescence index, used to distinguish the source of humic substances.
[0053] Fluorescence index 20 (r1): I(250, 365) / I(250, 420). Protein-like / humic acid-like ratio.
[0054] Fluorescence index 21 (r2): I(300, 400) / I(350, 450). Content of medium molecular weight substances.
[0055] Fluorescence index 22 (r3): I(275, 310) / I(275, 340). Tyrosine / tryptophan ratio.
[0056] Fluorescence index 23(r4): I(340, 430) / I(340, 480). Blue-green light ratio, molecular size.
[0057] Fluorescence index 24 (r5): I(360, 450) / I(360, 520). Degree of organic matter degradation.
[0058] 3-5. Fluorescence index 25-49: Regional statistics.
[0059] For machine learning models such as random forests, the average value alone is insufficient; the "distribution pattern" of the data is also necessary. In the aforementioned five regions, five statistical measures were extracted, resulting in a total of 5 × 5 = 25 features, as detailed below:
[0060] Fluorescence index 25-29 (FI_I_std-FI_V_std): Regional standard deviation. Reflects the smoothness or heterogeneity of the spectrum in that region.
[0061] Fluorescence index 30-34 (FI_I_max-FI_V_max): Regional maximum value. Captures localized strong fluorescence peaks.
[0062] Fluorescence index 35-39 (FI_I_min-FI_V_min): Regional minimum. Captures baseline noise level.
[0063] Fluorescence index 40-44 (FI_I_sum-FI_V_sum): Regional summation. A more sensitive total index than the average.
[0064] Fluorescence index 45-49 (FI_I_range-FI_V_range): Regional range (maximum value - minimum value). Reflects the dynamic range of the spectrum.
[0065] Traditional methods only use the mean, but machine learning relies heavily on multi-scale statistics. The range and standard deviation can effectively distinguish between "broad and gentle heavy oil peaks" and "sharp and towering light oil peaks".
[0066] 3-6. Fluorescence index 50-54: Main peak position.
[0067] Fluorescence index 50 (max_ex_idx): The excitation wavelength index corresponding to the global maximum fluorescence intensity.
[0068] Fluorescence index 51 (max_em_idx): The emission wavelength index corresponding to the global maximum fluorescence intensity.
[0069] Fluorescence index 52 (max_ex_wavelength): The actual excitation wavelength of the maximum peak.
[0070] Fluorescence index 53 (max_em_wavelength): The actual emission wavelength of the maximum peak.
[0071] Fluorescence index 54 (max_intensity): The global maximum fluorescence intensity value.
[0072] The most direct way to classify petroleum is by where the strongest fluorescence peak is located. The peak of light oil (low molecular weight) is usually in the low wavelength region (blue shift), while the peak of heavy oil (high molecular weight, high degree of condensation) is red-shifted to the high wavelength region.
[0073] 3-7. Fluorescence index 55-59: Matrix statistics.
[0074] Fluorescence index 55 (matrix_mean): mean fluorescence intensity of the full fluorescence matrix.
[0075] Fluorescence index 56 (matrix_std): Standard deviation of fluorescence intensity of the full fluorescence matrix.
[0076] Fluorescence index 57 (matrix_median): Median fluorescence intensity of the full fluorescence matrix (strong anti-interference ability).
[0077] Fluorescence index 58 (matrix_skew): Skewness. Reflects the tailing characteristics of fluorescence distribution (highly affected by Raman / Rayleigh scattering artifacts).
[0078] Fluorescence index 59 (matrix_kurtosis): Kurtosis (subtract 3 to get excess kurtosis). Reflects the sharpness of the fluorescence peak.
[0079] This set of features is primarily used to assess data quality (such as whether there is severe scattering that has not been subtracted), while also providing a mathematical description of the overall fluorescence "saturation" of different oil products.
[0080] S4. The Z-score normalization is performed on the 59 fluorescence index values extracted above.
[0081] X_norm=(X-μ) / σ
[0082] Where X is the initial eigenvalue; μ is the global mean; σ is the global standard deviation; and X_norm is the standardized eigenvalue.
[0083] S5. The fluorescence index data standardized by Z-score (standard score) is imported into the random forest algorithm. The random forest algorithm is used to train the petroleum sample classification model, and the data is mapped one-to-one with the category number, thus obtaining the trained petroleum sample classification model. This includes:
[0084] 5-1. Random Sampling: Bootstrap is a random sampling method with replacement. Its principle is to assume the training set has N samples, randomly select one sample each time, record it, and then replace it. This process is repeated N times to obtain a new training subset of size N. .
[0085] 5-2. Constructing a decision tree: For each new training subset Train a decision tree independently, and randomly select m=7 features ( M=59 is the total characteristic number.
[0086] Random forests reduce variance and prevent overfitting by ensembles multiple trees. For most medium-sized datasets, the model's prediction accuracy typically plateaus when the number of trees increases to 100-200. Using 100 trees, while maintaining classification performance, and combined with "compact" optimization, ensures that the prediction script won't crash due to memory overflow when loading a MATLAB ".mat" file on a regular computer.
[0087] S6. Using the petroleum sample classification model constructed above, the collected petroleum samples are classified, and the source of the petroleum sample is obtained according to the classification table. In this embodiment, 426 randomly collected petroleum samples are classified. First, the three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of the 426 petroleum samples are obtained. The obtained three-dimensional fluorescence spectrum data is preprocessed to obtain standardized three-dimensional fluorescence spectrum data. 59 fluorescence index values of each oil sample are extracted. The data are imported into the above-trained random forest model and compared with the petroleum samples in the oil sample source category table in the model (i.e., the correspondence between category number, oil field name, platform name and well number has been established) to find the category with the highest similarity as the predicted category. For the above 426 samples, the identification accuracy rate reaches more than 80%. The petroleum sample identification results are shown in Table 2 (including continued Table 2(1) to continued Table 2(14)).
[0088] Table 2. Petroleum Sample Identification Results
[0089]
[0090] Continued from Table 2 (1)
[0091]
[0092] Continued from Table 2 (2)
[0093]
[0094] Continued from Table 2 (3)
[0095]
[0096] Continued from Table 2 (4)
[0097]
[0098] Continued from Table 2 (5)
[0099]
[0100] Continued from Table 2 (6)
[0101]
[0102] Continued from Table 2 (7)
[0103]
[0104] Continued from Table 2 (8)
[0105]
[0106] Continued from Table 2 (9)
[0107]
[0108] Continued from Table 2 (10)
[0109]
[0110] Continued from Table 2 (11)
[0111]
[0112] Continued from Table 2 (12)
[0113]
[0114] Continued from Table 2 (13)
[0115]
[0116] Continued from Table 2 (14)
[0117]
[0118] Table 2 shows the predictions for 426 samples. Only 17 samples (underlined data in the table) with sample numbers 6, 7, 55, 65, 74, 141, 150, 195, 203, 205, 283, 299, 301, 307, 396, 399, and 401) had incorrect predictions. The accuracy of identifying the category of petroleum samples, i.e., the corresponding oil spill source, in this embodiment is approximately 96%.
[0119] Although the present invention has been described above, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are preferred application examples that demonstrate the core technical ideas of the present invention, and are merely illustrative and not restrictive. Those skilled in the art can make many improvements and changes under the guidance of the present invention without departing from the spirit of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A rapid identification method for the source of marine oil spills based on three-dimensional fluorescence spectroscopy data, characterized in that, Includes the following steps: Step 1: Construct a petroleum sample classification model, including: The three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of oil samples from known offshore oil fields were measured. Based on the three-dimensional fluorescence spectrum and ultraviolet fluorescence spectrum data of oil samples from oil wells, a category table representing the source of the oil sample was established. The category table includes category number, oil field name, platform name and oil well number. For each oil sample, the three-dimensional fluorescence spectral data is preprocessed using the ultraviolet fluorescence spectral data, including: removal of Rayleigh and Raman scattering, conformal interpolation of missing data, correction of internal filtration effect and norm normalization. The ultraviolet fluorescence spectral data is used to calculate the internal filtration effect correction factor. The correction factor is used to recover the true fluorescence intensity of the three-dimensional fluorescence spectrum of the petroleum sample, thereby obtaining the standardized three-dimensional fluorescence spectral data of the oil sample. Fifty-nine fluorescence index data were extracted from the standardized three-dimensional fluorescence spectral data of all oil samples and Z-score normalized. The fluorescence index data after Z-score standardization is imported into the random forest algorithm and mapped one-to-one with the category number to obtain the trained petroleum sample classification model. Step 2: Use the petroleum sample classification model constructed above to identify the category of the collected petroleum sample, and obtain the source of the petroleum sample according to the category table.
2. The rapid identification method for the source of marine oil spills according to claim 1, characterized in that, In step 1, the excitation wavelength Ex of the three-dimensional fluorescence spectrum is set to 200-580 nm with a scan step size of 5 nm; the emission wavelength is set to 250-580 nm with a scan step size of 5 nm; and the ultraviolet fluorescence spectrum wavelength is set to 200-580 nm with a scan step size of 5 nm.
3. The rapid identification method for the source of marine oil spills according to claim 1, characterized in that, In step 1, the 59 fluorescence index data include regional fluorescence integral, classic characteristic fluorescence peaks, key optical ratio index, regional texture and statistical features, global main peak location features, and global matrix distribution features.