Machine Learning Models for Rapid Geochemical Oil Spill Origin Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods lack a well-defined methodology for accurately and efficiently determining the origin of oil spills at sea, particularly in production areas with similar compositional characteristics, which is crucial for exploratory and forensic oil research.
Innovation Solution
A method using multivariate data analysis and machine learning algorithms to build classification models for identifying the oil complex of origin from spill samples, involving data preprocessing, exploratory analysis, and application of various machine learning algorithms to achieve high accuracy in oil field prediction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional geochemical analysis methods are used to identify oil spill origins, then comprehensive chemical characterization can be achieved, but the analysis time is excessive and human subjectivity affects results
Solution Approach 1:
The patent applies preliminary action by pre-processing geochemical data and training machine learning models with historical oil samples before actual spill analysis. The system pre-establishes classification models using techniques like Principal Component Analysis (PCA) and trains algorithms such as Random Forest and SVM with labeled training datasets, so that when a real spill occurs, the pre-trained model can rapidly classify the oil origin without requiring time-consuming manual analysis of each parameter.
Solution Approach 2:
The patent replaces the mechanical system of manual geochemical analysis with an automated machine learning system. Instead of researchers manually comparing chromathograms and making subjective judgments, the system uses automated algorithms (Random Forest, SVM, k-NN) to process geochemical data, extract features, and classify oil origins objectively. This substitution eliminates human subjectivity and significantly reduces analysis time while maintaining or improving accuracy.
2Reliability
If manual geochemical analysis is performed by experts, then interpretative results can be obtained, but human subjectivity reduces reliability and consistency
Solution Approach 1:
The patent applies self-service by designing a system that performs automated self-classification of oil spills. The machine learning models automatically process input chromathogram data, extract relevant features, compare them against the trained database, and generate classification results without requiring continuous human intervention. The system serves itself by autonomously making decisions based on the trained models, eliminating human subjectivity and ensuring consistent, reproducible results across different users and time periods.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously validates its classifications against the trained database and provides confidence scores for each prediction. The model learns from the feedback of classification accuracy and can be retrained with new data to improve performance. This feedback loop ensures high reliability by continuously optimizing the system's ability to correctly identify oil origins while maintaining methodological complexity management through automated validation protocols.
3Measurement precision
If comprehensive geochemical parameters are analyzed, then accurate oil characterization is achieved, but the complexity of data processing increases significantly
Solution Approach 1:
The patent applies the extraction principle by selectively extracting the most discriminative features from comprehensive geochemical datasets. Using feature extraction techniques and dimensionality reduction methods like PCA, the system identifies and extracts key chromathogram parameters that are most relevant for distinguishing between different oil origins. This extraction process removes redundant and less informative parameters, reducing data processing complexity while preserving the accuracy needed for reliable oil characterization and classification.
Data Source
AI summary
The present disclosure is directed to embodiments of a method that aims at improving the process of recognizing the origin of oil spills, sampled as orphan spots on the sea surface, especially due to the time spent nowadays on these activities (hours and/or days for data analysis and interpretation) and given the difficulty of obtaining such accurate/reliable results, given the subjectivity inherent to human resources. The method described herein aims at significantly contributing to the geochemical research through the use of mathematical routines and machine learning for the generation of classification models, by means of pattern recognition, serving as a decision-making instrument in exploratory biases.


