Unbalanced ocean oil and gas data processing and classifying method and electronic equipment
By constructing a hypergraph to represent the high-order local manifold relationships between minority class samples, new boundary samples are identified and generated, solving the problem of imbalanced marine oil and gas data and improving the decision performance and recognition accuracy of marine oil and gas data classifiers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-18
- Publication Date
- 2026-04-14
AI Technical Summary
The imbalance of marine oil and gas data hinders the application of artificial intelligence technology in critical business scenarios, especially in reservoir discovery and safety early warning, where the identification accuracy is low, and traditional methods miss valuable reservoirs or cannot effectively learn a very small number of failure modes.
By constructing a hypergraph to represent the higher-order local manifold relationships between minority class samples, we can identify boundary samples and generate new boundary samples, thus expanding the minority class sample set. Combined with ensemble classification model training methods, this improves data balance and classifier decision performance.
It effectively expands and balances unbalanced marine oil and gas data, improves the decision boundary of the classifier, provides a more discriminative training set, and improves the identification accuracy of reservoir identification and safety monitoring.
Smart Images

Figure CN121859002A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of marine oil and gas data processing technology, and in particular to an unbalanced marine oil and gas data processing and classification method and electronic equipment. Background Technology
[0002] Marine oil and gas data faces extensive challenges in "data governance," primarily manifested in the following aspects: First, the long business chain and severe data silos: Marine oil and gas operations encompass the entire chain from upstream exploration and midstream transportation to downstream refining. Each link is operated by different departments using independent systems, resulting in information barriers between data from reservoirs, wellbores, pipelines, and platforms, making it difficult to achieve collaborative optimization and closed-loop control throughout the entire process. Second, high technical requirements and difficulties in data acquisition and communication: The high pressure and corrosiveness of the marine environment place extremely high demands on sensors and communication equipment. Underwater communication mainly relies on acoustic communication, but suffers from low bandwidth, long latency, and high energy consumption, limiting real-time, high-capacity data transmission. Data acquisition itself also faces challenges, such as difficulties in accurate measurement and a lack of unified sampling standards, leading to poor data comparability. Third, incomplete and missing data: Due to equipment failures, unrecorded shut-in periods, and other reasons, key time-series data such as flow rate and pressure are prone to being missing, directly and seriously affecting the quality of reservoir management decisions. Fourth, data bias: Early studies have shown that there are systematic biases in the estimation of offshore oil and gas field reserves. For example, the reserves of small oil fields are often overestimated, while the reserves of large oil fields are underestimated.
[0003] These challenges have led to the problem of imbalanced marine oil and gas data. As one of the core challenges in marine oil and gas data, data imbalance directly hinders the successful application of artificial intelligence technology in critical business scenarios (such as reservoir discovery and safety early warning). Summary of the Invention
[0004] In view of this, embodiments of this application provide an unbalanced marine oil and gas data processing and classification method and electronic device.
[0005] According to a first aspect of this application, embodiments of this application provide a method for processing unbalanced marine oil and gas data, including:
[0006] An unbalanced marine oil and gas data sample set was obtained. The unbalanced marine oil and gas data sample set includes a first type of data sample set and a second type of data sample set. The first type of data sample set includes multiple first type data samples, and the second type of data sample set includes multiple second type data samples. The number of first type data samples is greater than the number of second type data samples. A hypergraph is constructed based on the second type of data sample set. The hypergraph includes multiple hyperedges, and each hyperedge includes at least two second type data samples whose similarity meets the similarity threshold. Based on the second type of data sample set and the first type of data sample set, construct the nearest neighbor sample set for each second type of data sample; Based on the set of nearest neighbors for each second-class data sample, determine the boundary samples in the second-class data sample set; New boundary samples are generated based on boundary samples and hypergraphs to obtain a new second type of data sample set; Based on the first type of data sample set and the new second type of data sample set, a processed unbalanced marine oil and gas data sample set is obtained.
[0007] Optionally, a hypergraph is constructed based on the second type of data sample set, including: For each second-class data sample, determine the first similarity between the second-class data sample and other second-class data samples; Based on the first similarity, multiple first target second category data samples that are similar to the second category data samples are identified from other second category data samples; Based on a similarity threshold, second-target second-class data samples that are similar to second-class data samples are determined from multiple first-target second-class data samples; Based on the second type of data samples and the second target second type of data samples, a hyperedge is generated to obtain the hypergraph.
[0008] Optionally, based on the second type of data sample set and the first type of data sample set, a nearest neighbor sample set is constructed for each second type of data sample, including: For each second-class data sample, determine the second similarity between the second-class data sample and other second-class data samples, as well as between each first-class data sample; Based on the second similarity, multiple nearest neighbor samples similar to the second type of data samples are determined from other second type of data samples and multiple first type of data samples, thus obtaining the nearest neighbor sample set of the second type of data samples.
[0009] Optionally, based on the set of nearest neighbors for each second-class data sample, boundary samples in the second-class data sample set are determined, including: For each second-class data sample, if it is determined that the corresponding nearest neighbor sample set includes both first-class and second-class data samples, then the second-class data sample is determined as the boundary sample in the second-class data sample set.
[0010] Optionally, new boundary samples are generated based on the boundary samples and the hypergraph, including: Based on the weight allocation strategy, the weights corresponding to the boundary samples and the safe region samples are determined respectively; the safe region samples are the data samples in the second type of data samples excluding the boundary samples; the weight of the boundary samples is greater than the weight of the safe region samples. Based on the weights corresponding to the boundary samples and the safe area samples, the selection probabilities of the boundary samples and the safe area samples are determined respectively. Base samples are selected from boundary samples and safe zone samples based on selection probability; New boundary samples are generated based on base samples and hypergraphs.
[0011] Optionally, new boundary samples are generated based on the base samples and the hypergraph, including: For each base sample, a reference sample corresponding to the base sample is determined based on the hypergraph, and the reference sample and the base sample are connected by a hyperedge; New boundary samples are generated based on base samples, reference samples, and interpolation strategies.
[0012] According to a second aspect of this application, embodiments of this application provide a classification model training method, including: A processed set of unbalanced marine oil and gas data samples is obtained. The processed set of unbalanced marine oil and gas data samples is generated by the unbalanced marine oil and gas data processing method described above. The processed set of unbalanced marine oil and gas data samples includes a first type of data sample set and a new second type of data sample set. The first type of data sample set includes multiple first type data samples and their corresponding labels. The new second type of data sample set includes multiple second type data samples and their corresponding labels. Based on multiple first-class data samples and multiple second-class data samples, the first base learner and the second base learner are trained respectively to obtain the trained first base learner and the trained second base learner; the first base learner and the second base learner are of different types. Based on the first output data during the training of the first base learner, the second output data during the training of the second base learner, and the corresponding labels, the meta-learner is trained to obtain the trained meta-learner. An ensemble classification model is obtained based on the first base learner, the second base learner, and the meta-learner after training.
[0013] Alternatively, classification model training methods may also include: Based on the hypergraph generated during the process of generating the processed unbalanced marine oil and gas data sample set, the validity information of the hypergraph structure is determined. Based on multiple Class I data samples, multiple Class II data samples, and the prediction function of the ensemble classification model, the feature importance analysis information of the features in the multiple Class I data samples and multiple Class II data samples during the training of the ensemble classification model is determined. The contribution information of the ensemble classification model is determined based on the first performance evaluation metric of the ensemble classification model, the second performance evaluation metric of the first base learner after training, and the third performance evaluation metric of the second base learner after training. Based on the effectiveness information of the hypergraph structure, the feature importance analysis information, and the contribution information of the ensemble classification model, decision analysis information for the ensemble classification model is formed.
[0014] Optionally, based on the hypergraph generated during the processing of the unbalanced marine oil and gas data sample set, the validity information of the hypergraph structure is determined, including: Based on the hypergraph generated during the process of generating the processed imbalanced marine oil and gas data sample set, the average size of the hyperedge is determined. The average size of the hyperedge represents the coverage of the local neighborhood between the second type of data samples. The larger the average size of the hyperedge, the denser the distribution of the minority class samples with high similarity. The quality of hyperedges is determined based on the hypergraph generated during the process of generating the processed unbalanced marine oil and gas data sample set. The validity information of the hypergraph structure is determined based on the average size and / or quality of the hyperedges.
[0015] Optionally, based on multiple first-class data samples, multiple second-class data samples, and the prediction function of the ensemble classification model, the feature importance analysis information of the features in the multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model is determined, including: Based on multiple Class I data samples, multiple Class II data samples, and the prediction function of the corresponding ensemble classification model, determine the contribution score of each feature of each data sample in the training process of the ensemble classification model. Based on contribution scores, the global importance information of each feature in multiple Class I data samples and multiple Class II data samples is determined during the classification model training process. Based on global importance information, feature importance analysis information is determined for features in multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model.
[0016] According to a third aspect of this application, embodiments of this application provide a method for classifying marine oil and gas data, including: Marine oil and gas data to be classified was obtained; the marine oil and gas data to be classified includes first-class data and second-class data, with the number of first-class data being greater than the number of second-class data. The ensemble classification model is used to process the marine oil and gas data to be classified, and the classification results corresponding to the marine oil and gas data to be classified are obtained. The ensemble classification model is trained using the classification model training method described above.
[0017] According to a fourth aspect of this application, embodiments of this application provide an electronic device, including: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the unbalanced marine oil and gas data processing method described above, or the classification model training method described above, or the marine oil and gas data classification method described above.
[0018] According to a fifth aspect of this application, embodiments of this application provide a computer-readable storage medium storing computer instructions for causing a computer to perform the unbalanced marine oil and gas data processing method, the classification model training method, or the marine oil and gas data classification method as described above.
[0019] The imbalanced marine oil and gas data processing method, classification model training method, marine oil and gas data classification method, electronic device, and readable storage medium provided in this application's embodiments obtain an imbalanced marine oil and gas data sample set; construct a hypergraph based on a second type of data sample set; construct a nearest neighbor sample set for each second type of data sample based on the second type of data sample set and the first type of data sample set; determine the boundary samples in the second type of data sample set based on the nearest neighbor sample set for each second type of data sample; generate new boundary samples based on the boundary samples and the hypergraph to obtain a new second type of data sample set; and obtain... The processed imbalanced marine oil and gas data sample set is used to represent the higher-order local manifold relationships between minority class samples through the hypergraph structure of the similarity matrix. This accurately identifies the boundary samples in the minority class sample set that have a significant impact on the classifier's performance. Hyperedge connections can then be used to constrain the generation process of new boundary samples in the minority class sample set, avoiding the generation of false boundary samples that deviate from the true distribution and obtaining a larger number of new boundary samples. This expands the second type of data sample set, strengthens the sample distribution in the boundary region, improves the decision boundary of the classifier, and provides a more balanced and discriminative training set for subsequent ensemble classification model training.
[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating an unbalanced marine oil and gas data processing method according to an embodiment of this application. Figure 2 This is a flowchart illustrating a classification model training method in an embodiment of this application; Figure 3This is a flowchart illustrating a marine oil and gas data classification method in an embodiment of this application. Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] Imbalanced data refers to a significant difference in the number of samples from different categories (especially the target variable) in a dataset used for machine learning modeling. This phenomenon is very common in the offshore oil and gas sector because economically valuable reservoirs or high-risk accident conditions are inherently low-probability events. Imbalanced data severely impacts the performance of AI models, causing them to become "lazy," meaning they tend to correctly predict the majority class while ignoring the minority class, or performing poorly in predicting / classifying minority class samples. For example, in reservoir identification, if non-reservoir data accounts for 99% and reservoir data only 1%, a model that "does nothing" can achieve 99% accuracy, but it has practically no useful value.
[0024] Imbalance issues are evident in multiple core business scenarios of offshore oil and gas: In reservoir identification, the majority of samples are ordinary rock formations (such as mudstone and tight layers), while the minority of samples are commercially viable natural gas hydrate / oil and gas reservoirs. This data imbalance causes traditional methods to miss valuable reservoirs, resulting in low identification accuracy. In lithology identification, common and background lithologies constitute the majority of samples, while key lithologies constitute the minority. This data imbalance leads to poor model identification of geologically significant minority lithologies, resulting in prediction bias. In safety monitoring and anomaly detection, monitoring data under normal operating conditions constitute the majority of samples, while accident data such as pipeline leaks and corrosion constitute the minority. This data imbalance prevents models from effectively learning the very few fault modes, leading to missed detections and potentially catastrophic consequences.
[0025] In summary, the imbalance in marine oil and gas data directly hinders the successful application of artificial intelligence technology in critical business scenarios (such as reservoir discovery and safety early warning). To address this issue, it is necessary to combine data sampling techniques such as SMOTE (Synthetic Minority Oversampling) and advanced machine learning algorithms with business scenarios to achieve data augmentation, obtain balanced data, and thus help improve the classification decision-making performance and interpretability of the model.
[0026] Therefore, embodiments of this application provide a method for processing unbalanced marine oil and gas data, such as... Figure 1 As shown, it includes: S101, Obtain an unbalanced marine oil and gas data sample set. The unbalanced marine oil and gas data sample set includes a first type of data sample set and a second type of data sample set. The first type of data sample set includes multiple first type data samples, and the second type of data sample set includes multiple second type data samples. The number of first type data samples is greater than the number of second type data samples.
[0027] In this embodiment, the unbalanced marine oil and gas data sample set is a sample set in which the data distribution is significantly biased towards one category (majority class), while the sample size of the other category (minority class) is relatively small. The first type of data sample set is the majority class sample set, and the second type of sample set is the minority class sample set.
[0028] Imbalanced marine oil and gas data sample sets can serve as reservoir identification sample sets, with the majority of samples being ordinary rock formations that are not reservoirs (such as mudstone and tight layers), and the minority of samples being natural gas hydrate / oil and gas reservoirs with commercial exploitation value.
[0029] Imbalanced marine oil and gas data sample sets can be, for example, lithology identification sample sets, where common and background lithologies are the majority class samples, while key lithologies are the minority class samples.
[0030] Imbalanced marine oil and gas data sample sets can be used as safety monitoring and anomaly detection sample sets, where monitoring data under normal operating conditions constitute the majority class of samples, while accident data such as pipeline leaks and corrosion constitute the minority class of samples.
[0031] S102, construct a hypergraph based on the second type of data sample set. The hypergraph includes multiple hyperedges, and each hyperedge includes at least two second type of data samples whose similarity satisfies the similarity threshold.
[0032] In this embodiment, due to the sparse distribution of minority class samples and the complex local manifold structure between samples, traditional oversampling methods struggle to accurately capture high-order neighborhood relationships between minority class samples, easily generating synthetic samples that deviate from the true distribution. Therefore, a hypergraph can first be constructed based on cosine similarity to represent the high-order neighborhood relationships between the second class of data samples.
[0033] In this embodiment, hyperedges can be constructed based on a similarity threshold to form a hypergraph, thereby identifying similar second-class data samples.
[0034] S103, Based on the second type of data sample set and the first type of data sample set, construct the nearest neighbor sample set for each second type of data sample.
[0035] In this embodiment, the boundary samples in the second type of data sample set are usually located near the classification decision surface and have a significant impact on the classifier performance. These samples exhibit obvious boundary characteristics in the data distribution, that is, they are tightly surrounded by majority class samples in the feature space and occupy a critical position at the decision boundary in the class distribution. Therefore, in order to expand the second type of data samples and improve the boundary decision performance of the classification model, the data samples can be expanded based on the boundary samples in the second type of dataset to obtain a larger number of new boundary samples. This requires identifying the boundary samples of the second type of data sample set. To achieve the identification of boundary samples in the second type of data sample set, a nearest neighbor sample set for each second type of data sample can be constructed based on the second type of data sample set and the first type of data sample set. The nearest neighbor sample set can include K nearest neighbor samples.
[0036] In this embodiment, the set of nearest neighbors for each second-class data sample can be determined by the K-nearest neighbor algorithm, and / or the set of nearest neighbors for the second-class data sample can be determined based on the similarity between the second-class data samples and between the second-class data samples and the first-class data samples.
[0037] S104, Based on the set of nearest neighbors of each second-class data sample, determine the boundary samples in the second-class data sample set.
[0038] In this embodiment, minority class samples located near the decision boundary can be identified based on the K-nearest neighbor class distribution, thereby obtaining boundary samples in the second type of data sample set.
[0039] S105, generate new boundary samples based on boundary samples and hypergraph to obtain a new second type of data sample set.
[0040] In this embodiment, a hypergraph-guided sampling strategy can be used to preferentially select boundary samples as base samples and select reference samples from their hypergraph neighborhoods to generate new boundary samples through interpolation. The newly generated boundary samples are then merged with the original set of second-class data samples to obtain a new set of second-class data samples.
[0041] S106. Based on the first type of data sample set and the new second type of data sample set, the processed unbalanced marine oil and gas data sample set is obtained.
[0042] In this embodiment, the first type of data sample set and the new second type of data sample set can be merged to obtain a processed imbalanced marine oil and gas data sample set. This results in a balance between the number of first and second type data samples in the processed imbalanced marine oil and gas data sample set, facilitating subsequent training of the classification model and improving the model's accuracy in identifying minority class data.
[0043] The imbalanced marine oil and gas data processing method provided in this application involves: acquiring an imbalanced marine oil and gas data sample set; constructing a hypergraph based on a second type of data sample set; constructing a nearest neighbor set for each second type of data sample based on the second type of data sample set and the first type of data sample set; determining the boundary samples in the second type of data sample set based on the nearest neighbor set for each second type of data sample; generating new boundary samples based on the boundary samples and the hypergraph to obtain a new second type of data sample set; and obtaining the processed imbalanced marine oil and gas data sample set based on the first type of data sample set and the new second type of data sample set. In this way, the hypergraph structure of the similarity matrix represents the high-order local manifold relationship between minority class samples, and accurately identifies the boundary samples in the minority class sample set that have a significant impact on classifier performance. Therefore, hyperedge connections can be used to constrain the generation process of new boundary samples in the minority class sample set, avoiding the generation of false boundary samples that deviate from the true distribution. This expands the second type of data sample set and strengthens the sample distribution in the boundary region, improving the decision boundary of the classifier and providing a more balanced and discriminative training set for subsequent ensemble classification model training.
[0044] In an optional embodiment, step S102, constructing a hypergraph based on the second type of data sample set, includes: For each second-class data sample, determine the first similarity between the second-class data sample and other second-class data samples; based on the first similarity, determine multiple first target second-class data samples that are similar to the second-class data samples from other second-class data samples; based on the similarity threshold, determine second target second-class data samples that are similar to the second-class data samples from multiple first target second-class data samples; based on the second-class data samples and the second target second-class data samples, generate a hyperedge to obtain a hypergraph.
[0045] In practice, a hypergraph can be constructed based on cosine similarity. A hyperedge in the hypergraph can connect any number of vertices, allowing it to better represent the relationships between minority class samples. Suppose a binary classification problem involves an imbalanced marine oil and gas data sample set. The feature vector of each sample It is d A dimensional vector space, that is: , For the firsti Class labels for each sample. Construct a hypergraph. , where the vertex set Each vertex in the array corresponds to a sample, and the hyperedge set Each hyperedge in the dataset contains a set of highly similar samples, while the weight set... This represents the average similarity level between samples within each hyperedge. (Using...) The set representing the minority class of samples. The set representing the majority class samples. Since the generation of hyperedges is based on the local similarity relationship between samples, therefore, for each First, calculate its comparison with all other minority class samples. The cosine similarity is denoted as . As shown in equation (1).
[0046] (1).
[0047] Based on the calculated similarity matrix, select with Most similar k A minority class sample, for each sample Build k Neighbor candidate set, denoted as As shown in equation (2).
[0048] (2).
[0049] Subsequently, a similarity threshold filtering mechanism is applied to obtain the final hyperedge, denoted as . As shown in equation (3).
[0050] (3).
[0051] Among them, parameters This is a similarity threshold used to ensure that samples within the hyperedge have sufficient similarity. In a special case, when there are no samples in the nearest neighbor candidate set that meet the threshold condition, the hyperedge contains only samples from that sample. itself.
[0052] To measure the importance of different hyperedges in the overall structure, each hyperedge is assigned a specific value. Assign weights The calculation method is shown in Equation (4). This weight reflects the average similarity level of samples within the hyperedge. The higher the weight value, the closer the relationship between the samples within the hyperedge.
[0053] (4).
[0054] In this embodiment, a hypergraph is constructed for minority class samples based on cosine similarity, where hyperedges represent clusters of highly similar samples, and the density of clusters is measured by weights, which enables the hypergraph to better represent high-order neighborhood relationships between minority class samples.
[0055] In an optional embodiment, step S103, based on the second type of data sample set and the first type of data sample set, constructs a nearest neighbor sample set for each second type of data sample, including: For each second-class data sample, determine the second similarity between the second-class data sample and other second-class data samples and each first-class data sample; based on the second similarity, determine multiple nearest neighbor samples similar to the second-class data sample from other second-class data samples and multiple first-class data samples to obtain the nearest neighbor sample set of the second-class data sample.
[0056] In practice, for each second type of data sample This allows us to calculate the second similarity between it and other second-class data samples, as well as between it and each first-class data sample. Then, we select... Most similar k The nearest neighbor samples are used to form the nearest neighbor set of the second type of data samples. .
[0057] In this embodiment, the set of nearest neighbor samples for each second-class data sample can be quickly determined through similarity calculation.
[0058] In an optional embodiment, step S104, determining the boundary samples in the second type of data sample set based on the nearest neighbor sample set of each second type of data sample, includes: For each second-class data sample, if it is determined that the corresponding nearest neighbor sample set includes both first-class and second-class data samples, then the second-class data sample is determined as the boundary sample in the second-class data sample set.
[0059] In practice, for each second-class data sample, the class distribution of its nearest neighbor sample set can be analyzed to determine whether the sample is located in the boundary region. If there are samples of different classes among the K nearest neighbors of a sample, the sample is identified as a boundary sample, formally expressed as equation (6), where, and Representing samples respectively and The category label, the nearest neighbor sample set is .
[0060] (6).
[0061] Based on the above expression, the boundary sample set B can be defined as equation (7).
[0062] (7).
[0063] In this embodiment, the local neighborhood distribution of samples in the feature space is effectively utilized, which can accurately identify samples located in the category boundary region of the second type of data sample set.
[0064] In an optional embodiment, step S105, generating new boundary samples based on boundary samples and the hypergraph, includes: Based on the weight allocation strategy, the weights corresponding to the boundary samples and the safe region samples are determined respectively; the safe region samples are the data samples in the second type of data samples excluding the boundary samples; the weight of the boundary samples is greater than the weight of the safe region samples; based on the weights corresponding to the boundary samples and the safe region samples respectively, the selection probabilities corresponding to the boundary samples and the safe region samples are determined respectively; base samples are selected from the boundary samples and the safe region samples based on the selection probabilities; new boundary samples are generated based on the base samples and the hypergraph.
[0065] In some implementations, new boundary samples are generated based on base samples and the hypergraph, including: For each base sample, a reference sample corresponding to the base sample is determined based on the hypergraph, and the reference sample and the base sample are connected by a hyperedge; based on the base sample, the reference sample, and the interpolation strategy, a new boundary sample is generated.
[0066] In specific implementation, firstly, based on the distribution location of the samples in the feature space, the samples in the second type of data sample set are divided into boundary samples and safe region samples. Boundary samples are located near the class decision boundary, while safe region samples are located inside the minority class distribution, surrounded by samples of the same class. Since boundary samples have a greater impact on the classifier's decision, they are given higher sampling weights. This application embodiment assigns weights to boundary samples... Set the weights of samples in the safe region Twice as of the weights. Based on the weight allocation, the sample The probability of being selected as a base sample is denoted as As shown in equation (8). Specifically, when When it is a boundary sample, .when When the sample is from a safe area, .
[0067] (8).
[0068] Secondly, samples with a probability greater than 0.3 are selected as base samples. For each selected base sample, samples of the same type connected to the base sample through hyperedges in the hypergraph are preferentially selected as reference samples. These samples, connected by hyperedges, have high similarity. Let... For Hypergraph Includes base samples The set of all hyperedges, denoted as the reference sample set. As shown in equation (9). Wherein, For superedge, For reference sample, The class labels of the base samples, The category label for the reference sample.
[0069] (9).
[0070] The generation of new boundary samples employs an improved interpolation strategy. Base samples are selected based on weight distribution, with interpolation preferentially performed from minority class reference samples within the hypergraph neighborhood. Appropriate Gaussian noise is added to increase the diversity of the new boundary samples. The calculation formula for the synthesized new boundary samples is shown in equation (10), where... For new boundary samples; As the base sample, The interpolation coefficients are uniformly distributed. Used to control noise intensity.
[0071] (10).
[0072] In this embodiment, based on the accurate identification of boundary samples, a hypergraph-guided boundary region sample sampling strategy is further proposed. By combining the hypergraph and boundary characteristics of the samples, the generated new boundary samples are ensured to maintain the original data distribution characteristics and strengthen the classification boundary, ultimately achieving the synthesis of high-quality samples in the boundary region.
[0073] This application also provides a classification model training method, such as... Figure 2 As shown, it includes: S201, Obtain the processed unbalanced marine oil and gas data sample set; The processed unbalanced marine oil and gas data sample set is generated by the unbalanced marine oil and gas data processing method described above; The processed unbalanced marine oil and gas data sample set includes a first type of data sample set and a new second type of data sample set. The first type of data sample set includes multiple first type data samples and corresponding labels; The new second type of data sample set includes multiple second type data samples and corresponding labels.
[0074] S202, based on multiple first-class data samples and multiple second-class data samples, train the first base learner and the second base learner respectively to obtain the trained first base learner and the trained second base learner; the first base learner and the second base learner are of different types.
[0075] S203, based on the first output data during the training of the first base learner, the second output data during the training of the second base learner, and the corresponding labels, the meta-learner is trained to obtain the trained meta-learner.
[0076] S204, based on the trained first base learner, the trained second base learner, and the trained meta-learner, obtains an ensemble classification model.
[0077] In this embodiment, firstly, two base learners are independently trained on the processed imbalanced marine oil and gas data sample set, and the prediction results of the base learners are used to construct the training set of the meta-learner. The meta-learner uses the prediction probabilities of each base learner as input features, learns the mapping relationship between the prediction probabilities of each base learner and the true labels through maximum likelihood estimation, and adaptively determines the linear weight of the prediction probabilities of each base learner in the final decision, thereby achieving weighted fusion of prediction results.
[0078] In practice, a stacking ensemble strategy can be adopted, where Random Forest and Extreme Gradient Boosting (XGBoost) are trained independently as base learners to generate predicted probabilities. and As meta-features input into the logistic regression meta-learner, the decision function of the meta-learner is: ; in, and These are the weighting coefficients. b is the bias term, and sigmoid is the activation function.
[0079] The classification model training method provided in this application involves: acquiring an imbalanced marine oil and gas data sample set; constructing a hypergraph based on a second type of data sample set; constructing a nearest neighbor set for each second type of data sample based on the second type of data sample set and the first type of data sample set; determining boundary samples in the second type of data sample set based on the nearest neighbor set of each second type of data sample; generating new boundary samples based on the boundary samples and the hypergraph to obtain a new second type of data sample set; and obtaining the processed imbalanced marine oil and gas data sample set based on the first type of data sample set and the new second type of data sample set. Thus, through phase... The hypergraph structure of the similarity matrix is used to represent the higher-order local manifold relationships between minority class samples, and the boundary samples in the minority class sample set that have an important impact on the classifier performance are accurately identified. Thus, the generation process of new boundary samples in the minority class sample set can be constrained by the hyperedge connection, resulting in a larger number of new boundary samples. This expands the second class data sample set and strengthens the sample distribution in the boundary region, which can improve the decision boundary of the classifier. This provides a more balanced and discriminative training set for training the ensemble classification model. In this way, an ensemble classification model that can classify imbalanced datasets more accurately can be obtained.
[0080] In an optional embodiment, the classification model training method further includes: Based on the hypergraph generated during the processing of the imbalanced marine oil and gas data sample set, the effectiveness information of the hypergraph structure is determined. Based on multiple first-class data samples, multiple second-class data samples, and the prediction function of the ensemble classification model, the feature importance analysis information of the features in the multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model is determined. Based on the first performance evaluation index of the ensemble classification model, the second performance evaluation index of the trained first base learner, and the third performance evaluation index of the trained second base learner, the contribution information of the ensemble classification model is determined. Based on the effectiveness information of the hypergraph structure, the feature importance analysis information, and the contribution information of the ensemble classification model, the decision analysis information of the ensemble classification model is formed.
[0081] In some implementations, the validity information of the hypergraph structure is determined based on the hypergraph generated during the process of generating the processed imbalanced marine oil and gas data sample set. This includes: determining the average size of the hyperedges based on the hypergraph generated during the process of generating the processed imbalanced marine oil and gas data sample set. The average size of the hyperedges characterizes the coverage of the local neighborhood between the second type of data samples. A larger average size of the hyperedges indicates a denser distribution of minority class samples with high similarity; determining the quality of the hyperedges based on the hypergraph generated during the process of generating the processed imbalanced marine oil and gas data sample set; and determining the validity information of the hypergraph structure based on the average size and / or quality of the hyperedges.
[0082] In some implementations, based on multiple first-class data samples, multiple second-class data samples, and the prediction function of the ensemble classification model, the feature importance analysis information of the features in the multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model is determined. This includes: determining the contribution score of each feature of each data sample in the multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model based on the prediction function of the corresponding ensemble classification model; determining the global importance information of each feature in the multiple first-class data samples and multiple second-class data samples during the training of the classification model based on the contribution score; and determining the feature importance analysis information of the features in the multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model based on the global importance information.
[0083] In this embodiment, to explain the transparency and reliability of the ensemble classification model in imbalanced data classification tasks, this application embodiment performs an interpretability analysis on the classification results from three aspects: hypergraph structure effectiveness, feature importance, and model ensemble contribution after model training. This interpretability analysis does not change the original classification process, but rather performs a post-hoc analysis of the decision-making mechanism of the ensemble classification model based on the classification results.
[0084] In practice, the role of hypergraphs in imbalanced data processing is first evaluated by analyzing their structure. The size of the hyperedge reflects the density of sample clustering, and the vertex set... Corresponding data samples, hyperedge set This indicates that each hyperedge contains a set of highly similar samples, and the weight set... This reflects the average similarity level between samples within each hyperedge. E | represents the total number of superedges, | For the first i The number of vertices contained in a hyperedge. The average size of a hyperedge is defined as shown in equation (14), denoted as . Se .
[0085] (14).
[0086] The hyperedge quality is evaluated by a combination of intraclass consistency and distance compactness, and its calculation formula is shown in Equation (15). The balance coefficient between the first and second types of data samples. For the class purity of samples within the hyperedge, For the sake of compactness, Let be the average Euclidean distance between samples within the hyperedge. When the hyperedge quality... A value greater than 0.8 proves that the hypergraph construction is effective.
[0087] (15).
[0088] Then, to quantify the contribution of each input feature to the model output, this embodiment uses the SHAP (SHapley Additive exPlanations) method for feature importance analysis. The SHAP value is based on the Shapley value in game theory, assigning a contribution score to each feature to reflect its marginal contribution to a given prediction result. The sample... Features j The SHAP value is denoted as The calculation formula is shown in equation (16).
[0089] (16) in, F For the complete feature set, S For those without features j a subset of For the sample The model prediction function. This formula calculates the features by iterating through all possible combinations of feature subsets. j At that time, the model is on the sample The differences in the predicted values are calculated and weighted according to the subset size to obtain the features. j The marginal contribution of the prediction for this sample. The feature can be obtained by calculating the SHAP value of all samples. j The global importance of is shown in Equation (17).
[0090] (17).
[0091] in, The total number of features, determined by feature importance. The features are ranked to ultimately analyze their importance.
[0092] Finally, the performance of each model was cross-validated, and the average improvement of the ensemble classification model compared to the base classifier was calculated, as shown in Equation (19).
[0093] (19).
[0094] in, M This refers to model evaluation metrics, which may include precision, recall, the harmonic mean of precision and recall (F1-Score), accuracy, and AUC. To improve the performance of the integrated classification model, The performance of the two base classifiers are respectively. This represents the difference between the ensemble learning and the best base classifier. The larger the value, the better the ensemble learning classifier performs compared to a single classifier.
[0095] In some embodiments, to systematically evaluate the correctness and effectiveness of the imbalanced marine oil and gas data processing method and classification model training method of this application for classifying imbalanced data, multiple sets of experiments were designed. The imbalanced marine oil and gas data processing method and classification model training method of this application (denoted as HyperGraph-guided Oversampling for Imbalanced Data Classification) were tested on nine public datasets (numbered D1-D9) and one real well logging dataset (numbered D10). The proposed classification method (HGODC) was compared with nine mainstream oversampling methods (SMOTE (synthetic minority class oversampling), CURE-SMOTE (an improved oversampling method combining CURE clustering and SMOTE), G-SMOTE (geometric SMOTE oversampling), KMeans-SMOTE (an improved oversampling method combining KMeans clustering and SMOTE), LVQ-SMOTE (an improved oversampling method combining learned vector quantization and SMOTE), MSMOTE (an improved SMOTE oversampling method), SMOTETomek (a hybrid sampling method), SMOTE-Cosine (an improved SMOTE oversampling method), and SMOTE-IPF (a classic improved SMOTE oversampling method)). The effectiveness and correctness of the proposed method were verified in terms of classification performance. The base classifiers for the classification model training method in this application are random forest and XGBoost, and the meta-classifier is logistic regression. The nine comparison methods use random forest as their classifier.
[0096] Table 1 shows the basic information of the dataset used.
[0097] Table 1
[0098] The classification results of the 10 methods on 10 real datasets are shown in Table 2.
[0099] Table 2
[0100] As shown in Table 2, the proposed method demonstrates the most significant advantage in the AUC metric, which measures the overall classification performance of the model. Compared to traditional SMOTE variants and their improved algorithms, HGODC, through its hypergraph-guided boundary identification sampling strategy, can more accurately identify and enhance the distribution of minority class samples near the decision boundary, thereby significantly improving the identification ability of the minority class while maintaining the majority class classification accuracy. Furthermore, the adaptive adjustment of interpolation coefficients ensures stable performance on different datasets, enabling HGODC to not only surpass clustering-based algorithms such as KMeansSMOTE, but also exhibit more balanced and superior classification performance compared to boundary-focused methods such as MSMOTE and SMOTE_IPF, providing a more reliable solution to imbalanced data classification problems encountered in practical applications.
[0101] This application also provides a method for classifying marine oil and gas data, such as... Figure 3 As shown, it includes: S301, Obtain marine oil and gas data to be classified; the marine oil and gas data to be classified includes first-class data and second-class data, with the number of first-class data being greater than the number of second-class data.
[0102] S302, The marine oil and gas data to be classified is processed based on the ensemble classification model to obtain the classification results corresponding to the marine oil and gas data to be classified; the ensemble classification model is trained using the classification model training method described above.
[0103] The marine oil and gas data classification method provided in this application involves: acquiring an imbalanced marine oil and gas data sample set; constructing a hypergraph based on a second type of data sample set; constructing a nearest neighbor sample set for each second type of data sample based on the second type of data sample set and the first type of data sample set; determining the boundary samples in the second type of data sample set based on the nearest neighbor sample set for each second type of data sample; generating new boundary samples based on the boundary samples and the hypergraph to obtain a new second type of data sample set; and obtaining the processed imbalanced marine oil and gas data sample set based on the first type of data sample set and the new second type of data sample set. Thus, the minority class is represented by a hypergraph structure of a similarity matrix. The model identifies higher-order local manifold relationships between samples and accurately identifies boundary samples in the minority class sample set that significantly impact classifier performance. Hyperedge connections can then be used to constrain the generation process of new boundary samples in the minority class sample set, resulting in a larger number of new boundary samples. This expands the second-class data sample set and strengthens the sample distribution in the boundary region, improving the classifier's decision boundary and providing a more balanced and discriminative training set for the ensemble classification model. Consequently, an ensemble classification model capable of more accurately classifying imbalanced datasets can be obtained. Therefore, when this ensemble classification model is used to classify marine oil and gas data, the results are accurate.
[0104] According to embodiments of this application, this application also provides an electronic device and a readable storage medium.
[0105] Figure 4 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0106] like Figure 4 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0107] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0108] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as imbalanced marine oil and gas data processing methods, classification model training methods, or marine oil and gas data classification methods. For example, in some embodiments, the imbalanced marine oil and gas data processing methods, classification model training methods, or marine oil and gas data classification methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the unbalanced marine oil and gas data processing method, classification model training method, or marine oil and gas data classification method described above can be performed. Alternatively, in other embodiments, computing unit 801 can be configured to perform the unbalanced marine oil and gas data processing method, classification model training method, or marine oil and gas data classification method by any other suitable means (e.g., by means of firmware).
[0109] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0110] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0111] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0112] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0113] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0114] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0115] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0116] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0117] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for processing unbalanced marine oil and gas data, characterized in that, include: An unbalanced marine oil and gas data sample set is obtained, which includes a first type of data sample set and a second type of data sample set. The first type of data sample set includes multiple first type data samples, and the second type of data sample set includes multiple second type data samples. The number of first type data samples is greater than the number of second type data samples. A hypergraph is constructed based on the second type of data sample set. The hypergraph includes multiple hyperedges, and each hyperedge includes at least two second type data samples whose similarity satisfies a similarity threshold. Based on the second type of data sample set and the first type of data sample set, construct a set of nearest neighbor samples for each second type of data sample; Based on the set of nearest neighbors for each second type of data sample, determine the boundary samples in the second type of data sample set; Based on the boundary samples and the hypergraph, new boundary samples are generated to obtain a new second type of data sample set; Based on the first type of data sample set and the new second type of data sample set, a processed unbalanced marine oil and gas data sample set is obtained.
2. The method for processing unbalanced marine oil and gas data according to claim 1, characterized in that, Constructing a hypergraph based on the second type of data sample set includes: For each second type of data sample, determine the first similarity between the second type of data sample and other second type of data samples; Based on the first similarity, a plurality of first target second type data samples similar to the second type data samples are determined from the other second type data samples; Based on a similarity threshold, a second target second type data sample that is similar to the second type data sample is determined from multiple first target second type data samples; Based on the second type of data sample and the second target second type of data sample, a hyperedge is generated to obtain the hypergraph.
3. The method for processing unbalanced marine oil and gas data according to claim 1, characterized in that, Based on the second type of data sample set and the first type of data sample set, a nearest neighbor sample set is constructed for each second type of data sample, including: For each second type of data sample, determine a second similarity between the second type of data sample and other second type of data samples, as well as between each first type of data sample; Based on the second similarity, multiple nearest neighbor samples similar to the second type of data samples are determined from the other second type of data samples and multiple first type of data samples to obtain a set of nearest neighbor samples of the second type of data samples.
4. The method for processing unbalanced marine oil and gas data according to claim 1, characterized in that, Based on the nearest neighbor sample set of each second type of data sample, the boundary samples in the second type of data sample set are determined, including: For each second type of data sample, if it is determined that the corresponding nearest neighbor sample set includes both the first type of data sample and the second type of data sample, then the second type of data sample is determined as the boundary sample in the second type of data sample set.
5. The method for processing unbalanced marine oil and gas data according to claim 1, characterized in that, Generating new boundary samples based on the boundary samples and the hypergraph includes: Based on a weight allocation strategy, the weights corresponding to the boundary samples and the safe region samples are determined respectively; the safe region samples are the data samples in the second type of data samples excluding the boundary samples; the weight of the boundary samples is greater than the weight of the safe region samples. Based on the weights corresponding to the boundary samples and the safe area samples, the selection probabilities corresponding to the boundary samples and the safe area samples are determined respectively. Based on the selection probability, a base sample is selected from the boundary sample and the safe area sample; New boundary samples are generated based on the base samples and the hypergraph.
6. The method for processing unbalanced marine oil and gas data according to claim 5, characterized in that, Generating new boundary samples based on the base samples and the hypergraph includes: For each base sample, a reference sample corresponding to the base sample is determined based on the hypergraph, and the reference sample and the base sample are connected through the hyperedge; New boundary samples are generated based on the base samples, the reference samples, and the interpolation strategy.
7. A classification model training method, characterized in that, include: Obtain a processed sample set of unbalanced marine oil and gas data; The processed unbalanced marine oil and gas data sample set is generated by the unbalanced marine oil and gas data processing method as described in any one of claims 1-6; the processed unbalanced marine oil and gas data sample set includes a first type of data sample set and a new second type of data sample set, the first type of data sample set includes multiple first type data samples and corresponding labels; the new second type of data sample set includes multiple second type data samples and corresponding labels. Based on multiple first-type data samples and multiple second-type data samples, a first base learner and a second base learner are trained respectively to obtain a trained first base learner and a trained second base learner; the first base learner and the second base learner are of different types. Based on the first output data during the training of the first base learner, the second output data during the training of the second base learner, and the corresponding labels, the meta-learner is trained to obtain the trained meta-learner. An ensemble classification model is obtained based on the first base learner, the second base learner, and the meta-learner after training.
8. The classification model training method according to claim 7, characterized in that, Also includes: Based on the hypergraph generated during the process of generating the processed unbalanced marine oil and gas data sample set, the validity information of the hypergraph structure is determined. Based on multiple first-class data samples, multiple second-class data samples, and the prediction function of the ensemble classification model, the feature importance analysis information of the features in the multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model is determined; Based on the first performance evaluation metric of the ensemble classification model, the second performance evaluation metric of the trained first base learner, and the third performance evaluation metric of the trained second base learner, the contribution information of the ensemble classification model is determined. Based on the validity information of the hypergraph structure, the feature importance analysis information, and the contribution information of the ensemble classification model, decision analysis information for the ensemble classification model is formed.
9. The classification model training method according to claim 8, characterized in that, Based on the hypergraph generated during the process of generating the processed imbalanced marine oil and gas data sample set, the validity information of the hypergraph structure is determined, including: Based on the hypergraph generated during the process of generating the unbalanced marine oil and gas data sample set, the average size of the hyperedge is determined. The average size of the hyperedge represents the coverage of the local neighborhood between the second type of data samples. The larger the average size of the hyperedge, the denser the distribution of minority class samples with high similarity. Based on the hypergraph generated during the process of generating the processed unbalanced marine oil and gas data sample set, the quality of the hyperedge is determined. The validity information of the hypergraph structure is determined based on the average size and / or quality of the hyperedges.
10. The classification model training method according to claim 8, characterized in that, Based on multiple Class I data samples, multiple Class II data samples, and the prediction function of the ensemble classification model, feature importance analysis information is determined for features in the multiple Class I data samples and multiple Class II data samples during the training of the ensemble classification model, including: Based on multiple first-class data samples, multiple second-class data samples, and the prediction function of the corresponding ensemble classification model, determine the contribution score of each feature of each data sample in the training process of the ensemble classification model. Based on the contribution scores, the global importance information of each feature in multiple first-class data samples and multiple second-class data samples is determined during the classification model training process; Based on the global importance information, feature importance analysis information is determined for features in multiple first-class data samples and multiple second-class data samples during the training of the ensemble classification model.
11. A method for classifying marine oil and gas data, characterized in that, include: Obtain marine oil and gas data to be classified; The marine oil and gas data to be classified includes a first category of data and a second category of data, with the number of data in the first category being greater than the number of data in the second category. The marine oil and gas data to be classified is processed based on the ensemble classification model to obtain the classification result corresponding to the marine oil and gas data to be classified; the ensemble classification model is trained using the classification model training method as described in any one of claims 7-10.
12. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the unbalanced marine oil and gas data processing method as described in any one of claims 1-6, the classification model training method as described in any one of claims 7-10, or the marine oil and gas data classification method as described in claim 11.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the unbalanced marine oil and gas data processing method as described in any one of claims 1-6, the classification model training method as described in any one of claims 7-10, or the marine oil and gas data classification method as described in claim 11.