Automatic classification and grading method and system for sensitive data of oil-gas exploration and development

By performing feature engineering and label propagation algorithms on multi-source heterogeneous data in the data lake, combined with a rule-model hybrid mechanism, the problem of fragmented data security level assessment is solved, and high-accuracy data security level assessment and optimization are achieved, which is suitable for multi-type data processing in the data lake environment.

CN120687870APending Publication Date: 2025-09-23CHINA NATIONAL OFFSHORE OIL (CHINA) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510737334.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies cannot fully utilize the correlation between data to optimize the classification results, cannot achieve effective dissemination and optimization of security levels, lack support for data security level assessment in data lake environments, cannot improve the accuracy of data security level assessment, and cannot effectively deal with new security risks generated during data combination and circulation.

Method used

By performing data feature engineering on multi-source heterogeneous data in the data lake, constructing data feature vectors, and utilizing a hybrid mechanism of label propagation algorithms and rule models, automatic security level assessment of multi-type data is achieved. The label diffusion model is used for iterative optimization, combined with custom grading rules and expert knowledge, to adapt to different business scenarios.

Benefits of technology

It achieves high-accuracy data security level assessment with small sample sizes, supports unified processing of multiple data types, improves the interpretability and reliability of the results, and is suitable for complex heterogeneous data environments in data lakes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687870A_ABST
    Figure CN120687870A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data security assessment, and discloses an automatic classification and grading method and system for sensitive data of oil-gas exploration and development, and the method comprises the steps: carrying out the data feature engineering processing of multi-source heterogeneous data stored in a data lake, extracting name features, business ranges, data features and initial security grading information, constructing a data feature vector; based on the labeled data samples, calculating feature similarity of data feature vectors, constructing a relation network among data items, performing label propagation and security level evaluation, transmitting known security level information to unlabeled data items, and realizing automatic grading; a rule model mixing mechanism is adopted, a data security level classification result and a user feedback iterative optimization security classification result are output through label diffusion model evaluation, feedback data are applied to carry out label diffusion model retraining, and a classification result is obtained. According to the method, the security level automatic evaluation of multiple types of data can be realized through the label propagation algorithm based on the data association relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security assessment, and in particular to a method and system for automatically classifying and grading sensitive data in oil and gas exploration and development. Background Art

[0002] With the continuous evolution of information technology, data has become a fundamental strategic resource in modern society, and data security is increasingly becoming a key issue in information technology development. Data lakes, a new data management architecture, are widely adopted. They can store and manage heterogeneous data, including structured, semi-structured, and unstructured data, providing enterprises with flexible data storage and analysis capabilities.

[0003] The methods disclosed in the prior art have the following problems: they cannot fully utilize the correlation between data to optimize the grading results, cannot achieve effective dissemination and optimization of security levels, lack support for data security level assessment in a data lake environment, cannot fully utilize the propagation relationship between labels to improve the accuracy of data security level assessment, and cannot effectively deal with new security risks generated during data combination and circulation. Summary of the Invention

[0004] In response to the above problems, the purpose of the present invention is to provide a method and system for automated classification and grading of sensitive data in oil and gas exploration and development. Starting from the data association relationship, it can realize automatic assessment of the security level of multiple types of data through a label propagation algorithm, thus meeting the actual needs of data lake governance.

[0005] To achieve the above-mentioned objectives, in the first aspect, the present invention adopts the following technical solutions: a method for automatic classification and grading of sensitive data in oil and gas exploration and development, which includes: performing data feature engineering on multi-source heterogeneous data stored in a data lake, extracting name features, business scope, data features and initial security classification information, and constructing data feature vectors; based on labeled data samples, calculating the feature similarity of the data feature vectors, constructing a relationship network between data items, performing label propagation and security level assessment, and transferring known security level information to unlabeled data items to achieve automatic classification; adopting a rule-model hybrid mechanism, evaluating and outputting data security level classification results through a label diffusion model and iteratively optimizing the security classification results with user feedback, and retraining the label diffusion model using the feedback data to obtain the classification results.

[0006] Furthermore, multi-source heterogeneous data includes exploration data, development and production data, reserve data, marine engineering data, drilling and completion engineering data, and personnel information.

[0007] Furthermore, data feature engineering is performed on the multi-source heterogeneous data stored in the data lake, including: Unify the representation and processing of multi-type data in the data lake to ensure the standardization of heterogeneous data; It uses text embedding and multiple feature encoding strategies to establish semantic associations between data items and supports similarity calculations between complex data; multiple feature encoding strategies include one-hot encoding and label encoding.

[0008] Furthermore, based on the labeled data samples, the feature similarity of the data feature vectors is calculated, the relationship network between the data items is constructed, and label propagation and security level assessment are performed, including: Based on the core mechanism of the label propagation algorithm, combined with KNN or RBF kernel function, the similarity matrix between data items is calculated, and the parameter γ is configurable to optimize the similarity sensitivity; The configurable propagation parameter α is used to control the propagation intensity and stability of the security level information of known data items to unlabeled data items.

[0009] Furthermore, a rule-model hybrid mechanism is adopted to iteratively optimize the security classification results through label diffusion model evaluation, output data security level classification results, and user feedback. The feedback data is used to retrain the label diffusion model to obtain classification results, including: Build a rule model hybrid mechanism to integrate customized classification rules with the priority of data security level classification results, and allow expert knowledge and data-driven results to dynamically complement each other to adapt to different business scenarios; Conduct multi-dimensional verification of the security level assessment results to determine the precision, recall, and F1 score for each security level type; The small sample optimization strategy is adopted to obtain the classification results.

[0010] Furthermore, the customized grading rules are as follows: building a customized grading rule engine that supports the definition of rules in the format of "column name, pattern, target label".

[0011] Furthermore, the security level assessment results are verified in multiple dimensions to determine the precision, recall, and F1 score for each security level type, including: Calculate the overall indicators of data security level classification results, such as accuracy, precision, recall, and F1 score, to evaluate the performance of the label diffusion model; Calculate the precision, recall and F1 score for each safety level category respectively; Visualization technology is used to intuitively present data distribution and prediction results, supporting parameter optimization.

[0012] In the second aspect, the technical solution adopted by the present invention is: an automatic classification and grading system for sensitive data in oil and gas exploration and development, which includes: a data processing module, which performs data feature engineering on multi-source heterogeneous data stored in the data lake, extracts name features, business scope, data features and initial security classification information, and constructs data feature vectors; a label propagation and security level assessment module, which calculates the feature similarity of data feature vectors based on labeled data samples, constructs a relationship network between data items, performs label propagation and security level assessment, and transfers known security level information to unlabeled data items to achieve automatic classification; a classification output module, which adopts a rule-model hybrid mechanism, evaluates and outputs data security level classification results through a label diffusion model, iteratively optimizes the security classification results with user feedback, and uses the feedback data to retrain the label diffusion model to obtain the classification results.

[0013] In a third aspect, the technical solution adopted by the present invention is: a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device executes any one of the above methods.

[0014] In a fourth aspect, the technical solution adopted by the present invention is: a computing device, comprising: one or more processors, a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.

[0015] The present invention has the following advantages due to the adoption of the above technical solution: 1. The present invention can reduce the demand for labeled data and can still achieve high-accuracy security level assessment even with a small sample size.

[0016] 2. The present invention makes full use of data association and realizes efficient transmission of security level information through label propagation algorithm.

[0017] 3. This invention supports unified processing of multiple data types and is suitable for complex heterogeneous data environments in data lakes.

[0018] 4. The rule and model hybrid mechanism adopted in this invention improves the interpretability and reliability of the results.

[0019] 5. The modular design of the present invention makes the system have good scalability and configurability. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is an overall flow chart of the method for automatically classifying and grading sensitive data for oil and gas exploration and development according to an embodiment of the present invention; Figure 2This is a detailed flow chart of the method for automatically classifying and grading sensitive data for oil and gas exploration and development in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The prior art publication "Research on Automated Paths for Enterprise Data Classification and Grading" proposes a data grading method based on automated identification and labeling. This method uses keyword matching and machine learning algorithms to identify sensitive data, and combines word vectors and contextual information for classification, grading, and labeling. However, this method separates the classification and grading processes and only dynamically adjusts based on data volume thresholds, failing to fully utilize the relationships between data to optimize grading results.

[0022] The prior art paper "Power Plant Data Security Protection System Prototype Design" proposes a data security protection method based on semi-supervised learning. This method achieves data classification and protection through a pre-set hierarchical structure and deep learning models. However, this method fails to fully utilize the rich relationships between data in a data lake environment. Its pre-set hierarchical structure and independent processing mechanisms are difficult to adapt to the diversity and complexity of data in the data lake, and cannot effectively disseminate and optimize security levels.

[0023] The prior art publication "Method and System for Automated Data Identification, Classification, and Grading" proposes a data classification method based on recognition, classification, and grading models. This method preprocesses the data source, dividing it into static and dynamic data, then performs constraint and attribute condition judgments, respectively, and employs comparison, benchmarking, and calculation methods for classification. However, this method simply labels the graded data, failing to consider the correlation between data or leverage the propagation relationships between labels. This method, therefore, provides insufficient support for data security assessment in data lake environments.

[0024] The prior art document "Method and System for Automated Data Identification, Classification, and Grading" (CN118332407A) proposes a data classification method based on recognition, classification, and grading models. This method preprocesses data sources, sequentially identifying, classifying, and grading them, ultimately creating labels and performing catalog management. However, this method uses a linear processing flow, fails to consider data relevance, and uses labels only to identify results. Consequently, it fails to fully leverage the propagation relationships between labels to improve the accuracy of data security assessments.

[0025] The prior art publication "Data Asset Classification Modeling and Hierarchical Protection Method Based on Artificial Intelligence Technology" proposes a classification and grading method that supports access to multiple data sources. This method connects to data sources through various methods, such as MDBC and ODBC, employs regular expressions and machine learning to model data categories, and automatically adapts security policies. However, this method focuses excessively on data source access and sampling, while oversimplifying the assessment of data security levels. Sampling-based analysis may miss important security features, and the pre-set security policy lacks the ability to respond to dynamic changes in data usage scenarios.

[0026] The prior art publication "Artificial Intelligence-Based Power Grid Data Security Classification Method and System" proposes a security classification method for power grid data. This method categorizes power grid data into four security levels based on data type and implements access control through employee ID verification and password encryption. However, this method uses mechanical data type matching for security classification, ignoring the dynamic nature of data value and risk. Furthermore, it over-relies on static password access control mechanisms, making it ineffective in addressing new security risks arising from data combination and transfer.

[0027] In response to the above problems, the present invention proposes a method and system for automated classification and grading of sensitive data in oil and gas exploration and development. The method of the present invention has advantages in scenarios with small sample sizes; the reliability of the evaluation results is ensured by an iterative convergence mechanism and known label coverage. At the same time, the present invention realizes the unified representation and association establishment of various data types such as classification, numerical values, and text in the data lake through text vectorization and feature encoding strategies, and innovatively adopts a rule-model hybrid mechanism to combine custom grading rules with model predictions, significantly enhancing the interpretability of the results. In addition, the present invention also designs a hierarchical label propagation mechanism, which achieves more accurate security level assessment by considering the hierarchical relationship between different security levels and combining configurable propagation strength and kernel functions. The present invention can not only achieve high-accuracy data classification in small sample conditions, but also make full use of the correlation between data in the data lake environment, effectively solving the problem of fragmented and static data security level assessment in the existing technology.

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.

[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0030] In one embodiment of the present invention, a method for automatic classification and grading of sensitive data for oil and gas exploration and development is provided. In this embodiment, the present invention realizes automatic classification of data security through four core modules: data access, feature extraction, security classification, and model optimization. Based on the label propagation algorithm and rule mixing mechanism, it supports unified processing of heterogeneous data in a data lake environment, and improves the interpretability and accuracy of classification results. Figure 1 As shown, the method includes the following steps: 1) Perform data feature engineering on the multi-source heterogeneous data stored in the data lake, extract name features, business scope, data features, and initial security classification information, and construct data feature vectors; 2) Based on labeled data samples, the feature similarity of data feature vectors is calculated, a relationship network between data items is constructed, label propagation and security level assessment are performed, and known security level information is transferred to unlabeled data items to achieve automatic classification; 3) A rule-model hybrid mechanism is adopted to iteratively optimize the security classification results through label diffusion model evaluation, output data security level classification results and user feedback to enhance the interpretability of the classification results; and the feedback data is used to retrain the label diffusion model to obtain the classification results.

[0031] In this embodiment, the data in each step is stored in the data layer. Specifically, the data processed by data feature engineering and the data subjected to label propagation and security level assessment are stored in the data feature library; and the classification results are stored in the data classification library.

[0032] In the above step 1), the multi-source heterogeneous data includes exploration data, development and production data, reserve data, marine engineering data, drilling and completion engineering data, and personnel information.

[0033] In step 1) above, data feature engineering is performed on the multi-source heterogeneous data stored in the data lake, including the following steps: 1.1) Unify the representation and processing of multi-type data in the data lake to ensure the standardization of heterogeneous data; Among them, multi-type data includes classification, numerical, text, etc.

[0034] 1.2) Use text embedding and various feature encoding strategies to establish semantic associations between data items and support similarity calculations between complex data; various feature encoding strategies include one-hot encoding and label encoding.

[0035] In step 2) above, based on the labeled data samples, the feature similarity of the data feature vectors is calculated, the relationship network between the data items is constructed, and label propagation and security level assessment are performed, including the following steps: 2.1) Based on the core mechanism of the Label Propagation algorithm, this algorithm combines the KNN or RBF kernel function to calculate the similarity matrix between data items. The parameter γ can be configured to optimize similarity sensitivity. The parameter γ refers to the gamma parameter in the RBF kernel function. By adjusting γ, the sensitivity of the similarity calculation between data points to distance can be finely controlled, thereby optimizing the performance of the Label Propagation algorithm.

[0036] The core label propagation mechanism employed in this embodiment is based on the principles of a mature and widely used semi-supervised learning algorithm. This mechanism first quantifies the similarity between data points based on the characteristics of the input data items using either the RBF kernel function or the KNN method, thereby constructing a similarity matrix. Specifically, when using the RBF kernel function, the parameter γ can be adjusted to fine-tune the sensitivity of the similarity calculation to the distance between data points, thereby optimizing the quality of the similarity matrix and providing a more reliable basis for subsequent effective label propagation.

[0037] 2.2) The intensity and stability of the propagation of security level information of known data items (classified as general, important, and core) to unlabeled data items is controlled through a configurable propagation parameter α, which refers to the degree of trust in the initial known security level.

[0038] In step 3) above, a rule-model hybrid mechanism is used to iteratively optimize the security classification results through label diffusion model evaluation, output data security level classification results, and user feedback. The feedback data is used to retrain the label diffusion model to obtain the classification results, including the following steps: 3.1) Build a rule-model hybrid mechanism to integrate custom grading rules with the priority of data security level classification results, improving classification interpretability. This also allows for dynamic complementarity between expert knowledge and data-driven results to adapt to different business scenarios. Among them, the custom grading rules are: building a custom grading rule engine, supporting the definition of rules in the format of "column name, mode, target label".

[0039] 3.2) Conduct multi-dimensional verification of the safety level assessment results to determine the precision, recall, and F1 score for each safety level type; In this embodiment, specifically, multi-dimensional verification of the security level assessment results is performed to determine the precision, recall, and F1 score of each security level type, including the following steps: 3.2.1) Calculate overall metrics such as accuracy, precision, recall, and F1 score of the data security level classification results to evaluate the performance of the label diffusion model; 3.2.2) Calculate the precision, recall, and F1 score for each safety level category, where the safety levels are general, important, and core. 3.2.3) Visualize data distribution and prediction results to support parameter optimization.

[0040] 3.3) A small sample optimization strategy is used to obtain the classification results. Specifically, a label diffusion algorithm is used to infer the data security level information based on a small number of labeled samples and assign it to the unlabeled data, thus completing the overall classification.

[0041] Aiming at the scenario where labeled samples are scarce in data lake environments, a semi-supervised learning parameter optimization scheme is designed. Experimental verification shows that a classification accuracy of over 89% can still be achieved under limited sample conditions (50% labeled).

[0042] Example: Assuming that an energy company is carrying out digital transformation activities in the process of oil exploration and production, it is necessary to classify and grade the sensitive data generated by the various types of data. The method provided by the present invention can be used to perform the operation. Figure 2 As shown, the present invention includes four core steps: data access, feature extraction, security classification and model optimization.

[0043] First, the system accesses and processes multi-source heterogeneous data from energy companies, including exploration data, development and production data, reserves data, marine engineering data, drilling and completion engineering data, and personnel information. Using a single-table field-level feature extraction module, it extracts name features, business scope, data features, and initial security classification information to create a data feature vector representation. Secondly, the intelligent data security classification module uses a small number of labeled data samples to calculate feature similarity and construct a relationship network between data items using KNN or RBF kernel functions. It then uses a label propagation algorithm to transfer known security level information (general / important / core) to unlabeled data items, achieving automatic classification. Thirdly, the system iteratively optimizes the security classification results through model evaluation, output results, and user feedback. It uses feedback data to retrain the model to continuously improve the classification accuracy. At the same time, the processing results are stored in the data feature library and data classification library. Finally, the classification results are applied to the formulation of data access control policies, and differentiated protection measures are implemented for data of different security levels.

[0044] Through the method provided by this invention, the energy enterprise can achieve efficient and automatic classification of massive amounts of heterogeneous data in exploration and production business scenarios, meet the compliance requirements of the "Data Security Law" and industry data security regulations, and ensure data security. The method of this invention is not only applicable to the energy industry, but also has broad application in industries such as finance, healthcare, and transportation that require strict data security management.

[0045] In one embodiment of the present invention, a system for automatically classifying and grading sensitive data in oil and gas exploration and development is provided, comprising: The data processing module performs data feature engineering on the multi-source heterogeneous data stored in the data lake, extracts name features, business scope, data features, and initial security classification information, and constructs data feature vectors; The label propagation and security level assessment module calculates the feature similarity of data feature vectors based on labeled data samples, constructs a relationship network between data items, performs label propagation and security level assessment, and transfers known security level information to unlabeled data items to achieve automatic classification; The classification output module adopts a rule-model hybrid mechanism to evaluate the label diffusion model, output the data security level classification results and iteratively optimize the security classification results based on user feedback, and then retrain the label diffusion model using the feedback data to obtain the classification results.

[0046] In the above embodiment, the multi-source heterogeneous data includes exploration data, development and production data, reserve data, marine engineering data, drilling and completion engineering data, and personnel information.

[0047] In this embodiment, data feature engineering is performed on multi-source heterogeneous data stored in a data lake, including: Unify the representation and processing of multi-type data in the data lake to ensure the standardization of heterogeneous data; It uses text embedding and multiple feature encoding strategies to establish semantic associations between data items and supports similarity calculations between complex data; multiple feature encoding strategies include one-hot encoding and label encoding.

[0048] In the above embodiment, based on the labeled data samples, the feature similarity of the data feature vectors is calculated, a relationship network between data items is constructed, and label propagation and security level assessment are performed, including: Based on the core mechanism of the label propagation algorithm, combined with KNN or RBF kernel function, the similarity matrix between data items is calculated, and the parameter γ is configurable to optimize the similarity sensitivity; The configurable propagation parameter α is used to control the propagation intensity and stability of the security level information of known data items to unlabeled data items.

[0049] In the above embodiment, a rule-model hybrid mechanism is adopted to iteratively optimize the security classification results through label diffusion model evaluation, output data security level classification results, and user feedback. The feedback data is used to retrain the label diffusion model to obtain the classification results, including: Build a rule model hybrid mechanism to integrate customized classification rules with the priority of data security level classification results, and allow expert knowledge and data-driven results to dynamically complement each other to adapt to different business scenarios; Conduct multi-dimensional verification of the security level assessment results to determine the precision, recall, and F1 score for each security level type; The small sample optimization strategy is adopted to obtain the classification results.

[0050] Among them, the custom grading rules are: building a custom grading rule engine, supporting the definition of rules in the format of "column name, mode, target label".

[0051] In this embodiment, the security level assessment results are verified in multiple dimensions to determine the precision, recall, and F1 score for each security level type, including: Calculate the overall indicators of data security level classification results, such as accuracy, precision, recall, and F1 score, to evaluate the performance of the label diffusion model; Calculate the precision, recall and F1 score for each safety level category respectively; Visualization technology is used to intuitively present data distribution and prediction results, supporting parameter optimization.

[0052] The system provided in this embodiment is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for specific processes and detailed contents, which will not be repeated here.

[0053] In one embodiment of the present invention, a computing device is provided. The computing device may be a terminal and may include: a processor, a communications interface, a memory, a display screen, and an input device. The processor, communications interface, and memory communicate with each other via a communications bus. The processor is configured to provide computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. When executed by the processor, the computer program implements the methods described in the above embodiments. The internal memory provides an environment for the operating system and computer program in the non-volatile storage medium to run. The communications interface is configured to communicate with an external terminal via wired or wireless communication. The wireless communication may be achieved via Wi-Fi, a network management service provider, NFC (near field communication), or other technologies. The display screen may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen layer covering the display screen, or may be buttons, a trackball, or a touchpad provided on the computing device housing, or may be an external keyboard, touchpad, or mouse. The processor may invoke logic instructions stored in the memory.

[0054] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0055] In one embodiment of the present invention, a computer program product is provided, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned method embodiments.

[0056] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores server instructions. The computer instructions enable a computer to execute the methods provided in the above embodiments.

[0057] The above embodiment provides a computer-readable storage medium, whose implementation principle and technical effects are similar to those of the above method embodiment, and will not be repeated here.

[0058] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0059] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for automatically classifying and grading sensitive data in oil and gas exploration and development, characterized in that: include: Perform data feature engineering on multi-source heterogeneous data stored in the data lake, extract name features, business scope, data features, and initial security classification information, and construct data feature vectors; Based on labeled data samples, the feature similarity of data feature vectors is calculated, a relationship network between data items is constructed, label propagation and security level assessment are performed, and known security level information is transferred to unlabeled data items to achieve automatic classification; A rule-model hybrid mechanism is adopted to evaluate the label diffusion model, output data security level classification results and iteratively optimize the security classification results based on user feedback. The feedback data is used to retrain the label diffusion model to obtain the classification results.

2. The method for automatically classifying and grading sensitive data for oil and gas exploration and development according to claim 1, characterized in that: Multi-source heterogeneous data includes exploration data, development and production data, reserve data, marine engineering data, drilling and completion engineering data, and personnel information.

3. The method for automatically classifying and grading sensitive data for oil and gas exploration and development according to claim 2, characterized in that: Perform data feature engineering on multi-source heterogeneous data stored in the data lake, including: Unify the representation and processing of multi-type data in the data lake to ensure the standardization of heterogeneous data; It uses text embedding and multiple feature encoding strategies to establish semantic associations between data items and supports similarity calculations between complex data; multiple feature encoding strategies include one-hot encoding and label encoding.

4. The method for automatically classifying and grading sensitive data for oil and gas exploration and development according to claim 1, wherein: Based on labeled data samples, the feature similarity of data feature vectors is calculated, a relationship network between data items is constructed, and label propagation and security level assessment are performed, including: Based on the core mechanism of the label propagation algorithm, combined with KNN or RBF kernel function, the similarity matrix between data items is calculated, and the parameter γ is configurable to optimize the similarity sensitivity; The configurable propagation parameter α is used to control the propagation intensity and stability of the security level information of known data items to unlabeled data items.

5. The method for automatically classifying and grading sensitive data for oil and gas exploration and development according to claim 1, characterized in that: A rule-model hybrid mechanism is used to evaluate the label diffusion model, output data security level classification results, and iteratively optimize the security classification results based on user feedback. The feedback data is used to retrain the label diffusion model to obtain classification results, including: Build a rule model hybrid mechanism to integrate customized classification rules with the priority of data security level classification results, and allow expert knowledge and data-driven results to dynamically complement each other to adapt to different business scenarios; Conduct multi-dimensional verification of the security level assessment results to determine the precision, recall, and F1 score for each security level type; The small sample optimization strategy is adopted to obtain the classification results.

6. The method for automatically classifying and grading sensitive data for oil and gas exploration and development according to claim 5, characterized in that: Custom grading rules are: building a custom grading rule engine that supports the definition of rules in the format of "column name, mode, target label".

7. The method for automatically classifying and grading sensitive data for oil and gas exploration and development according to claim 5, characterized in that: The safety level assessment results are verified in multiple dimensions to determine the precision, recall, and F1 score for each safety level type, including: Calculate the overall indicators of data security level classification results, such as accuracy, precision, recall, and F1 score, to evaluate the performance of the label diffusion model; Calculate the precision, recall and F1 score for each safety level category respectively; Visualization technology is used to intuitively present data distribution and prediction results, supporting parameter optimization.

8. An automated classification and grading system for sensitive data in oil and gas exploration and development, characterized by: include: The data processing module performs data feature engineering on the multi-source heterogeneous data stored in the data lake, extracts name features, business scope, data features, and initial security classification information, and constructs data feature vectors; The label propagation and security level assessment module calculates the feature similarity of data feature vectors based on labeled data samples, constructs a relationship network between data items, performs label propagation and security level assessment, and transfers known security level information to unlabeled data items to achieve automatic classification; The classification output module adopts a rule-model hybrid mechanism to evaluate the label diffusion model, output the data security level classification results and iteratively optimize the security classification results based on user feedback, and then retrain the label diffusion model using the feedback data to obtain the classification results.

9. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any one of the methods of claims 1 to 7 .

10. A computing device, characterized in that include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods according to claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for automatically identifying, classifying and grading data

    CN118332407A