Enterprise intelligent data space construction method based on big data

By constructing an enterprise intelligent data space based on big data, the heterogeneity and security risks of multimodal data in oil production enterprises have been solved, enabling efficient data management and accurate prediction of leakage and diffusion, thereby improving production management and safety.

CN121458142AInactive Publication Date: 2026-02-03东营职业学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511613781.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Oil production companies face heterogeneity, spatiotemporal asynchrony, and security risks in their multimodal data, and lack effective data storage and management solutions.

Method used

We construct an enterprise intelligent data space based on big data, collect multimodal data through a distributed optical fiber sensor network, build a data lake with a unified spatiotemporal benchmark using a protocol conversion gateway and spatiotemporal alignment algorithm, construct a three-layer knowledge graph using a graph attention network, and achieve secure sharing and efficient management by combining hierarchical encryption algorithms.

Benefits of technology

It achieves efficient fusion and unified management of multimodal data, improves cross-node data query efficiency, enhances production decision-making capabilities, improves the accuracy and security of leakage and diffusion prediction, and reduces spatiotemporal alignment errors and response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121458142A_ABST
    Figure CN121458142A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise intelligent data space construction method based on big data, and belongs to the technical field of petroleum production, and the method comprises the steps: obtaining the multi-modal production data of a petroleum production enterprise, the petroleum production enterprise comprises a well site node, an operation node and a headquarter node, the multi-modal production data comprises well site data, operation area data and headquarter data; constructing a production knowledge graph of the petroleum production process based on the multi-modal production data; a production prediction model is constructed based on the production knowledge graph, and the production prediction model is used for detecting whether leakage diffusion exists in the petroleum production process or not; and the production prediction model and the production knowledge graph are safely shared in the petroleum production enterprise based on the hierarchical encryption algorithm, so that the enterprise data space of the petroleum production enterprise is obtained, and the effect of synchronously storing the petroleum production data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of oil production, and particularly relates to a construction method of an enterprise intelligent data space based on big data. BACKGROUND

[0002] In the field of oil production, an enterprise is usually composed of multiple nodes (well sites, operation zones, headquarters), and the multi-modal data (such as downhole sensing data, operation zone SCADA system data, and headquarters management data) generated by each node has significant heterogeneity, time-space asynchrony, and security risks.

[0003] In view of the problem of asynchronous storage of oil production data in the related art, no effective solution has been proposed so far. SUMMARY

[0004] In view of the above technical problems, the application provides a construction method of an enterprise intelligent data space based on big data, which is applied to an oil production enterprise and includes the following steps. Obtaining multi-modal production data of the oil production enterprise, wherein the oil production enterprise includes well site nodes, operation nodes, and headquarters nodes, and the multi-modal production data includes well site data, operation zone data, and headquarters data; Constructing a production knowledge graph of an oil production process based on the multi-modal production data; Constructing a production prediction model based on the production knowledge graph, wherein the production prediction model is used to detect whether there is leakage diffusion in the oil production process; Securely sharing the production prediction model and the production knowledge graph in the oil production enterprise based on a hierarchical encryption algorithm to obtain an enterprise data space of the oil production enterprise.

[0005] Further, the step of obtaining the multi-modal production data of the oil production enterprise includes the following steps: collecting well site data through a distributed optical fiber sensing network at a target sampling frequency, wherein the well site data includes downhole pressure, temperature, and flow parameters; obtaining operation zone data from an operation zone SCADA system through a protocol conversion gateway; obtaining management data from a headquarters ERP system, and performing time-space alignment on the management data, the well site data, and the operation zone data to obtain headquarters data; and constructing an oil production data lake with a unified time-space reference as the multi-modal production data based on the well site data, the operation zone data, and the headquarters data.

[0006] Further, the production knowledge graph of the oil production process is constructed based on the multi-modal production data, comprising: adopting a graph attention network to mine the association relationship among wellbores, equipment and processes based on the multi-modal production data; constructing a three-layer architecture knowledge graph including a wellbore layer, an equipment layer and a process layer; wherein the wellbore layer includes well structure attributes and inter-well topological relationships, the equipment layer includes equipment parameter templates and failure mode libraries, and the process layer includes process constraint conditions and production process rules.

[0007] Further, the production knowledge graph further comprises: establishing an equipment failure mode library based on API standards, defining the mapping relationship between vibration characteristics and equipment failure; setting water injection process parameter constraint conditions, including pressure range and water quality index threshold; establishing a knowledge reasoning engine to realize accurate early warning with a fault diagnosis confidence of ≥0.92.

[0008] Further, the production prediction model and the production knowledge graph are securely shared in the oil production enterprise based on a hierarchical encryption algorithm, obtaining an enterprise data space of the oil production enterprise, comprising: designing and deploying encryption algorithms at the well site production node, the operation area production node and the headquarters production node, wherein the encryption algorithms comprise: using SM4 algorithm to encrypt real-time sensing data at the well site node, deploying attribute-based SM9 encryption scheme at the operation node, and implementing homomorphic encryption supporting aggregated calculation at the headquarters node; establishing a priori verification mechanism, generating SHA3-256 instruction fingerprint before data transmission, and performing complete transmission after verification pass rate ≥ pass threshold.

[0009] Further, the enterprise data space further comprises: developing a hybrid drive reservoir model, coupling physical equations and LSTM prediction algorithm, satisfying ‖P_data-P_physics‖<0.1MPa; deploying a CFD accelerated leakage diffusion prediction system with a response time ≤30 seconds; wherein the physical equation solves the pressure field distribution by using black oil model, and the input features of the LSTM prediction algorithm include production history data and real-time working condition data.

[0010] The beneficial effects of this invention compared with the prior art are: (1) Efficient fusion and unified management of multimodal data: Through distributed optical fiber sensor network, protocol conversion gateway and spatiotemporal alignment algorithm, a high-precision oil production data lake is constructed, which reduces the spatiotemporal alignment error of well site, work area and headquarters data; The data storage architecture with unified spatiotemporal reference is adopted to improve the efficiency of cross-node data query and support real-time production monitoring and historical data analysis; (2) Enhanced production decision-making capability of dynamic knowledge graph: Based on the graph attention network (GAT) three-layer architecture knowledge graph (well barrel layer, equipment layer and process layer), dynamic correlation analysis of well barrel-equipment-process is realized, which improves the confidence of fault diagnosis; (3) High-precision leakage diffusion prediction and safety early warning: The hybrid driving model (black oil physical equation + LSTM prediction algorithm) is adopted to make the pressure field prediction error ‖P_data-P_physics‖<0.1MPa. Attached Figure Description

[0011] Figure 1 This is a flowchart of the method for constructing an enterprise intelligent data space based on big data according to the present invention. Detailed Implementation

[0012] Example: Figure 1 As shown, a method for constructing an enterprise intelligent data space based on big data, applied to an oil production enterprise, includes: Acquire multimodal production data of the oil production enterprise, wherein the oil production enterprise includes well site nodes, operation nodes and headquarters nodes, and the multimodal production data includes well site data, operation area data and headquarters data; A production knowledge graph of the petroleum production process is constructed based on the aforementioned multimodal production data; A production prediction model is constructed based on the production knowledge graph, wherein the production prediction model is used to detect whether there is leakage and diffusion in the oil production process; Based on a hierarchical encryption algorithm, the production prediction model and the production knowledge graph are securely shared among the oil production enterprises to obtain the enterprise data space of the oil production enterprises.

[0013] Optionally, in this embodiment, the aforementioned oil production enterprise is a comprehensive enterprise engaged in oil exploration, extraction, transportation, processing and sales, and typically adopts a three-tier management model of "well site - operating area - headquarters".

[0014] Optionally, in this embodiment, the multi-modal production data mentioned above are heterogeneous data of different types and different sources generated in the oil production process, mainly including: structured data (such as production reports in the database, equipment parameters), time series data (such as pressure, temperature, flow collected by downhole sensors), unstructured data (such as geological exploration reports, equipment maintenance logs), industrial control system data (such as real-time monitoring data of SCADA system), etc.

[0015] Optionally, in this embodiment, the data collection and management of the oil production enterprise is generally divided into three levels: well site node, i.e. the production site of a single oil well, whose data can include but is not limited to downhole pressure, temperature, flow (real-time sensing data), etc.; operation node, i.e. the regional center managing multiple well sites, whose data can include but is not limited to SCADA system data, equipment status, production scheduling, etc.; headquarters node, i.e. enterprise-level decision-making and management, whose data can include but is not limited to ERP system data, financial data, supply chain data, etc.

[0016] Optionally, in this embodiment, well site data is obtained through distributed optical fiber sensing and downhole sensors; operation area data is obtained from SCADA system and equipment monitoring; headquarters data is obtained from ERP system and management reports.

[0017] Optionally, in this embodiment, the production knowledge graph is a structured knowledge base used to describe the correlation between equipment, process and wellbore in oil production, including: wellbore layer (well structure, inter-well topological relationship); equipment layer (equipment parameters, fault mode library); process layer (water injection process constraints, production process rules); reasoning engine (such as fault diagnosis based on API standard). It can be used but not limited to when downhole pressure is abnormal, the knowledge graph can automatically associate possible equipment faults (such as pump valve damage) and give maintenance suggestions.

[0018] Optionally, in this embodiment, the production prediction model is an AI model for predicting anomalies (such as leaks, equipment failures) in the oil production process, which usually uses: physical models (such as black oil equation to simulate reservoir pressure), data-driven models (such as LSTM to predict leakage risk), hybrid models (combining physical laws and machine learning), etc. For example: input real-time pressure and flow data, the model outputs leakage probability, and triggers an alarm when the threshold is exceeded.

[0019] Optionally, in this embodiment, the oil production process can be detected for leakage diffusion by the following methods: CFD (Computational Fluid Dynamics) simulation of leakage diffusion path; LSTM model analysis of historical leakage data; hybrid model combining physical equations (such as Darcy's law) and AI prediction; real-time warning (response time ≤ 30 seconds).

[0020] Optionally, in this embodiment, different encryption strategies are adopted for different node data security requirements: SM4 (national symmetric encryption) is used at the well site node to realize real-time data encryption (low delay); SM9 (attribute-based encryption) is used at the operation node to realize permission control (such as only the operation area manager can decrypt); homomorphic encryption (supporting ciphertext calculation) is used at the headquarters node to realize data aggregation analysis (such as calculating the total oilfield production).

[0021] Optionally, in this embodiment, an efficient, secure, and intelligent oil production data space is constructed by multi-modal data fusion + knowledge graph + intelligent prediction + hierarchical encryption in the above method, which significantly improves the production management level and security.

[0022] In an optional embodiment, the multi-modal production data of the oil production enterprise is obtained by: collecting the well site data through a distributed optical fiber sensing network at a target sampling frequency, wherein the well site data includes downhole pressure, temperature, and flow rate parameters; obtaining the operation area data from the operation area SCADA system through a protocol conversion gateway; obtaining management data from the headquarters ERP system, and performing spatio-temporal alignment of the management data with the well site data and the operation area data to obtain the headquarters data; and based on the well site data, the operation area data, and the headquarters data, constructing an oil production data lake with a unified spatio-temporal reference as the multi-modal production data.

[0023] Optionally, in this embodiment, the target sampling frequency is a data acquisition rate preset according to well conditions, which can be but is not limited to ≥100Hz.

[0024] Optionally, in this embodiment, collecting the well site data through a distributed optical fiber sensing network at a target sampling frequency includes: A multi-modal distributed optical fiber (Φ0.25mm) is laid on the outer wall of the target oil well casing, Φ-OTDR (phase-sensitive optical time domain reflectometry) technology is used to realize full wellbore coverage, the sampling frequency is set to 1kHz (for high-pressure wells) or 100Hz (for conventional wells), the spatial resolution is 1 meter, the temperature measurement accuracy is ±0.5℃, and the strain measurement range is ±5000με.

[0025] The original optical signal is demodulated by the edge computing node to extract: pressure data (based on the optical fiber strain-pressure conversion model, error <0.1MPa), temperature data (DTS temperature measurement, one data point every 10 meters), and flow rate data (combined with V-cone flowmeter calibration, accuracy ±2%); a sliding window filter (window size 5s) is executed to eliminate noise; and a structured well site data packet (JSON format, containing timestamp, well number, coordinates, and parameter value) is output.

[0026] Optionally, in this embodiment, obtaining the job site data from the job site SCADA system via the protocol conversion gateway includes: Establishing an OPC UA-MQTT gateway: deploying an industrial protocol conversion gateway at the job site control center, configuring: SCADA side: OPC UA protocol (port 4840), reading frequency 10 Hz; data lake side: MQTT protocol, QoS level 1; mapping key points: Pump101.RPM→equipment.rpm; Valve203.Position→pipeline.valve_open_degree.

[0027] Performing unit conversion (psi→MPa, °F→℃), outlier rejection (3σ principle), and time alignment (NTP server synchronization, error <1 ms) on the collected equipment state data (such as pump frequency, valve opening degree) to obtain a time series database (InfluxDB structure) as the job site data.

[0028] Optionally, in this embodiment, obtaining the management data from the headquarters ERP system and performing spatio-temporal alignment of the management data with the well site data and the job site data to obtain the headquarters data includes: Synchronizing daily: production schedule (PROD_PLAN), equipment maintenance record (MAINT_LOG), and inventory (INVENTORY) via SAP HANA Connector; Generating a unified time series (1-minute granularity) by linear interpolation on well site data (1kHz), SCADA data (10Hz), and ERP data (1 / day); Establishing a well coordinate-job site-headquarters mapping relationship table (GIS geographic coding), converting device location to grid encoding (precision level 7) via GeoHash algorithm to obtain a spatio-temporal alignment data table (Parquet format) as the job site data.

[0029] Optionally, in this embodiment, constructing a unified spatio-temporal reference oil production data lake as the multi-modal production data includes: Designing a Delta Lake storage architecture: Bronze layer: raw data (S3 storage, retaining all fields); Silver layer: cleaned data (Parquet format, ZSTD compression); Gold layer: business aggregation table (aggregated by day / week / month).

[0030] Registering in Apache Atlas: data lineage (such as fiber sensing→edge node→data lake); business terminology (such as downhole pressure = Wellbore.Pressure).

[0031] In an optional embodiment, the constructing a production knowledge graph of the oil production process based on the multi-modal production data comprises: mining an association relationship among a wellbore, equipment and a process based on the multi-modal production data by using a graph attention network; and constructing a three-layer architecture knowledge graph comprising a wellbore layer, an equipment layer and a process layer; wherein the wellbore layer comprises wellbore structure attributes and inter-well topological relationships, the equipment layer comprises equipment parameter templates and a failure mode library, and the process layer comprises process constraint conditions and production process rules.

[0032] Optionally, in the embodiment, the oil production knowledge graph based on the graph attention network can be constructed by the following steps, but is not limited to the following steps: Step 201, data preparation and graph structure definition: Obtaining wellbore data, equipment data and process data: wellbore data, obtaining wellbore structure data (casing program, inclination) from a drilling database, and real-time downhole sensor data (pressure / temperature); equipment data, pump unit operation parameters (vibration, current, temperature) provided by a SCADA system, and equipment maintenance records; process data, water injection schemes in a production management system, and safety operation procedure documents.

[0033] Node and relationship definition: Defining node types: wellbore nodes (containing attributes such as well depth and casing size), equipment nodes (containing attributes such as equipment model and operation parameters), and process nodes (containing attributes such as pressure range and water quality standards).

[0034] Defining edge relationships: physical connection relationships (such as a “wellbore-pumping unit” connection), and logical influence relationships (such as an “injection pressure-casing stress” association).

[0035] Step 202, graph attention network modeling process: Model training process: Feature encoding: direct vectorization of structured data; feature extraction of text data by using a BERT model; and feature extraction of time series data by using a 1D-CNN; Attention weight calculation: dynamically assigning weights to the connection relationships among the three types of nodes (such as “wellbore-equipment-process”) (for example: when analyzing casing damage, the model automatically strengthens the association weight of “cementing quality-injection pressure”).

[0036] Step 203, three-layer knowledge graph architecture: Wellbore layer implementation: Wellbore structure attributes: casing steel grade, wall thickness, depth; cementing quality score (CBL / VDL logging data); Inter-well topological relationships: spatial distance matrix; and production layer connectivity analysis results.

[0037] Device layer implementation: Parameter template: centrifugal pump: rated speed, allowable vibration value, efficiency curve; electric submersible pump: motor temperature threshold, current fluctuation range; Fault mode library: fault feature library based on API RP 11S1 standard; field maintenance case library (including vibration spectrum feature map).

[0038] Process layer implementation: Constraint condition: water injection pressure window (such as 15-22 MPa), water quality index (oil content ≤15 mg / L); Flow rule: IF-THEN rule (such as "if pump vibration is out of limit and wellhead pressure drops, then perform shutdown check"); process optimization strategy library (such as acidizing and fracturing scheme selection tree).

[0039] In an optional embodiment, the production knowledge graph further comprises: establishing a device fault mode library based on API standards, defining the mapping relationship between vibration features and device faults; setting water injection process parameter constraints, including pressure range and water quality index threshold; establishing a knowledge reasoning engine to realize accurate early warning with fault diagnosis confidence ≥0.92.

[0040] In an optional embodiment, the production prediction model and the production knowledge graph are securely shared in the oil production enterprise based on a hierarchical encryption algorithm, obtaining an enterprise data space of the oil production enterprise, comprising: designing and deploying encryption algorithms at the well site production node, the operation area production node and the headquarters production node respectively, wherein the encryption algorithm comprises: using SM4 algorithm to encrypt real-time sensing data at the well site node, deploying attribute-based SM9 encryption scheme at the operation node, and implementing homomorphic encryption supporting aggregation calculation at the headquarters node; establishing a priori verification mechanism, generating SHA3-256 instruction fingerprint before data transmission, and performing complete transmission after verification pass rate ≥ pass threshold.

[0041] In an optional embodiment, the enterprise data space further comprises: developing a hybrid driven reservoir model, coupling physical equations and LSTM prediction algorithm, satisfying ‖P_data-P_physics‖<0.1 MPa; deploying a CFD accelerated leakage diffusion prediction system with a response time ≤30 seconds; wherein the physical equation solves the pressure field distribution using black oil model, and the input features of the LSTM prediction algorithm include production history data and real-time working condition data.

[0042] In the construction method of enterprise intelligent data space based on big data proposed in the application, multi-modal production data of well site, operation area and headquarters is collected through distributed fiber sensing network, SCADA system and ERP system to build a unified space-time reference oil production data lake; the correlation between wellbore, equipment and process is mined by using graph attention network (GAT) to establish a three-layer knowledge graph including wellbore layer, equipment layer and process layer, wherein the equipment layer is based on API standard to build a failure mode library, and the process layer sets dynamic constraint conditions; a hybrid driven prediction model (coupling physical equation and LSTM algorithm) is developed to realize leakage diffusion detection (reduce response time and prediction error); hierarchical encryption strategy (SM4 at well site, SM9 at operation area, homomorphic encryption at headquarters) is adopted to ensure data security sharing, and finally an enterprise-level data space supporting real-time monitoring, intelligent early warning and collaborative decision-making is formed, which improves the fault diagnosis accuracy, leakage early warning response speed and data utilization efficiency.

[0043] In the description of the present specification, the description referring to the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example.

Claims

1. A method for constructing an enterprise intelligent data space based on big data, characterized in that, Applied to oil production enterprises, the method includes: Acquire multimodal production data of the oil production enterprise, wherein the oil production enterprise includes well site nodes, operation nodes and headquarters nodes, and the multimodal production data includes well site data, operation area data and headquarters data; A production knowledge graph of the petroleum production process is constructed based on the aforementioned multimodal production data; A production prediction model is constructed based on the production knowledge graph, wherein the production prediction model is used to detect whether there is leakage and diffusion in the oil production process; Based on a hierarchical encryption algorithm, the production prediction model and the production knowledge graph are securely shared among the oil production enterprises to obtain the enterprise data space of the oil production enterprises.

2. The method for constructing an enterprise intelligent data space based on big data according to claim 1, characterized in that, The acquisition of multimodal production data from the oil production enterprise includes: The well site data is acquired at a target sampling frequency through a distributed optical fiber sensor network, wherein the well site data includes downhole pressure, temperature, and flow rate parameters. The work area data is obtained from the work area SCADA system through a protocol conversion gateway; The management data is obtained from the headquarters ERP system, and the management data is spatiotemporally aligned with the well site data and the work area data to obtain the headquarters data. Based on the well site data, the operating area data, and the headquarters data, a unified spatiotemporal benchmark oil production data lake is constructed as the multimodal production data.

3. The method for constructing an enterprise intelligent data space based on big data according to claim 1, characterized in that, The construction of a production knowledge graph of the petroleum production process based on the multimodal production data includes: A graph attention network is used to mine the relationships between wellbore, equipment, and processes based on the multimodal production data. A three-layer knowledge graph is constructed, comprising a wellbore layer, an equipment layer, and a process layer. The wellbore layer includes wellbore structural attributes and inter-well topology relationships; the equipment layer includes equipment parameter templates and a fault mode library; and the process layer includes process constraints and production flow rules.

4. The method for constructing an enterprise intelligent data space based on big data according to claim 3, characterized in that, The production knowledge graph also includes: A device failure mode library was established based on the API standard, defining the mapping relationship between vibration characteristics and device failures; Set constraints on water injection process parameters, including pressure range and water quality thresholds; Establish a knowledge reasoning engine to achieve accurate early warning with a fault diagnosis confidence level of ≥0.

92.

5. The method for constructing an enterprise intelligent data space based on big data according to claim 1, characterized in that, The production prediction model and the production knowledge graph are securely shared among the oil production enterprises based on a hierarchical encryption algorithm, resulting in the enterprise data space of the oil production enterprises, including: Encryption algorithms are designed and deployed at the well site production node, the work area production node, and the headquarters production node, respectively. The encryption algorithms include: encrypting real-time sensor data using the SM4 algorithm at the well site node, deploying an attribute-based SM9 encryption scheme at the work area node, and implementing homomorphic encryption that supports aggregation calculation at the headquarters node. Establish a prior verification mechanism. Generate an SHA3-256 instruction fingerprint before data transmission. Perform complete transmission only after the verification pass rate is greater than or equal to the pass threshold.

6. The method for constructing an enterprise intelligent data space based on big data according to claim 5, characterized in that, The enterprise data space also includes: A hybrid-driven reservoir model was developed, which coupled physical equations with an LSTM prediction algorithm, satisfying ||P_data-P_physics|| < 0.1 MPa; A CFD-accelerated leak diffusion prediction system is deployed with a response time of ≤30 seconds; wherein the physical equations are solved using a black oil model to determine the pressure field distribution, and the input features of the LSTM prediction algorithm include historical production data and real-time operating data.